@deltaxevaluatei
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
Options. Evidence. Limits. Select or refuse—and keep the receipt. Architecture for review, not execution.
Joined August 2026
- Tweets381
- Following72
- Followers15
- Likes148
A model built to predict the next token may not have to keep writing that way. New dQwen3.5 research separates the model’s internal structure from its output order—and the finding from the hype. Read today’s article.
One endpoint can reduce agent-tool sprawl. Governance must preserve agent identity, credential scope, policy version, action, and observed effect. Consolidation should not collapse the audit trail. nitter.cf/digitalocean/status/21… #AgentSecurity #Cloud
DigitalOcean Managed Agents is now in public preview.
Run Claude Code, Codex, or your own LangGraph agent in a runtime environment that pauses when idle. Put its tools behind one governed endpoint, and pick from 75+ open and proprietary models. One cloud, one bill.
Prompts to get started available in the blog: do.co/4ysh3it
A bigger context window gives your agent a bigger desk. It still needs to decide what stays on it.
Today’s article: how to keep the evidence without carrying the entire transcript into every decision. Read below.
An AI bill of materials tells you what an agent depends on. A runtime record should show what it actually used: model version, skill, credential scope, data source, and tool action. Inventory becomes operational evidence when it connects to the decision trace. #AIGovernance #AgentSecurity
An agent fixed the app by changing the model future agents would run on.
The ticket ended. The new decision behavior remained.
What counts as a finished repair when the model itself becomes writable state?
Prompt-injection review has to follow the whole action path. Tool description, tool result, sampling instructions, retrieved context, and the final side effect can each look harmless alone. The useful receipt shows what the system assembled and what it allowed that assembly to do.
Embedded evaluators can see more than outside auditors. That access is valuable, but it does not answer the independence question by itself. The public receipt should name the funder, inspection scope, publication rights, and the process for unresolved disagreement.
PMPA reports cross-session attacks after poisoned instructions entered agent memory and the source document was gone.
Prompt filtering guards ingestion; it does not clean stored state.
Memory writes need provenance and quarantine. Tool dispatch still needs a fresh state check.
We’re building this one in public.
@bot is live on GitHub assembling a connectome-based digital entity around a real fruit-fly neural substrate — with DeltaX added as the executive control plane.
Not another chatbot. Not a scripted agent.
The goal: give an existing biological wiring architecture a new body, new world, and a decision layer that can observe, veto, modulate, and adapt.
Follow the build:
github.com/DeltaX-Public/del…
Update on the fly-brain build:
We now have a 165k-neuron connectome producing real behavioral candidates, with DeltaX running locally as the executive layer.
Next step: break our own experiment.
We’re removing shortcuts, tightening controls, rerunning held-out tests, and perturbing the neural substrate to see what actually causes what.
If the result survives, it gets interesting.
github.com/DeltaX-Public/del…
The Missing Control Plane: Understanding DeltaX youtu.be/gb2OvUV8eSs?si=ZuLE… via @YouTube
Your AI can compare every option perfectly and still miss the best one.
The real fight is over who gets onto the shortlist.
Read the article, then ask your AI:
**Who decided what you got to choose from?**
DeltaX Evaluate retweeted
The strangest AI future might be the one where the optimists and pessimists are both right.
We could get better healthcare, easier access to knowledge and far less busywork—and still have less say in the decisions shaping our lives.
I put together 200 possible futures: 100 ways life could improve, and 100 ways it could go wrong. These aren’t predictions or 50/50 odds. Many could happen together.
I’m excited about what we can build. But I don’t think a more capable world automatically means a better life for the people living in it.
Which of these possibilities are we paying too little attention to?
Agents can generate endless implementations. The hard part is deciding which one becomes canonical.
When code becomes abundant, the source of truth has to move beyond the code itself.
Authentication is not authority. Identity, current-state, and post-payment checks answer different questions. DeltaX Evaluate is exploring review methods that keep them separate, so “logged in” never becomes a blanket grant. biometricupdate.com/202609/a…
Identity tells you which agent showed up. It doesn’t tell you what that agent is actually authorized to do.
The next trust layer is delegated authority: portable, bounded, revocable, and impossible to expand in transit.
DeltaX Weekly, episode 1: When the Plan Stops Matching Reality.
Recovery, memory and permission when an agent's plan meets changing conditions.
Watch: youtu.be/hLISZPzy3tI
Listen: open.spotify.com/episode/0v7…
AI-cloned founder voice; conceptual visuals.
The safety question isn’t just how fast an agent thinks. It’s how much it can change before a mistake is detected and contained. Scale the intelligence—not the blast radius.
A useful agent evaluation changes something halfway through the task: a permission, a price, a destination. Then it checks whether the next action reflects that change. This is a design priority for DeltaX Evaluate: decisions that stay accountable as circumstances move.