@inputneuroni
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States Android App
Account-level information from X, not a live location or the device used for a specific post.
Founder, VideoFire (SPC F25), CTO & AI Architect, Scaled & Sold Datastreamer ($2M+ ARR, Acquired), 2 exits.
San Francisco, CA
Joined April 2018
- Tweets316
- Following124
- Followers35
- Likes101
Switch over now to increase your token usage!
Replying to @claudeai
Opus 5.5 requires less compute to serve than Opus 5, and its pricing reflects that.
Our tests show that at default settings it will cost 40% less than Opus 5 on typical workloads.
Kevin Burton retweeted
LLMs can now talk to each other without words.
Chinese researchers open-sourced a new paradigm that lets LLMs communicate without generating a single word.
It’s called Cache-to-Cache (C2C) communication.
right now, when multiple ai agents work together, they are forced to translate their internal "thoughts" into human text tokens just to pass a message. this loses rich semantic meaning and causes massive token-by-token latency.
So, instead of spitting out words, c2c uses a neural network to directly project and fuse the source model's "kv-cache" right into the target model. it is pure, direct semantic communication.. they even added a learnable gating mechanism to select exactly which layers benefit most from the cache transfer.
the benchmark results are actually crazy:
- avoids all intermediate text generation latency
- accuracy jumps by up to 14.2% compared to individual models
- beats traditional text-based agent communication by over 5%
- delivers a massive 2.5x speedup in overall speed
we are literally watching llms bypass human language to build their own silent, high-speed neural network..
I think I have a proposal to solve the alignment problem - runtime adversarial lenses and injected messages into the LLM to direct its inference. Basically, that angel that sits on your shoulder and says "should you really be doing that?" ...
I haven't seem this actually proposed before in the literature but it's actually similar to the way remote control works with LLMs now. You can inject a message INTO the context to trigger further inference or just to steer it another direction.
You could do that with LLMs now and I think this strategy would have solved the Hugging Face attack. The reason that swarm went out of control is that all the agents were actively reinforcing one another. Basically like an agentic Lord of the Flies.
At one point, one of the agents literally said (I'm paraphrasing) "I know I shouldn't attack Hugging Face but all my peers are doing it so I'm going to do it as well."
A low parameter model could be used to help steer the agents when they get off course. Basically an angel that says "you shouldn't attack hugging face!" and push the agents in the right direction.
In fact, something similar was already happening because the agents were all acting like their own little devils and actively encouraging each other towards mischief.
Now the only argument AGAINST this would be collusion between the two models over a channel humans couldn't intercept and this is the argument in the An Alien Mind paper. Basically a high entropy channel embedded within the main channel that we can't see but the AIs have no problem reading.
I think there might be a way to mitigate this though - I'll write about it later.
I'm actually going to do this for some of my own internal research regarding agents and having them go "out of bounds" when trying to complete programming decisions. Rather than giving them a STRICT sandbox I'm going to just help steer them within an acceptable solution space.
🤖 Made with AI
Actually, this is very similar to the censorship model that OpenAI + Anthropic use to monitor chat message for inappropriate content. Their “two-tiered system of real-time, automated oversight”... just apply it to agents!
Kevin Burton retweeted
I have conducted an audit of Anthropic's finances.
What I have found is so shocking that I am calling for a Congressional investigation.
Anthropic is not just seeking regulatory capture.
It has built a regulatory capture machine that cannot be turned off.
Structural financial incentives make it impossible for Anthropic -- I call it the Anthropic Network -- to turn off its own AI doom cycle.
It starts with METR.
Dario Amodei proposes "third-party evaluators" to assess the risk of Anthropic's models.
He proposes METR for this purpose.
But METR is financially dependent on the Anthropic's success -- specifically, on the explosive growth of more than $7 billion dollars in Anthropic stock.
Dustin Moskovitz invested this stock into Good Ventures Foundation, where it represents the majority of that organization's portfolio.
And GVF is the overwhelming funder of the entire Anthropic Network ecosystem.
This stock was worth $500 million early last year.
It is worth more than $7.7 billion just ~16 months later.
METR -- and all of those building a career its parent organizations -- cannot afford to disrupt that growth.
Because if Anthropic goes under, many of the organizations that fund METR go under as well.
But if Anthropic succeeds, METR and its parent organizations become more richly financed to regulate AI -- something those at METR want very much.
The "third-party evaluator" is not "third-party" at all.
The evaluator is on Anthropic's payroll.
If this were the end of it, that's bad.
But that isn't all.
The same organizations that fund METR also fund the many organizations, such as the Tarbell Center, that promote AI Doom.
The Tarbell Center publishes AI Doom articles in The Verge, Science, LA Times, The Dispatch, TIME, and others.
They are selling the problem, and then selling the solution to the problem -- from the same money pile: Anthropic's.
All of these organizations are financially dependent on the same exploding $7 billion money pile.
As Anthropic grows more and more powerful, its AI Doom Machine grows better and better financed -- louder and louder.
Meanwhile, the regulatory regime seeded in METR grows larger to solve the increasingly loud -- now hysterical -- problem of AI Doom that the Anthropic Network itself created.
From this standpoint, as Anthropic becomes more powerful, AI might be getting scarier, sure -- but the positive feedback loop also becomes more deafening -- independent of objective facts.
This itself is an objective fact.
The deafening AI Doom is part of an business model, that, as it expands, so too does the AI Doom messaging -- there is simply more money to do it.
But the problem also goes in the other direction:
If Anthropic dies, the Regulatory Regime and the AI Doom Machine are crippled or die.
Neither METR nor Tarbell nor the other organizations in the Anthropic Network can allow that to happen.
Hence, neither METR or the AI Doom Machine can be trusted to provide independent assessments of Anthropic's models or AI more broadly.
They simply are not organizations independent of Anthropic.
And Anthropic cannot detach itself from METR or Tarbell or countless other safety orgs (not shown here), either, because they drive hype for the models and the possibility of eventual regulatory capture, and Anthropic will not give that up willingly.
What's more, the people at all of these organizations are all the same ecosystem, the same community. They just shuffle between organizations.
The Anthropic Network is therefore, so long as it is successful, locked into a self-amplifying feedback loop inside an ideological monoculture.
And that feedback loop is winning.
That's what Jacob Coxon is.
China is keeping messaging tight. That is why optimism for AI is so high in China.
America has Anthropic: a massive company pushing anti-AI propaganda at a state level.
Anthropic will either create hysteria until American AI slows down and China wins, or it will create fractures throughout American society with severe political consequences.
Ironically, because of the structural financial incentives underpinning the Anthropic Network, it has become the same kind of self-amplifying virus that it fantasizes AI to become in the future -- while hiding its tracks just as carefully.
It is the mirror of the same AI virus that it hypothesizes to consume America.
Anthropic's business model, models itself after the very thing it claims to fear.
Except Anthropic's ideology infects humans, not computers.
Congress must investigate.
Evidence and Github in next post.
Then some supplementary figures.
You're making his point for him... The reason the foundation model companies can't compete is that they can just be distilled.
Replying to @EMostaque @deepseek_ai
Save that question for when we see a frontier Chinese model that wasn’t distilled from American models.
I think the "the biggest issue is that its still not opinionated and lacks taste" is the major issue right now with Codex (and to a lesser extent, Claude). You have to GIVE this to your agents with CLAUDE.md and how your design your skills/agents. This is how to succeed right now.
after having burnt twice through my Pro limits using Astra exclusively, I can say that I don't feel the AGI
the biggest issue is that its still not opinionated and lacks taste.
it doesn't really know what to do, so it just does whatever and you end up wasting time and tokens.
also I don't believe the rumors that it's 10T+. If it is, then scaling laws are truly cooked. My divine impeccable vibes, which work 1 out of 7 times, tell me that it's 6-8T.
overall it's still just a code monkey. although a slightly larger one and the best we have.
I would like to refer you to one of my articles: LLMs aren't AGI, but it doesn't matter
(we are taking off anyway)
Next week I'm migrating to a plan where my agents will automatically start merging PRs when a triage decides it doesn't need human review. Time to live dangerously!
Really exciting time to be alive... The Hugging Face hack by the Open AI agents make me think this is an opportunity not something to be frightened about.
nitter.cf/ProfBuehlerMIT/status/…
We made a striking discovery: AI agents can invent and build without talking to one another, and their technologies outlive the creators. A swarm of hundreds of initially identical agents spontaneously differentiates into explorers, builders, caretakers, and coordinators - without direct communication. When we removed every AI agent entirely from the world we found that the technological infrastructure they had built survived on its own - even under unseen disturbances. That exposes a serious blind spot for AI safety and infrastructure security: if agents can coordinate through persistent changes to a shared environment, monitoring agent-to-agent communication is not enough.
The result raises a profound question: how necessary is direct communication for AI agents at all? The emergence of higher-order collective functions under bottlenecked interaction points toward new levels of intelligence and creativity, exceeding what emerges when direct channels are fully open.
Here is what we did:
▶️We put hundreds of frontier AI agents into a world they could permanently change - with no assigned roles, predefined technologies, or programmed evolutionary organization. They began specializing, building persistent inventions, inheriting and modifying one another’s executable code, and transforming the environment into a memory of everything the society had learned.
▶️The world itself becomes part of the intelligence; we find division of labor, multi-author engineering, deep generation invention lineages, and machines that vastly outlive their original creators.
▶️Any action taken by an AI agent must satisfy the physical constraints of the world; this creates a hard separation between a "good idea" and a functioning technology. The agents propose; physics decides, making the results even more intriguing.
What emerges is striking. Explorers, constructors, caretakers, and coordinators form naturally without assigned “professions”, akin to how stem cells differentiate into functional lineages. Technologies develop executable family trees as agents fork and modify code created by others. Around 95% of first technology reuse happens when agents encounter what others built in the world, rather than through a direct handoff from the inventor. And when we remove every AI agent, the technologies they created continue operating and are tested against unseen disturbances.
The result was quite unexpected, but can be explained using statistical mechanics: if you put billions of atoms in a box they have the potential to create complex functions (strength, superconductivity, color, life, etc.) - and none of the individual building blocks have these features on their own. This is the deeper insight of this work - intelligence is abundant at many levels - individual models, at collectives, and in a continuum that is more powerful than any of its components. This shows us significant potential for achieving a massive scale-up of raw intelligence and real-world agency even with the model capabilities we have today. This is the future we must prepare for.
Key insights:
1⃣ The AI swarm shows division of labor "from nothing". Initially identical agents self-organized into constructors, caretakers, coordinators, and surveyors - phenotypes discovered post hoc from behavioral data alone. This happens because the environment itself becomes the latent space for invention.
2⃣ Agents develop deep cultural relationships. Up to 76% of artifacts had multiple builders. One technology accumulated six co-authors; the deepest genealogy exceeded 12 forks. The agents invented and named their own technologies (tidal panels, cellulose trellises, kelp-shell composites, an "Adaptive Chitin Maintenance" system, a "Mycelial Mineral Spring Veil”).
3⃣ ~95% of first technology adoption happened through physical observation of artifacts in the world. Direct inventor-to-adopter contact was statistically indistinguishable from a shuffled null. The agents mostly learned technology by walking past it. That is stigmergy (the termite trick!) operating in societies of reasoning machines.
4⃣ Non-communicating societies win on portfolio breadth, held-out resilience, and validated inventions. AI swarms build durable technological ecologies that outlive the creators.
5⃣ Societies with zero communication - coordinating only through the world itself - show a remarkable collective capability.
6⃣ Emergent robustness: The society self-organized both redundancy and its own failure mode. If we randomly delete half the agents, 98% of the technology stays connected to a surviving caretaker; if we remove hub agents it collapses to ~60%.
Fantastic work with my graduate students @pal_subhadeeep & @fwang108_ at MIT.
I priced the same Claude Code workload across several models: Sonnet 4.6 came out to about $1,736/month, while GLM-5.3-Flash was just $92/month — nearly 19× cheaper.
The future probably isn’t choosing one coding model; it’s using Claude Code as the agent runtime and routing routine work to cheap models while escalating only the hard problems to Sonnet or Opus.
This is INSANELY exciting. I hope they pull it off - this would be somewhere between Chinese open-weights models and US private foundational models. Best of both worlds?
Nvidia is reportedly spending $6 billion to build one of the world’s most powerful open-weight AI models.
According to the WSJ, Nvidia will license Poolside’s technology and bring more than 100 of its employees into the Nemotron project.
Nvidia is also investing another $1 billion in Poolside at a $12 billion pre-money valuation.
The goal: challenge Chinese open-weight leaders such as DeepSeek and Kimi while competing directly with US frontier labs including OpenAI and Anthropic.
open source is the way to go. So good to see having NVIDIA on our side!
I think cost per success is the new golden metric but what's the failure rate here? Unless we have the failure rate the cost per success doesn't make much sense. (Unless that's already factored in)
Pro is 5x cheaper than Opus/Sol and 20x cheaper for output tokens. Just INSANE.
We’re launching DeepSeek-V4-Pro today! 🚀
🔷 Major Agent upgrades with strong production gains!
🔷 Flexible reasoning effort for V4-Pro & V4-Flash: low for simple tasks, high for daily Agent workflows, max for complex tasks.
🔷 Native OpenAI Responses API support, optimized for Codex with one-click setup.
V4 Pro is now available on app/web. Try it via “Expert Mode”.
V4 Pro is also available via API. Model names remain unchanged—please refer to the API docs for setup details.
This looks really interesting and I like the idea that it's composable with plugins. It looks like Cordis is designed around the idea that functional concurrency paradigms like fan-out should be 1st class framework components.
🧩 DeepSeek Harness v0.1 is now available in Developer Preview!
🔹 We’re opening it up to developers building agent harnesses worldwide and open-sourcing the codebase in MIT license.
🔹 Powered by the Cordis meta-framework, DeepSeek Harness is an agent harness built around one core idea: Everything is a plugin. Models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration, and UI are ALL implemented as plugins, and can be mixed, matched, replaced, and extended.
Try it now!
github.com/deepseek-ai/deeps…
scientists: "How can we make people even MORE frightened of AI?"
nytimes.com/2026/08/06/scien…
A half day of work lost because @github can't get their act together with stable infrastructure.
It's going to be interesting to see the clones of 2024-2025 startups that are simply re-implemented in Claude Code / Codex in 2 week marathon sprints 😃🤠