nano bio.md

California, USA
Joined May 2026
joseph retweeted
🚀 Introducing XOR An open source, multimodal Jev-like decision model. Being developed with data sovereignty & enterprise grade in mind. ❶ Multimodal: great for text + visual classification ❷ Much bigger 260k context window ❸ Qwen3.6 based huggingface.co/juspay/xor ↓↓↓ See it in action
17
101
16
667
40,276
joseph retweeted
Dario wrote a 3800 words article on pacing the frontier only to drop a new SOTA model 11 days later. Safe to say, no one is actually doing it.
83
292
33
7,152
196,927
Just watched a frontier LLM debug a website by literally turning my wifi off and on - but ofc it couldn’t actually turn the wifi back on because its connection was severed
156
196
52
10,846
217,120
Replying to @github
1
1
1,065
33,360
joseph retweeted
Is that like the Brad Pitt of you virgins?
11
6
4
495
39,240
state of the art motherfuckers
747
740
249
21,982
2,036,179
not now honey, I'm doing frontier AI research
48
110
18
1,922
101,584
joseph retweeted
636606729769440499166579950236036751749912014371509557713570027508971809534551913252252094954941974952859310861988904737359709200557919 is a factor of RSA-896 saweis.net/posts/rsa-896.htm…
225
990
322
9,510
3,902,515
joseph retweeted
Introducing Step 5 Preview: Advancing the Pareto Frontier. Step 5 Preview is our new flagship model for agentic work, delivering frontier-level performance across software engineering and professional knowledge work, with particular strength in finance. - 600B total / 27B active MoE, with 1M context + Vision - Substantially lower task cost at comparable intelligence - Broad software engineering capabilities with sustained execution over long horizons Try Step 5 Preview: platform.stepfun.ai Model page: stepfun.com/step-5-preview Open weights on Oct 15.
200
280
163
1,876
433,012
I got early access to @typesafeai’s Jev—the new “System One” model that doesn’t generate text. I tested the live API. Headline: 50 semantic judgments in 226 ms. One state, one request, all answers together. The parallelism looks real. 🧵
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
1
51
Where this is useful: • route tickets/events/docs • score risk, urgency, relevance • audit agent traces and claims • gate cheap model → expensive model → human • monitor huge streams and wake an agent only when a semantic condition hits
1
4
The intelligence looks good but isn’t independently proven. The calibration needs scrutiny. But the interface + speed are real. “Stop hiring a novelist for every if-statement” is the best explanation I have. Live model tested was `jev-1.13.0`
8
Last month I wrote about how we can build a positive and safe future for everyone: meta.com/thefutureisforevery… Every lab has the responsibility and incentive to move at the pace required to train its models safely, and the ability to take its own actions to ensure that happens. The reality is: - People won't want to use agents that are misaligned with them and that don't do what they ask, so labs have a strong natural incentive to make their models more aligned. There is a lot of debate about slowing progress on capabilities until alignment catches up. My view is that trust and alignment are quickly becoming the most important capabilities that will differentiate agents and models. Any lab that doesn't focus on alignment will fall behind. - Labs face significant liability if their models cause harm, so they have a strong incentive to prevent this as well. Meta delayed shipping Muse for several months to focus on safety and security. We didn't call for everyone else to do this before we would. We just did it as part of our day-to-day work because it was clearly the right thing for people and for us. I'm proud of the security foundations we've built. - Engaging independent evaluators and advisors is industry best practice. MSL already does this today in several areas because it helps produce better work. Other labs can just do this too. In general, it would be helpful for there to be a larger and more diverse ecosystem of evaluators. - Committing the significant majority of compute towards serving people rather than racing towards recursive self-improvement is one of the best ways to ensure we develop this technology safely. Meta has made this commitment and other labs can do this as well. I believe the key to building a positive future for everyone is maintaining the right balance of power. This is within our power to do.
1,927
2,858
1,545
28,463
7,715,748
Dario has written that we need to “pace the frontier,” and Sam has agreed. People may be surprised by my response: go ahead. You guys are the frontier. By any reasonable metric — market share, revenue growth, model capability — the two of you have a duopoly on frontier intelligence. You’ve also claimed the lead is widening because of recursive self-improvement. I don’t see what you see in the lab. If the unreleased models are scary enough that you think you should slow down, I support your decision to be responsible. But stop pretending you need anyone else’s permission. Stop pretending antitrust law has to be suspended so you can form a cartel. Stop pretending you need a regulatory approval process that supersedes product liability. Stop pretending METR is independent when it is intertwined with Anthropic’s investors and staff. Stop pretending you need those same evaluators to police competitors who aren’t even at the frontier. Most of all, stop pretending the motivation to slow down is purely altruistic. You face massive product-liability exposure if your products enable a truly damaging cyberattack. The market already punishes models that behave in unpredictable or unauthorized ways. After the Hugging Face episode, it is simply good business for OpenAI and Anthropic to trade some raw power for reliability and predictability. Call it alignment if you want. It is also just giving customers what they want. Pacing the frontier would also create breathing room for a more intelligent conversation about regulation than Bernie Sanders’ “shut it all down.” China is very unlikely to join a global agreement, as you know, and that has to be taken into account as well. So go ahead and pace the frontier. You are the ones setting it. The easiest way not to build superintelligence is for you to agree not to build it. Demanding your preferred regulatory framework as the price of that will look like blackmail of the public and the political system. So just do it. If you do, you’ll buy goodwill for the next conversation. If you don’t, we’ll know this was just another bid for regulatory capture — or an election-season psyop.
3,431
12,353
2,702
72,183
9,578,427