@jssmithi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
Emphasis on probabilities makes Jev natural for something like auto-approval / reviewer escalation. Follow @JasonSteving for more on this.
Started out playing a game with Jev but the capability unlock is serious. It's neat to see how easily it drops into Temporal's experimental Agent Harness. Will share more soon...
Where are all the Temporalites on X? Drop your handle below. I want to follow you. @temporalio
Johann Schleier-Smith retweeted
I am giving a keynote next Tuesday Sept 22 at the "Rows & Columns Summit" on the torrid history of hybrid transaction/analytical processing (HTAP) database management systems: Broken Dreams, Broken Promises, Broken Marriages.
rowsandcolumnssummit.com/
Johann Schleier-Smith retweeted
ok @Muse + 📞 is maybe the worst kept secret in the industry but also maybe my fave Muse feature.
for those starting to get access, send feedback/requests! we're ramping sloooowly as we keep a close eye on quality/reliability.
🚀🚀🚀
We did it again. Temporal just raised a $550M Series E round at a $12.55B valuation.
Learn more about how we got here: temporal.io/blog/temporal-ra…
Johann Schleier-Smith retweeted
Today we're launching the Agents API, a brand new way to build Agents in the cloud, backed by the Codex harness. Bring along all your favorite tools and connectors, connect it to any sandbox, and let Astra cook.
Can't wait to see what you whip up 👨🍳
openai.com/index/introducing…
Johann Schleier-Smith retweeted
My @aiDotEngineer talk about the universal standard we need for clients to drive agent harnesses, and why I think it's Agent Client Protocol
youtube.com/watch?v=YkNulwcc…
Johann Schleier-Smith retweeted
We've sponsored Crystal Palace's men's team since earlier this year. Starting this weekend, we're extending that to @cpfc_w; same club, same commitment.
More on why: temporal.io/blog/temporal-pa…
Johann Schleier-Smith retweeted
I did not expect this. Anthropic just published a Lean proof of Fermat's Last Theorem.
The proof is +13M LoC, more than 5 times the size of Mathlib.
Kevin Buzzard's post: xenaproject.wordpress.com/20…
The proof: github.com/anthropics/fermat…
Anthropic's post: anthropic.com/research/forma…
Johann Schleier-Smith retweeted
AI for cyber is about to go vertical. The models increasingly becoming insanely good at finding and exploiting vulnerabilities. Frontier models are ahead, but we’re already seeing that open weights is not far behind.
Most enterprises are already inundated with cyber discoveries, so now that’s only going to multiply. Triaging and automating the fixes with more AI -along with human oversight- is essentially the only way forward.
In case you were wondering what jobs AI was going to create, it’s definitely going to great time to be in security.
As we prepare to release Astra, we’re focused on making increasingly capable AI safe and broadly accessible.
Astra represents a significant advance in cybersecurity capability, reaching the Critical threshold under our Preparedness Framework.
We're previewing how we evaluated the model, how its safeguards have advanced alongside its capabilities, and what we'll continue to learn and improve.
openai.com/index/path-to-ast…
Extremely proud of this article I wrote for the legendary @lennysan! This is the culmination of months of research into why most AI designs are slop, and my top techniques for fixing this.
LLMs can be super creative, but as we post-train them more and more, we stifle that creativity. We train them to predict tokens that balance everyone’s preferences. But good design does the opposite: it’s risky, bold, and opinionated. You have to put your LLM in a different state of mind to get this kind of design from it. You have to break out of the average slop, into the fringes of the bell curve where unseen ideas live.
This post is all about how to do that. Hope it helps, and let me know if I should go deeper on any of this for future posts!
open.substack.com/pub/lenny/…
“8. Never automate writing-as-thinking. Instead, automate writing-as-reporting: status updates, email summaries, etc. But never outsource the writing you think with: start the doc yourself and end it yourself, using AI in the middle only for research, data, and pushback.” - @tarstarr as summarized by @lennysan
My biggest takeaways from @tarstarr, @OpenAI's ChatGPT Work product lead:
1. The future of work is steering, not rowing. As AI agents take on more of the execution work, humans will shift toward “steering”: making the call for where to go next. In particular, the taste-driven dimension of steering—choosing a direction because you believe the world should look a certain way.
2. “Are you mainlining it yet?” OpenAI’s product culture runs on three internal questions: Are we being as ambitious as possible? Is this maximally accelerated? And are you mainlining it yet (i.e. using your own product all day, every day)? Tara credits the Codex vibe shift over the past few months to this long-held discipline: the team’s user obsession, tight iteration loops, and adjusting quickly once they see how the market reacts.
3. Build for where the models will be in two to three months. Build for current model capabilities, and your product will be outdated by the time it ships. Build for capabilities 12 months out. Tara’s heuristic is to “build for two to three months ahead of the model.”
4. PMs are now in the business of elevating ambition. Tyler Cowen has noted how powerful it is for a leader to look at someone’s work and ask, “Could you do this faster? Could this be 10x bigger?” Tara sees this as a core function of the product role now. When engineers, designers, or stakeholders propose a scope or timeline, a key PM intervention is raising the possibility ceiling: “How could we 10x this? Couldn’t we try this faster?” The OpenAI internal memes (Is this maximally accelerated? Are you mainlining it?) encode the same instinct.
5. Most knowledge work can’t be verified like code, which means it won’t be replaced anytime soon. Coding is output-oriented—you can run tests and see if it works—but with knowledge work, the process itself is how you discover (and trust) the solution. Talking to customers, trying out ideas, seeing the market’s reaction. This is also why it’s important for AI products to surface in-progress work, citations, and chain of thought—so users can go on the journey with the model and actually believe the end result.
6. AI makes clear thinking even more important. Building faster is a gift and a risk. The gift is the ability to iterate at much greater speed. The risk is that you can now travel very far in entirely the wrong direction before anyone notices. If ideation and hypothesis quality do not keep pace with execution speed, teams “blow off course way quicker” than they ever would have before. Speed without a clear hypothesis will just compound errors faster.
7. Empirical beats theoretical. At Stripe, Tara spent tens of hours writing rigorous strategy documents because the market was established enough to reason from first principles. At OpenAI, the market changes too fast for a 12-month roadmap to be meaningful. The right response is to move from academic to empirical: identify your sharpest hypothesis, then test it with users as quickly as possible. Long reasoning documents rarely make sense anymore. Instead, get to something real people can try as quickly as you can.
8. Never automate writing-as-thinking. Instead, automate writing-as-reporting: status updates, email summaries, etc. But never outsource the writing you think with: start the doc yourself and end it yourself, using AI in the middle only for research, data, and pushback. Share docs at 70% complete so collaborators can poke holes and polish with you. At OpenAI it’s now “mocks, not docs”—prototypes and A/B results communicate better than long documents, because AI has made a long doc a meaningless signal of rigor. Tara still writes hundreds of docs—but for herself, not as the shareable artifact.
9. Someone still has to be the DRI, even when roles dissolve. Tara has always liked almost no boundaries between engineer, PM, and designer. Everyone can pick up the work now. But someone needs to be accountable. Someone still has to own the outcome.