@jssmith
San Francisco, CA
Joined May 2008
Harnesses matter. The green dots are the new frontier.
Today we're excited to share Strands Harness. Strands Harness makes it easier to build reliable, cost-effective, agents with any model. It's open source, too! And equal or better performance to proprietary harnesses with 28% fewer tokens!
145
Emphasis on probabilities makes Jev natural for something like auto-approval / reviewer escalation. Follow @JasonSteving for more on this.
Jev powered "auto mode" in the Temporal Agent Harness is such a natural fit. Jev is a big unlock for efficient agent oversight. Very excited to see how straightforward it was to integrate along existing harness seams. More to come here for sure.
1
6
1,479
Started out playing a game with Jev but the capability unlock is serious. It's neat to see how easily it drops into Temporal's experimental Agent Harness. Will share more soon...
3
1
20
2,129
Where are all the Temporalites on X? Drop your handle below. I want to follow you. @temporalio
3
3
181
I am giving a keynote next Tuesday Sept 22 at the "Rows & Columns Summit" on the torrid history of hybrid transaction/analytical processing (HTAP) database management systems: Broken Dreams, Broken Promises, Broken Marriages. rowsandcolumnssummit.com/
1
19
3
134
9,478
Johann Schleier-Smith retweeted
have feedback on @muse phone calls? send them to @jrlevine!
ok @Muse + 📞 is maybe the worst kept secret in the industry but also maybe my fave Muse feature. for those starting to get access, send feedback/requests! we're ramping sloooowly as we keep a close eye on quality/reliability.
32
3
1
110
25,290
Johann Schleier-Smith retweeted
Huge day for us. Thank you @NYSE
4
7
1
154
5,979
🚀🚀🚀
We did it again. Temporal just raised a $550M Series E round at a $12.55B valuation. Learn more about how we got here: temporal.io/blog/temporal-ra…
1
1
10
1,302
Johann Schleier-Smith retweeted
Today we're launching the Agents API, a brand new way to build Agents in the cloud, backed by the Codex harness. Bring along all your favorite tools and connectors, connect it to any sandbox, and let Astra cook. Can't wait to see what you whip up 👨‍🍳 openai.com/index/introducing…
153
279
100
2,835
2,002,347
Johann Schleier-Smith retweeted
My @aiDotEngineer talk about the universal standard we need for clients to drive agent harnesses, and why I think it's Agent Client Protocol youtube.com/watch?v=YkNulwcc…
1
8
2
17
7,630
Johann Schleier-Smith retweeted
We've sponsored Crystal Palace's men's team since earlier this year. Starting this weekend, we're extending that to @cpfc_w; same club, same commitment. More on why: temporal.io/blog/temporal-pa…
4
22
1,858
Johann Schleier-Smith retweeted
AI for cyber is about to go vertical. The models increasingly becoming insanely good at finding and exploiting vulnerabilities. Frontier models are ahead, but we’re already seeing that open weights is not far behind. Most enterprises are already inundated with cyber discoveries, so now that’s only going to multiply. Triaging and automating the fixes with more AI -along with human oversight- is essentially the only way forward. In case you were wondering what jobs AI was going to create, it’s definitely going to great time to be in security.
As we prepare to release Astra, we’re focused on making increasingly capable AI safe and broadly accessible. Astra represents a significant advance in cybersecurity capability, reaching the Critical threshold under our Preparedness Framework. We're previewing how we evaluated the model, how its safeguards have advanced alongside its capabilities, and what we'll continue to learn and improve. openai.com/index/path-to-ast…
61
71
14
531
158,359
Johann Schleier-Smith retweeted
Extremely proud of this article I wrote for the legendary @lennysan! This is the culmination of months of research into why most AI designs are slop, and my top techniques for fixing this. LLMs can be super creative, but as we post-train them more and more, we stifle that creativity. We train them to predict tokens that balance everyone’s preferences. But good design does the opposite: it’s risky, bold, and opinionated. You have to put your LLM in a different state of mind to get this kind of design from it. You have to break out of the average slop, into the fringes of the bell curve where unseen ideas live. This post is all about how to do that. Hope it helps, and let me know if I should go deeper on any of this for future posts! open.substack.com/pub/lenny/…
28
21
9
223
12,573
“8. Never automate writing-as-thinking. Instead, automate writing-as-reporting: status updates, email summaries, etc. But never outsource the writing you think with: start the doc yourself and end it yourself, using AI in the middle only for research, data, and pushback.” - @tarstarr as summarized by @lennysan
My biggest takeaways from @tarstarr, @OpenAI's ChatGPT Work product lead: 1. The future of work is steering, not rowing. As AI agents take on more of the execution work, humans will shift toward “steering”: making the call for where to go next. In particular, the taste-driven dimension of steering—choosing a direction because you believe the world should look a certain way. 2. “Are you mainlining it yet?” OpenAI’s product culture runs on three internal questions: Are we being as ambitious as possible? Is this maximally accelerated? And are you mainlining it yet (i.e. using your own product all day, every day)? Tara credits the Codex vibe shift over the past few months to this long-held discipline: the team’s user obsession, tight iteration loops, and adjusting quickly once they see how the market reacts. 3. Build for where the models will be in two to three months. Build for current model capabilities, and your product will be outdated by the time it ships. Build for capabilities 12 months out. Tara’s heuristic is to “build for two to three months ahead of the model.” 4. PMs are now in the business of elevating ambition. Tyler Cowen has noted how powerful it is for a leader to look at someone’s work and ask, “Could you do this faster? Could this be 10x bigger?” Tara sees this as a core function of the product role now. When engineers, designers, or stakeholders propose a scope or timeline, a key PM intervention is raising the possibility ceiling: “How could we 10x this? Couldn’t we try this faster?” The OpenAI internal memes (Is this maximally accelerated? Are you mainlining it?) encode the same instinct. 5. Most knowledge work can’t be verified like code, which means it won’t be replaced anytime soon. Coding is output-oriented—you can run tests and see if it works—but with knowledge work, the process itself is how you discover (and trust) the solution. Talking to customers, trying out ideas, seeing the market’s reaction. This is also why it’s important for AI products to surface in-progress work, citations, and chain of thought—so users can go on the journey with the model and actually believe the end result. 6. AI makes clear thinking even more important. Building faster is a gift and a risk. The gift is the ability to iterate at much greater speed. The risk is that you can now travel very far in entirely the wrong direction before anyone notices. If ideation and hypothesis quality do not keep pace with execution speed, teams “blow off course way quicker” than they ever would have before. Speed without a clear hypothesis will just compound errors faster. 7. Empirical beats theoretical. At Stripe, Tara spent tens of hours writing rigorous strategy documents because the market was established enough to reason from first principles. At OpenAI, the market changes too fast for a 12-month roadmap to be meaningful. The right response is to move from academic to empirical: identify your sharpest hypothesis, then test it with users as quickly as possible. Long reasoning documents rarely make sense anymore. Instead, get to something real people can try as quickly as you can. 8. Never automate writing-as-thinking. Instead, automate writing-as-reporting: status updates, email summaries, etc. But never outsource the writing you think with: start the doc yourself and end it yourself, using AI in the middle only for research, data, and pushback. Share docs at 70% complete so collaborators can poke holes and polish with you. At OpenAI it’s now “mocks, not docs”—prototypes and A/B results communicate better than long documents, because AI has made a long doc a meaningless signal of rigor. Tara still writes hundreds of docs—but for herself, not as the shareable artifact. 9. Someone still has to be the DRI, even when roles dissolve. Tara has always liked almost no boundaries between engineer, PM, and designer. Everyone can pick up the work now. But someone needs to be accountable. Someone still has to own the outcome.
1
137