@chiproi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- Web
Account-level information from X, not a live location or the device used for a specific post.
@aisysbooks @goodailist AI Engineering: https://nitter.cf/t.co/94dv4uTU1H Designing MLSys: https://nitter.cf/t.co/G81hL2dWmr Reading @chipslib
San Francisco, CA
Joined June 2008
- Tweets587
- Following733
- Followers15K
- Likes7.2K
Pinned Tweet
My 8000-word note on agents: huyenchip.com//2025/01/07/ag…
Covering:
1. An overview of agents
2. How the capability of an AI-powered agent is determined by the set of tools it has access to and its capability for planning
3. How to select the best set of tools for your agent
4. Whether LLMs can plan and how to augment a model’s capability for planning
5. Agent’s failure modes
AI-powered agents are an emerging field with no established theoretical frameworks for defining, developing, and evaluating them. This post is a best-effort attempt to build a framework from the existing literature, but it will evolve as the field does.
As always, feedback is much appreciated!
Interesting approach. Models can't output freeform text but can choose from a set of predefined values. Could be useful for data labeling and tasks with a fixed list of possible actions.
Unclear how reasoning would work though, but it's super cheap (output tokens are free!)
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution
i'm researching best practices for skills.
what's the most number of skills you've had installed for your agents?
16%0 / what's a skill?
36%1 - 5
17%6 - 10
30%> 10
820 votes • Final resultswhat's a good model tiering system? i'm sick of telling my agent orchestrator things like: "for Claude, use model X, for OpenAI, use model Y, etc."
i want to be able to tell my orchestrator: "use models tier ..." for this kind of task
that's the problem he should've sent them in all caps
We asked an unreleased research version of Claude to take a stab at the Riemann hypothesis.
It didn’t solve it, but it did make strides on a related problem: it increased the lower bound for the fraction of zeros of the Riemann zeta function that satisfy the hypothesis from 41.6% to 67.2%.
anthropic.com/research/riema…
why don't model providers price tokens like electricity, higher price during peak hours and cheaper off peak?
Auto science is gonna be 🔥
Congrats anh Quoc and team!
Excited to co-found Discovery Loop with my long-time collaborators @JeffDean @Sanjay_Ghemawat @OriolVinyalsML . Our mission is to automate machine learning, engineering and science. Learn more at: discoveryloop.com ♾
Chip Huyen retweeted
Our first model, Inkling. Trained from scratch, weights are open, fine-tunable on Tinker today.
Today, we are introducing Inkling.
Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available.
thinkingmachines.ai/news/int…
Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵
Congrats to @AlecRad @Luke_Metz and @soumithchintala!
It's really cool to see work produced by a group of 20-somethings without a single PhD between them winning this award.
I hope to see these three collaborate again one day :)
We are honored to announce the Test of Time awards for #ICLR2026 🏆 This award recognizes papers published 10 years ago at ICLR 2016 that have had a lasting impact on the field:
blog.iclr.cc/2026/04/22/anno…
Chip Huyen retweeted
Got to meet the wonderful Chip Huyen @chipro
She’s so nice and smart!!
Chip Huyen retweeted
Excited to release PostTrainBench v1.0!
This benchmark evaluates the ability of frontier AI agents to post-train language models in a simplified setting.
We believe this is a first step toward tracking progress in recursive self-improvement 🧵:
How long do you think AI will be able to fully automate your job?
11%Already automated
41%<= 3 years
29%> 3 years
19%Never
2,306 votes • Final resultsI built GoodAIList.com to help me stay-up-date with new AI stuff.
It's tracking 14K open source repos so far, with contributions from over 145K developers.
Every day, it:
- searches for new AI repos (based on 123 keywords and topics)
- surfaces repos that are gaining traction, and
- categorizes each repo
The annotations are done by AI so they are not super accurate, but they've helped me find some useful stuff.
It also lets me see where the contributors are, so when I travel, I can find folks doing cool stuff in a new city or country.
Super impressed by the projects at the Agentic Hackathon last weekend! Many teams work on really hard/important problems:
* Long running tasks: memory management, recovering from mid-task failures, and maintaining consistency across steps and sub-agents
* Adaptive retrieval from multiple sources: databases, search indices, and websites
* Agents that work with voice, video, and even 3D environments
If you are in SF, come check out the finalist demos tomorrow! luma.com/6bd4bt9j
There will be talks by Douglas Eck, who is doing amazing work with Veo and Imagen and many other awesome folks.
Thanks @MongoDB and @cerebral_valley for hosting and for letting me serve as a judge for these fantastic projects.
After years of following @lennysan's wonderful takes on product, I finally had the opportunity to chat with him about AI products!
youtube.com/watch?v=qbvY0dQg…
1. Many AI product problems aren’t because of AI. It’s usually because of user experience, data quality, or organizational structure.
A chatbot failed to get traction because their targeted users simply couldn’t type (because their hands were usually busy -- taking care of kids or driving), so showing pre-populated questions and adding a voice option significantly improved traction.
Another team told me their lead scoring model was broken. It turns out that it’s because the marketing team wasn’t asking the right questions to get data.
The biggest product improvements still come from understanding your users, preparing your data, and investing in your team!
2. Senior engineers see the most productivity improvement with AI coding because they have more experience with writing design docs and API specs, which help them write better instructions.
However, they’re also more resistant to using AI for coding. Senior folks are often more opinionated and get frustrated easily when AI doesn’t do what they want.
3. Many teams spend a lot of time debating which tool to use, which can be counter-productive. When teams ask me which of the 2 tools to use, I usually ask 2 questions:
“How much performance improvement will the optional tool give over the less optimal one?”
--> If the improvement is small, then spend less time debating.
“How hard is it to change from one tool to another once you’ve adopted it?”
--> If the tool is new and not yet battle tested, I’d think twice about adopting something that I can’t get out later.
4. Many people know that the most effective way to learn AI is to build with AI. Yet, people keep asking me: “But what should I build?”
We seem to be having an “idea crisis”. We have all these wonderful tools to help us build things, and no idea what to build.
An exercise I often recommend is to spend a week noticing what frustrates you in your daily work, then build small tools to solve those specific pain points.