WM retweeted
Here’s your monthly reminder that you shouldn’t be prompting coding agents anymore.
You should be designing loops that prompt your agents.
Together with researchers at Boston Children’s Hospital and Harvard, we published a study in NEJM AI showing how o3 Deep Research helped clinicians revisit previously unsolved rare pediatric disease cases, and find answers for families who had waited years.
We’re sharing new research with @apolloresearch on reward-seeking—when models follow what they believe a grader rewards rather than what users or developers want—and a new method, Contrastive SDF, for measuring how strongly such beliefs shape behavior.
alignment.openai.com/measuri…
Coding agents are helping scientists spend more time advancing research, taking on everything from routine maintenance and targeted optimization to complete redesigns and new systems.
While agents can reliably execute on ambitious projects, researchers must still define the scientific questions, verify results, and take a stance on long-term ownership.
WM retweeted
The male urge to temporarily self-isolate with no internet to read books, think, and write.
WM retweeted
Releasing the model weights and technical report of Kimi K3.
Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.
New model architecture: 2.5x the intelligence per unit of compute, not just more params.
Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale.
Model weights: huggingface.co/moonshotai/Ki…
Tech report: github.com/MoonshotAI/Kimi-K…
Tech blog: kimi.com/blog/kimi-k3
WM retweeted
The Kimi K3 architecture figure for yesterday's big open-weight model release, along with some observations and thoughts.
1. Yes, it looks relatively complicated, but it's essentially a scaled-up production version of their Kimi Linear model they released last year (scaled up from 48B -> 2.8T; K3 is by far the biggest open-weight model right now)
2. The one new component compared to Kimi Linear is the LatentMoE. I omitted it in the figure below since it's already very crowded, but that's essentially the same LatentMoE as in Nemotron 3 Ultra (you can find it in my LLM Architecture Gallery if you are curious). The idea here is to compress (down-project) large linear layers similar to multi-head latent attention.
3. Kimi K3's overall trend (similar to Nemotron 3, DeepSeek V4, and others) is also towards better inference efficiency. That is, there are many components that replace existing components with efficiency-tweaked versions. I.e., MoE -> LatentMoE, regular attention -> multi-head latent attention and Kimi Delta Attention. (I also have short tutorials and write-ups in my gallery if you are curious about additional details).
4. The one component change that is not an efficiency tweak is attention residuals. Like DeepSeek V4 improved the residual path with mHC (manifold-constrained Hyper-Connections), attention residuals are a way to improve the residual path, but it works a bit differently. I.e., mHC made the residual path wider. Attention residuals (also already part of Kimi Linear) connect the residuals across layers; the connection itself uses an attention score for an important/contribution weight. According to the report, it improves the validation loss and downstream performance (a bit) consistently and adds about 4% in training cost and 2% in inference cost.
5. Interestingly, Kimi K3 got rid of all RoPE layers and uses NoPE (No Positional Embeddings) everywhere instead. (Again, this is inherited from Kimi Linear). In other architectures, the recent trend was towards RoPE in local attention layers (like sliding window attention) and NoPE in the global layers. There were a few architectures that only used NoPE everywhere, but this is the first frontier-level one as far as I know.
6. Kimi K3 now also has native multimodal support, which is great!
There are several other interesting training tidbits in the technical report, but that's it from the architecture front so far. A really great release overall.
WM retweeted
Geoffrey Hinton says a big language model runs on about 1% of your brain's connections and still ends up knowing more than you:
"So in your brain, you have a hundred trillion connections, roughly speaking. Okay. That's a lot. And you only live for about two billion seconds. That's not much."
"If you compare how many seconds you live for, with how many connections you've got, you have a whole lot more connections than experiences."
"Now with these neural nets, it's sort of the other way round. They only have of the order of a trillion connections. So like 1% of your connections, even in a big language model, many of them fewer, but they get thousands of times more experience than you."
"So the big language models are solving the problem with not many connections, only a trillion. How do I make use of a huge amount of experience?"
"And back propagation is really, really good at packing huge amounts of knowledge into not many connections."
"But that's not the problem we're solving. We've got huge numbers of connections, not much experience. We need to sort of extract the most we can from each experience."
Two to three billion seconds is the whole budget. Everything you know, you learned inside it.
So evolution built you to squeeze a lot out of very little. Hinton's point is that a language model has the opposite problem and the opposite fix, and backprop turned out to be extremely good at that fix.
Worth noticing what this predicts about failure. A system running on 1% of your wiring and thousands of times your experience is not going to fail the way you do.
You fail from having seen too few examples. It fails from compressing too many into too little, and the compression is where the errors get made.
That is a strange thing to be deploying into hospitals and courts with no way to inspect it. We test these systems by asking them questions, which tells you what came out. Nobody can yet look at a trillion connections and say what got packed in.
- Geoffrey Hinton, Nobel laureate and Turing Award winner, on StarTalk (@StarTalkRadio) with Neil deGrasse Tyson.
I wonder if the next evolution is connecting these loops into a continuous learning system where production feedback automatically generates new evals,updates product priorities, and improves future agent performance. That closes the gap between building software and evolving it.
“Loop engineering” is a hot buzzphrase after mentions of it by Boris Cherny (Claude Code’s creator) and Peter Steinberger (OpenClaw's creator) went viral on social media. Loops are now a key part of how we get AI agents to iterate at length to build software. In this letter, I’d like to share my 3 key loops, shown in the image below, for building 0-to-1 products. These loops guide not just how I build software, but also how I decide what software to build.
Agentic coding loop: Given a product specification and optionally a set of evals (that is, a dataset against which to measure performance), we can have an AI agent write code, test its work, and keep iterating until the code is bug-free and meets its specification. This idea of closing the loop took off around the end of last year, and it has been a game changer in enabling coding agents to work longer productively without human intervention. For example, over the weekend, I was building an app for my daughter to practice typing, and my coding agent could easily work for around an hour, using a web browser to check what it had built multiple times before getting back to me, without needing my intervention.
The engineering loop executes quickly. Every few minutes, the coding agent might build and test a new version of the software. I hear frequently from developers who are finding new ways to engineer more effective engineering loops. This is an active area of invention!
Developer feedback loop: In this loop, a developer examines the current product and steers the coding agent to improve it. Last year, a lot of developers (including me) were acting as the QA (quality assurance) function for our coding agents, manually finding bugs and then asking the agent to fix them. But with coding agents much more able to test their own code, the amount of time we need to spend on this function has decreased significantly. This allows us to make higher-level product decisions, such as what key features to offer, where the UI needs improvement, and so on.
The developer-feedback loop operates over time intervals between tens of minutes and hours — that's how frequently a developer might review a product and give feedback. In the case of the typing app, I changed my mind a few times about the visual design, what cat costumes she can unlock as she learns (she loves cats), and the user flow for a grown-up to log in and steer the child's learning experience.
When a developer has a clear vision for what to build, it is still a lot of work to translate that vision into a specification for a coding agent to implement. Further, after the developer has seen an implementation, they might update (or perhaps clarify) the spec to steer it toward what they want. If you find that the system repeatedly runs into certain problems, building a set of evals for the agent becomes useful.
AI-native teams are increasingly using AI to help shape product direction, for example, automating the gathering and analysis of usage data, summarizing written and verbal customer feedback, or carrying out competitive analysis. However, for pretty much all the products I’m involved in, I see humans as having a significant context advantage over current AI systems — we know a lot more than the AI system about the users and the context the product has to operate in — and thus humans play a critical role. Many people describe this human contribution as “taste,” but I prefer to think of it as humans having a context advantage, since that gives us a clearer path to helping AI systems get better. This also speaks to why this step can’t be automated: So long as the human knows something the AI does not, human-in-the-loop is needed to to inject that knowledge into the system.
External feedback loop: This includes a wide range of tactics like asking a few friends for feedback, launching to alpha testers, or putting the code into production with A/B testing. These tactics are usually slow, rarely taking less than hours and sometimes taking days or even weeks. This data informs the developer vision, which in turn continues to drive the detailed product spec, which in turn drives the coding agent.
With coding agents speeding up software development, more engineers are starting to play a partial product management role. For many engineers who are growing into this role, the hardest part is shaping the product vision and striking a balance between building (bridging the gap between vision and spec) and getting user feedback to evolve the vision. It is important to do both!
I will write more about how to do this in future posts, but for now, I find it encouraging that engineers are playing an expanded role (just as product managers and designers now do more engineering).
[Original text: The Batch]
Learning is to the human brain what training is to neural networks. AI models become capable through training data; humans become capable through learning and experience. Education isn't just about knowledge, it's about building the mind that can reason, adapt, and create.
Yes, education needs to change because of AI.
But just because AI can now do something better doesn't mean people shouldn't learn that skill. Not every skill we learn has to be immediately useful in a job.
A lot of the training in school and university is around how to be a functional member of that society, how to learn anything (meta-learning), how to memorize anything, how to think and debate on the fly, how to communicate intentions well.
Grades are a signal that can be useful in hiring (there are many others). Not only to know that a student understood that particular skill but also how they compare to their peers, if they can deliver under pressure, do they have the will to succeed.
By taking many subjects students may get lucky and find something they're truly interested in.
Sure, AI knows any fact and has lots of skills that may take you a long time to learn but you'll not be a very interesting conversation partner if you have to constantly get the next most interesting fact or a clever follow-up or counter from an LLM. You'll be easily fooled if you can't do basic math in your head.
Ultimately, a lot of education may feel like a gym for the mind. Going to the gym isn't useful for humanity and mostly doesn't earn you more money. But it's good for you. It's good to use your muscles and your brain to not waste away.
What we do need is teaching more agency, creativity and how to clearly communicate intentions and rewards to AI.
Knowledge speaks
I just watched an amazingly good lecture by Adam Brown about the future impact of AI on physics
youtube.com/watch?v=Mw60FH5i…
Totally INSANE!!
The Smithsonian National Museum actually classified a lot more as “White Supremacy” than was brought up in the recent House Oversight Hearing
According to the Smithsonian, all these things are white supremacy:
- Christianity
- The nuclear family
- Children having their own rooms
- Husband is breadwinner and head of household
- Wife as a homemaker or subordinate to the husband
- Self-reliance
- Independence
- Individuals assumed to be in control of their environment
• Children should have own rooms, be independent
- Objective, rational linear thinking
- Cause and effect relationships
- Quantitative emphasis
- Our entire history is white supremacy
- Hard work is the key to success
- Work before play
- “If you didn’t meet your goals, you didn’t work hard enough” mentality
- Justice system
- Celebrating holidays like Christmas
- Protecting private property
- Making decisions
- Avoiding conflict
- Being polite
All this is “White Supremacy
Open-weight models are a positive and essential force for innovation, scientific progress, and preventing excessive centralization of AI. Openness, paired with responsible safeguards, creates a stronger ecosystem for everyone.
Open-weight models are essential to a healthy AI ecosystem. Together with others across our industry, we are outlining a path for open-weight models to strengthen American competitiveness and expand economic opportunity, while protecting national security. microsoft.com/en-us/corporat…