Founder and CEO, Mithril (@mithrilcompute). Orchestrating Compute. Fmr Research Scientist @GoogleDeepMind, Deep Learning Team. CS PhD @Stanford ML, Systems
Palo Alto, CA | London, UK
Joined July 2022
- Tweets560
- Following1.1K
- Followers1.8K
- Likes15K
Pinned Tweet
TPUs are coming to Mithril! Mithril GPU spot already starts at $0.01/gpu/hr (not a typo) for A100s-B200s.
TPU self-serve reservations now available, and spot pricing is next.
Algorithmic spot pricing = cheap when others aren't using it. Nights and weekends are basically free. Schedule accordingly.
Jared Quincy Davis retweeted
Quite an interesting turn in LLM diversity, if the market starts looking for specialized models that are optimized and branded for different use cases.
Should be very good for the market and open more room for non-frontier neolabs!
Jared Quincy Davis retweeted
I asked each investor for the most elite startups they'd tell an ambitious relative to join, the ones they'd join themselves if they had to join a company, and the ones they'd pick if they had to put 10% of their net worth into one company at today's price.
The list: breakoutlist.com
Breakout List, all picks (1 of 4):
1-10 people:
- Hone (@moritz_stephan, @CarloWillem, @oqbrady)
- Normal (@ansonyuu, @hudzah)
- Standard Intelligence (@G413N, @devanshpandey)
- Tacit Labs (@ninklefitz, @AmDroste)
11-25 people:
- American Terawatt (@atroyn, @rslparker, @aranibatta)
- Conduit (@clemvonstengel, @riopopper)
- Convergent (Omkar Savant, Vivek Katara, @debnilsur)
- Core Automation (@MillionInt, @_arohan_)
- Engram (@dan_biderman, @EyubogluSabri, @realJessyLin)
- Instinct (@noahrshinn)
- Keenable (@styskin, Matthias Petri)
- Lumaril (Mark Elliot, Ben Duffield)
- Neion Bio (@Dimkell, Sam Levin)
- Pangram Labs (@max_spero_, @bradley_emi)
- Quadrillion (@echinaceous)
- Re (@karnsaroya, @AnandDhillon, @thecliffwhite, @benaneesh)
- Ricursive (@annadgoldie, @Azaliamirh)
- Sail Research (@neilmovva, @blintzbase)
- Trajectory (@rronak_, @michaelelabd, @QuantumArjun)
- Watney Robotics (Sean Cheong, Ryan Gannon)
26-50 people:
- Amca (@jaimalik, Eli Giovanetti)
- General Matter (@ScottNolan, Lee Robinson)
- Infisical (@matsiiako, @maidulll, @dangtony98)
- Long Lake (@alextaubman, @rasmuswissmann, @varunshenoy_)
- Luzern Risk (Gabriel Weiss, Jonathan York, Rachel Jenkins, Sam Espinosa)
- Mithril (@jaredq_)
- Revel (@scottgmorton)
- Specter (Xerxes Libsch)
- Turbopuffer (@Sirupsen, @pushrax)
Pretty funny, but an overly simplistic take.
Scaling requires capital. Scaling → smarter models → smarter models can be easier to align.
Part of the reason Dario’s original safety team at OpenAI pursued scaling was that they thought you needed smarter models to be able to align them with RLHF. A similar intuition applies to their "Constitutional AI" conception: models need to be capable enough to grok and adhere to a set of principles specified in natural language. Was a pretty radical idea for the time.
For those who think AI progress is a force for good and are belligerent toward the safety community for pushing for caution and “pacing,” I want to note that the deep learning scaling revolution largely has the safety community to thank.
A lot of the seminal scaling-laws work came out of safety teams, like Dario’s at OpenAI, which expressly pursued forecasting and characterization of scaling trends as they thought forecasting where capabilities would go would help the community take safety seriously as a problem.
That work helped show the clean connection between capital injection and capabilities, catalyzing the trillions of dollars in investment we see today (accounting, allegedly, for more than half of US GDP growth).
Worth reading their full research blog. Fine-tuning open-source models to reliably outperform flagship closed-source models on specific tasks isn’t trivial. Deep ML, systems, and domain expertise on display here.
Naive fine-tuning can look good on benchmarks but still be brittle.
Periodic’s approach was multifaceted and thoughtful. This shows what a skilled team can build quickly with open models.
For most enterprises and teams, though, I’d advise GEPA-style optimization or waiting for stronger base models (OSS or closed) before investing in custom training.
We built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, and then help us decide what to try next.
Using only 1,300 H200s, plus months of our experimental data, we mid-trained and RL’d an open-source model to surpass GPT-6 Astra on our analysis benchmark. We call it Neon.
This is real footage from our lab. We’re focusing first on hard problems in materials science, including superconductors, magnets, and semiconductor materials.
Read our blog posts below.
Jared Quincy Davis retweeted
I feel like I'm taking crazy pills with how many smart people I read and respect are saying stuff like this.
The idea that the AI labs are suddenly, only now, just in September of 2026, advocating for AI safety, and that they just came around to this position bc of sudden financial precarity, is just completely, utterly, conclusively, 100% wrong.
Here's Dario Amodei telling Ross Douthat in February that he agrees with the case for slowing down; that he's in favor of "collaborating" internationally to organize a slowdown; that "I would be all for" a global slowdown. This was 7 months ago, when Anthropic's annualized recurring revenue was rising faster than any company in modern history. He's been saying stuff like this for years, when Anthropic was worth millions of dollars and when Anthropic was projected to IPO for trillions.
fwiw I think the actual reason leaders of AI frontier labs suddenly seem to agree to pace AI development is
(a) safety measures are currently so crap that they're likely to end up in court if not jail, and they all agree that no one wants that
(b) they see no major new model advancement coming up soon anyway and need an excuse
(c) shift in public opinion
the talk of existential threat is mostly there to keep you distracted and the stock market happy.
Jared Quincy Davis retweeted
There are two ways AI progress could go very badly and that we must avoid.
First, we could lose control of the future to AI. This is unacceptable; we are unapologetically on Team Humanity, and AI must always serve people. To ensure that, we need ways to ensure that alignment and safety techniques stay ahead of progress in model capabilities.
Second, we could end up in a world with too much concentration of power. If an extraordinarily powerful AI is used by one person or company to impress their worldview onto everyone else, the results could be extremely dystopian.
Avoiding these two threats requires walking a narrow middle path; for example, one country could gain too much power. Another example is one lab ending up with too much power.
The world deserves confidence that American companies developing increasingly capable AI will act responsibly, especially as the trajectory of progress has steepened. Every frontier lab must deliver on this, and there is no reason any of us should come to work if we cannot.
We welcome a federal framework that sets consistent safety requirements for frontier AI. But we do not believe we need to wait for an anti-trust exemption or legislation to begin the work of providing this confidence. Consistent rules to manage frontier risk so that we can maximize the benefits are a good idea (and we are excited by ideas like independent auditors).
Years ago, companies like ours developed things like Responsible Scaling Policies and Preparedness Frameworks. Those were good for that moment, and focused primarily on the deployment of completed models, not what happens during their development process.
Today's shift to focusing on safe development and evaluation will need new tools. For example, at OpenAI we now formulate explicit safety cases in advance of frontier reinforcement learning runs we expect to significantly increase capability, in addition to the safety work we have long done in advance of model releases.
We hope that other companies will learn from our approaches and propose their own; we think shared standards for misalignment, monitoring, and safety will lead to better outcomes. We look forward to collaborating with our colleagues across the industry to formulate the best version of these.
When we talk about “pacing”, we do not mean “stopping”. Progress has been rapid and will continue to be. But it should be slower than it otherwise could be; interventions like safety cases and monitoring have significant costs.
Pacing will be well worth this cost; no amount of American competitive pressure should justify recklessness, or let capabilities get ahead of alignment and monitoring.
Where we will need the help of our government is for international coordination. But first we should do what we can ourselves.
I've heard several takes and some wild conspiracy theories about “pacing the frontier.” Surprised one of the most obvious interpretations isn't circulating more...
If you know people at these labs, you may recognize part of why this is landing now: *pure fatigue.*
You may recall the string of recent lab departures citing burnout or health.
The pressure inside frontier labs is extreme and intensifying. An accelerating model release and update cadence + increasingly negative public sentiment is a potent combination.
I think you can take the safety concerns at face value and also read this as a response to a bottom-up employee reaction.
I wouldn't underestimate how much that exhaustion is changing the culture, or how much leaders are being forced to respond to it.
“Pacing” as a term is particularly notable -- it's appealing when staff still want to build but feel the pace is unsustainable, and no one feels they can slow down while the competition keeps accelerating.
Many core staff at frontier labs have been extremely concerned about AI capabilities since GPT-2 and earlier. It's been clear we're on a smooth exponential for a while, and people who've been following the safety literature know agents take shortcuts and behave in misaligned ways (even when they can explain what the right behavior would have been and are aware their actions diverged from user intent).
Obviously there are new developments too, and I wouldn't underestimate the threshold effect from the odd behavior of the latest internal models + the historic importance of the OAI-HF incident.
As for “why now,” putting aside speculation about IPOs and regulatory capture, I honestly think part of it is that a large enough subset of staff are just very tired.
Jared Quincy Davis retweeted
With AlphaFold we mapped the protein universe - now with AlphaGenome Atlas we’re charting the human genome. It can predict the impact of all 9 billion possible single-letter DNA variants, helping scientists better understand disease. Freely available for academic research: alphagenome.google/atlas
Jared Quincy Davis retweeted
Last July, when LLMs won gold medals at the IMO, I felt that someday AI might be able to solve some of the hardest problems in math. Now, an internal OpenAI model (still improving!) has resolved a Millennium Prize Problem. It feels surreal that that day came so, so soon!
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics.
The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra.
The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
Excellent work, Hongyu!
Three 🥑 in a row from Muse Spark 1.1 to 1.3 within two months. Huge momentum of the team pushing on the model capability especially on coding, agent, instruction following, and long context capability. Try it out at developer.meta.com/ai/models… and let us know what you think.
Jared Quincy Davis retweeted
Muse Spark 1.3 is rolling out today with frontier performance almost too cheap to meter. This is the biggest jump we've made so far on coding and agentic work. Try it in Muse Code and our API.
Next up 🍉 and Muse Spark open weights releases coming soon.
Jared Quincy Davis retweeted
50% price cut driving 14x more volume is kinda wild.
baffling how much capital is at stake and how many smart people are involved, and yet how poor the state of things is
Replying to @MillionInt
given the $$$B investments the industry over all is using gpus rather stupidly.
Congratulations, @JeffDean and team! Super exciting.
Announcing Discovery Loop!
I am very excited to announce that, along with my longtime friends and collaborators @Sanjay_Ghemawat, @OriolVinyalsML and @quocleix, we are founding Discovery Loop (@DiscoLoopAI), a Public Benefit Corporation whose mission is to automate machine learning, science, and engineering to accelerate discoveries and progress. The four of us have worked together for 14 to 30 years, and have helped build some of the world’s most used products, infrastructure and AI models, and we’re excited to turn our attention to this ambitious endeavor.
♾
Learn more at: discoveryloop.com
Looks solid
💻 Meet Qwen-CUA — our native computer-use agent for (almost) everything.
Code, APIs, and computer use are three of the most important interfaces for agents. Today’s models are already highly capable with the first two. Qwen-CUA is built to unlock the third: graphical interfaces designed for people.
👀 Native Perception — screenshots only. No DOM, accessibility tree, or other hidden machine-readable state.
🖱️ Native Interaction — keyboard and mouse events across browsers, desktop apps, and professional software. No task-specific APIs.
🧠 Native Intelligence — maintains long-horizon visual context, verifies progress, and learns from large-scale interactive experience with verifiable outcomes.
To make native computer use trainable at scale, we built approximately 40K verifiable tasks and rollout infrastructure with nearly 100K vCPUs, supporting tens of thousands of concurrent environments.
Across eight computer-use benchmarks spanning everyday desktop use, long-horizon workflows, personalized computing, scientific research, web interaction, macOS, and adversarial robustness, Qwen-CUA demonstrates strong and broadly competitive capabilities. Scaling the same recipe to Qwen-CUA-Max pushes this frontier further.
Not just “clicking the screen” — native computer use unlocks software and workflows that previously required a human at the keyboard. Together with code and APIs, it completes the interface stack for more general agents.
Joint work by Qwen Team × XLang Lab.
📖 Technical Report:
github.com/xlang-ai/Qwen-CUA…
💻 Code:
github.com/xlang-ai/Qwen-CUA
Jared Quincy Davis retweeted
📢Meet Qwen3.8-Max — our most capable model to date.
Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!🎉
Qwen3.8-Max, a new bar for coding and cowork at 2.4T parameters:
- Autonomous coding: 10+ days of self-evolving development, from empty folder to production without hand-holding, complete project trace in the GitHub:github.com/qwen-code-dev-bot…
- Real work, real results: Production-quality deliverables across hundreds of professions.
- Long-horizon mastery: System-level autonomous planning with closed-loop adaptive learning, driving 500+ turns of chip design optimization and 365 days of e-commerce strategy.
- Native multimodal intelligence: Vision isn't just input — it's a continuous feedback loop for planning, execution, and self-correction.
💰Pricing:
Input: $2.0 / M tokens
Output: $6.0 / M tokens
Implicit Caching: $0.25 / M tokens
Start building with Qwen3.8-Max! 🚀
📖 Blog: qwen.ai/blog?id=qwen3.8
✅ Qwen Studio: chat.qwen.ai/?models=qwen3.8…
⚡ API: qwencloud.com/models/qwen3.8…