Monday's community meeting showcases several standout community projects, including:
• floki: an HTTP client for Mojo that works like Python's requests library, supported by the Modular Community Grant Program
• noeira: an end-to-end physical AI stack in Mojo - physics, learning, perception and deployment, from simulator to robot, supported by the Modular Community Grant Program
• warp: a look at building an async runtime to manage GPU synchronization calls across coroutines in Mojo
Join us via Zoom at 10 AM PT: luma.com/sep-modular
This week in Chicago, Mojo found The Bean, ate a pretzel half its size, and got to hang out with the whole Modular team at our off-site.
Want to come to the next one? We're hiring across Engineering, Product Management, Customer Engineering, and Developer Relations: modular.com/company/careers
Modular retweeted
.@clattner_llvm takes the #SnapdragonSummit stage for the first time as EVP Advanced AI Software @Qualcomm following the @Modular acquisition.
Introducing the audience to the stack and it's role in the broader ecosystem.
Modular retweeted
GPU architecture | LLM Inference Handbook
handbook.modular.com/kernel-…
Add this to your LLM learning resource bundle.
"Before writing or tuning GPU kernels, you need a working model of how a GPU runs code. Without it, suggestions like “increase occupancy” or "reduce shared memory bank conflicts" are just a set of rules to memorize. You don't fully understand when they apply and when they don't.
This section explains modern GPU architecture at the level needed for kernel work. The details lean toward NVIDIA hardware because CUDA dominates much of the LLM inference ecosystem today. However, the core concepts apply broadly to AMD GPUs and other parallel accelerators as well."
Optimizing large scale inference systems is what we do, so we decided to write down what we know.
Our LLM Inference Handbook is a free reference covering TTFT, TPOT, goodput, continuous batching, chunked prefill, prefix caching, KV cache math, prefill-decode disaggregation, quantization, and more.
It includes 20+ interactive visualizations, is updated continuously, and is open to PRs.
handbook.modular.com
Modular retweeted
My personal favourites :
1. handbook.modular.com/
2. jax-ml.github.io/scaling-boo…
3. baseten.co/inference-enginee…
Looking to get started with MAX?
At ModCon 2026, Ehsan M. Kermani @ehsanmok and Bingfeng Xia went layer by layer through the stack: MAX Serve, the framework, and the Mojo kernel library, ending with a live agent bringing up a new model end to end.
Start here: youtube.com/watch?v=3H8Orjaa…
Contributing guide, changelogs, and open issues are linked in the release post. If 26.6 breaks something for you, please let us know in GitHub Issues.
Read the full blog: modular.com/blog/modular-26-…
On the MAX side, audio joins text, vision, and image generation. The new audio_generation pipeline launches with MiniMax-Music3, producing 44.1 kHz stereo music from a style prompt and lyrics.
MAX performance improved, too: up to 4.8x faster Gemma 4 decode attention and 7.9x faster MoE routing on NVIDIA B200, and up to 6.6x faster decode attention projections on AMD MI355.
MAX changelog: max.modular.com/releases/v26…
We open sourced the Mojo compiler under Apache 2.0 at ModCon last month. Opening it to contributions was the top request we heard afterward, and it took some time to get the infrastructure right. It's ready now. We also migrated our internal issues to public GitHub issues, so you can read what the compiler team is working on this week.
Mojo changelog: mojolang.org/releases/v1.1.0…
Modular 26.6 is here. Mojo 1.1 opens the compiler to external contributions and adds developer experience improvements, while MAX 26.6 brings audio generation, new model architectures, and faster performance.
Release blog: modular.com/blog/modular-26-…
At ModCon this year, we asked five investors where the next wave of AI infrastructure capital is going: training, inference, or silicon? Five different answers, and one panelist said it's the wrong question to be asking.
The discussion also covered whether chipmakers absorbing AI software means the infrastructure is maturing or getting ahead of itself, and what open weight models do to the valuation of a compute-heavy startup.
Michelle Gonzalez (@M12vc), Liz Stein (@USITfund), Sam Fort (@dfjgrowth), Quentin Clark (@generalcatalyst) and Dave Munichiello (@GVteam) each closed with what they think founders should build right now.
Full recording: youtube.com/watch?v=5igY763x…
Fragmentation creates a tax across the entire AI ecosystem. At @PyTorch Conference this year, @clattner_llvm presents an alternative: one software stack to unite heterogeneous hardware.
Join us in San Jose on October 20-21: hubs.la/Q04v4SL60 #PyTorchCon
New chips are shipping faster than ever before, but the ecosystem is being held back by having to rewrite and then re-debug the same code over & over again.
@clattner_llvm CEO and Co-founder at @Modular and EVP of Advanced AI Software & Platforms at @Qualcomm, will deliver a keynote at PyTorch Conference North America about an open software platform for heterogeneous compute powered by Mojo and MAX.
PyTorch has always been the place where the best models come together, and now there's a way to get those models onto all kinds of hardware.
If you're interested in Al and compute, join us at the PyTorch Conference in San Jose, CA. Register now: hubs.la/Q04v4SL60
#PyTorchCon
Headed to Santa Clara today for #AIInfraSummit?
Don’t miss @alisterburt's talk at 10:30 AM PT in Expo Theater 2: "Modular: Open Source, Open Cloud, Open Silicon."
Stop by the Qualcomm booth (#206) anytime this week during expo hall hours to chat with the Modular team and catch a Modular Cloud demo.
If you looked at Mojo in 2024, liked it, and decided to check back later, later is now.
Last month, the language hit 1.0, beginning a new epoch of stability for Mojo. The full Mojo language also went open source under Apache 2.0, including the compiler and more.
Now is a great time to start building with Mojo. Clone the repo, build the compiler yourself, and learn from its unique MLIR internals and full commit history.
At ModCon 2026, Brad Larson (Staff PM, Mojo) and Denis Gurchenkov (Senior Director, Mojo Compiler Engineering) walked through why Modular built a language at all, what stability means in the 1.x series, and a hint of what's next: youtube.com/watch?v=a7YdrWOf…
Hippocratic AI cofounders @munjalshah and @debo_datta_ on the latency problem, from the ModCon keynote: youtube.com/watch?v=hUSMA2r2…
We partnered with Hippocratic AI to integrate MAX into their inference pipelines running on NVIDIA B300 GPUs.
Benchmarked against an existing SGLang deployment on 400B+ parameter models, MAX delivered sub-500ms mean time to first token and approximately 30% faster P99 end-to-end latency.
Every millisecond matters in real-time voice, and these gains compound rapidly at the scale of Hippocratic AI's Polaris system.
.@hippocraticai's health agents call tens of thousands of patients a day. Each conversational turn has to finish in about 800ms or the call stops feeling human.
That constraint is why they came to us.