@TimDarcet

LLMs @ FAIR, prev. DINO stuff @ INRIA & FAIR

Joined March 2021
Happy to release ✨ Muse Glimmer ✨ - level ~= Qwen 3.6-27B - Apache 2 - 30B dense - quantized to run in 17GB - quant + spec dec => 50 tok/s on macbook m5 max, interactive, smooth Enjoy!
Introducing Muse Glimmer, an open-weight 30B-parameter model optimized for local, always-on agent workflows. Muse Glimmer delivers strong performance on key agentic use cases and benchmarks compared with leading models in its size category, and is designed to run entirely on consumer hardware like a Mac or PCs with performant GPUs. In keeping with our long tradition of sharing fundamental AI research, we’re releasing model weights under a permissive Apache 2.0 license. 🧵👇
5
15
1
74
7,707
Achh Lucas beats me but also I physically cannot have a lower slop index with this number of papers so I'm happy lol
They call me Mr NoSlop, aye🫡
3
19
3,539
🔥 PhD in the DINO team 🔥 We have a PhD program opening in our team. Don't hesitate to reach out to me or Max if interested!
The DINO team at FAIR Paris is hiring a PhD student to push the frontier of visual representation learning. We offer hard research topics, a research and engineering team supporting you, and a stimulating PhD community. Help us shape the next generation of DINO models! Details👇
2
17
1
202
31,077
TimDarcet retweeted
The DINO team at FAIR Paris is hiring a PhD student to push the frontier of visual representation learning. We offer hard research topics, a research and engineering team supporting you, and a stimulating PhD community. Help us shape the next generation of DINO models! Details👇
5
21
8
246
58,462
> "Geoff, get a life" Fucking ouch
Some folks love MNIST more than anyone ever
21
2,901
I never understood the whole encoder/decoder naming in transformers. Everything is just self attention or cross attention with different attention masks with no typical volume change. The enc / dec naming made sense in VAEs where you bottleneck and learn appropriate posteriors.
encoder decoder is soooo back but is it really an encoder when it's causal 🌝
15
17
3
346
37,307
there was a time when you made your last matplotlib figure and you didn't know it was the last one
kids these days will never know the feeling of making matplotlib charts by hand
31
246
36
3,710
92,966
Want to bring this pet peeve to light: "decoder-only" transformer has never been a decoder, and never will be. It goes from data space to data space. If you want to say it's causal, just say it's causal
Disagree. An encoder = something translating a data point into a code, a decoder = something translating a code into a data point. Bc Vaswani et al used bidirectional attn in the enc and causal attn in the dec, part of the nlp community associated those as if they were synonyms
2
3
66
6,157
TimDarcet retweeted
Chinchilla writes loss as a sum of two independent terms: model size and data never interact. Kaplan's original law coupled them. But interaction has an impact, dropping it bends predictions exactly where you extrapolate. Skaling, our new law, puts it back with one exponent. 1/8
1
13
4
25
5,294
Meta AI researcher who created DINOv2 explains modern self-supervised learning! Also introduces a new approach called CAPI @TimDarcet gave a brilliant talk about his research at the @MedARC_AI journal club last month. Very clear and information-dense, I learned a lot from his talk. Give it a watch!
3
32
1
509
40,040
✨Excited to share the release of Muse Glimmer, a 30B dense open-weight model optimised for local agentic workflows (🧵) Muse Glimmer: - is built for end-to-end agentic task completion, combining reliable tool use, multi-step reasoning, failure recovery, and compatibility with agentic scaffolds to solve complex tasks from start to finish - offers multimodal reasoning and controllable effort, enabling it to interpret text, images, charts, and documents while balancing reasoning depth with speed - is tailored for local deployment to empower more and more people to realise their goals and aspirations 📊Blog post research.meta.ai/blog/introd… 🤗Huggingface huggingface.co/meta-models/M… 📚Developer Documentation dev.meta.ai/docs/muse-glimme… 🌱 Mark’s essay on AI meta.com/thefutureisforevery… Truly grateful to my dear colleagues for the rollercoaster ride Antoine Grosnit, Alexis Audran-Reiss, Arash Abtahizadeh, Adil Ahmed, @AsafAvrahamy, Svetlana Kurtsman, @LuciaCKun, Muna Aghamelu, Noam Levi, He Ye, Joy Chen, Gal Cohen, Onur Çelebi, @roccajo, @HugoTouvron, @CheeHauLeow, Igor Tufanov, Eduardo Sánchez, @PereLluisHC , Hela Momand, @BhavulGauri, Parth Pathak, Kelvin Niu, @kumarsujan, Anton Protopopov, Gaurav Chaurasia, @albertomariape , Jade Yu, @GagnonAudet, @_masoudjalili, @arreqe_ai, @ishitamed, Mattia Opper, Paris Giampouras, @rrmaura , @TimDarcet , Saba Nazir, Sandra Lefdal, @alex_h_miller , @rybolos , Luca Sbordone, Grant Gardner, @maryamfazel and, of course, the one and only Yoram Bacrach! Big thanks to the Meta AI leadership @NailaMurray , @rob_fergus , @LukeZettlemoyer , @skasriel @alexandr_wang , @finkd for making this happen 🙏
8
6
28
2,234
Muse Glimmer, an open-weight 30B agentic model that runs locally. It’s been incredibly fun to be a part of this journey and work alongside the amazing MSL team! ✨ Excited to see what the community builds with it! 👇 🤗: huggingface.co/meta-models/M… | 📝: research.meta.ai/blog/introd…
Today we're also opening the weights for Muse Glimmer, a great 30B parameter dense model that can run locally. Soon we'll also release the weights for Muse Spark 1.2, our latest foundation model. Meta is a strong supporter of open source and I'm proud of these releases. Congrats to @alexandr_wang and the MSL team for all your great work on these models.
7
4
26
2,013
Replying to @AIatMeta
@AIatMeta releases new open source LLMs! Exciting stuff!
Happy to release ✨ Muse Glimmer ✨ - level ~= Qwen 3.6-27B - Apache 2 - 30B dense - quantized to run in 17GB - quant + spec dec => 50 tok/s on macbook m5 max, interactive, smooth Enjoy!
2
5
359
TimDarcet retweeted
A year ago I was hoping for this !! Today it finally happened It may not be called Llama anymore, but seeing Meta release a frontier open weight model under an Apache license feels like history repeating itself. Huge win for the open ecosystem.
really hope this to happen soon
2
3
68
5,162
TimDarcet retweeted
Llama 5 basically. Super excited
Today we're also opening the weights for Muse Glimmer, a great 30B parameter dense model that can run locally. Soon we'll also release the weights for Muse Spark 1.2, our latest foundation model. Meta is a strong supporter of open source and I'm proud of these releases. Congrats to @alexandr_wang and the MSL team for all your great work on these models.
4
2
97
7,865
TimDarcet retweeted
Apache 2.0 🥹 first Gemma, now Meta we are so back
Today we're also opening the weights for Muse Glimmer, a great 30B parameter dense model that can run locally. Soon we'll also release the weights for Muse Spark 1.2, our latest foundation model. Meta is a strong supporter of open source and I'm proud of these releases. Congrats to @alexandr_wang and the MSL team for all your great work on these models.
4
8
1
206
13,600
Meet Muse Glimmer: an open-weight model built for always-on local agents. 30B parameters, Apache 2.0 and tuned for complex multi-step work so it plans, calls tools, hits errors, retries, and sees the task through long-horizon loops. Muse Glimmer is designed to balance capability against the memory and compute constraints of local hardware. We couldn’t be more excited to get this in the hands of developers. Download and start building now: bit.ly/4yYUKkS
58
188
30
1,568
114,372
TimDarcet retweeted
Thank you Mark, Alex and the whole Meta team for your contributions to open weight AI.
Today we're also opening the weights for Muse Glimmer, a great 30B parameter dense model that can run locally. Soon we'll also release the weights for Muse Spark 1.2, our latest foundation model. Meta is a strong supporter of open source and I'm proud of these releases. Congrats to @alexandr_wang and the MSL team for all your great work on these models.
78
164
7
3,891
347,598
We just shipped a very capable local agentic model. That can turn your gaming PC / MacBook Pro / 5090 into an agentic workstation!
1/ big announcement today: we will be releasing an open weight version of muse spark 1.2 soon. we also are releasing muse glimmer, a 30B agentic model with open weights under apache 2.0. muse glimmer can run on 24GB of VRAM without losing agentic reliability. 🧵
3
2
16
945
Meta is back with Muse Glimmer: a 30B open-source multimodal model built for local, agentic use. HF is shipping day-0 support and I built a few demos to see what it can do. First: we gave Glimmer tools and asked it to quantize itself.
9
20
6
125
51,696