@OpenEuroLLMi
iAccount based inSpain!
About this account
- Account based in
- Spain
- Connected via
- Web
! X says this location may be affected by a proxy or VPN.
Account-level information from X, not a live location or the device used for a specific post.
A series of foundation models for transparent AI in Europe
Joined June 2024
- Tweets26
- Following5
- Followers448
- Likes24
A series of technical reports from our project, like the one just posted by our colleague @kleiaaro, will try to provide answers to the many questions that LLM training teams cope with around the globe.
This one is about training recipes 👇
🚀 How do you choose the right LLM training recipe? We just released a tech report on the scaling-law pipeline of @OpenEuroLLM 🇪🇺, along with the corresponding checkpoints, data, and code.
arxiv.org/abs/2608.28308 1/N 🧵
Our 9B experimental WIP model now named
🎼 PRELUDE
is still training but we are sharing intermediate checkpoints for transparency:
🛑 WIP research artifact.
🛑 NOT cooled down yet.
✔️Include validation across 36 languages.
Full WIP model card👇huggingface.co/openeurollm/p…
Open models matter, we can't agree more!
Open weights is surely a necessary step.
The OpenEuroLLM project extends this view also to the data, training and evaluation aspects of current AI development.
#transparentAI #openScience
@HajicJan
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
images.nvidia.com/pdf/Open-W…
We are at 1/3 of training our 9B model! 🏃🏃🏃
✨Follow progress here: huggingface.co/spaces/openeu…
🪇See our public Weights & Biases training workspace here: wandb.ai/openeurollm-project…
More soon!
#TransparentAI #OpenScience #goOpenEuroLLM
Input, more input 🤖⚡
Just like Jonny 5 in Short Circuit, our baby model is reading every single token from its pretraining dataset.
So far: 10 trillion tokens, 36 languages + code & math as their own "languages" 📚🌍💻
We’re tracking progress & sharing it openly 👇
(1/2)
As of this morning:
🧠 425.49B tokens seen
📊 4.25% completed
This eager reader wants more input, one token at a time.
Follow along. 🔍
#PreTraining #LLM #MultilingualAI #TransparentAI
#goOpenEuroLLM
Pretraining launched!🚀
Our 9B/10TT baby model is making its first steps in Leonardo (CINECA). 🐣
All people involved are eager to see the results of the effort it took to get here and share them. 👀
And advancing to push hard for the next cycle. 🦾
#goOpenEuroLLM
Wrapping up our 3rd general meeting, hosted by @AISweden in sunny Stockholm ☀️
A full room makes the final decisions before training the first OpenEuroLLM model. Sharing updates, ideas, and future plans.
Two more days of tight collaboration.
Full speed mode. 🚀
#goOpenEuroLLM
HPLT is of the datasets we are sharing in our world-readable catalogue across HPCs. Interesting talk!
We are at #LREC2026 presenting HPLT v3 Datasets, monolingual, parallel, massive, highly curated. For depth info and analysis, please:
Join us in room Menorca 1 at 16:20!!!
All ready to share information about #OpenEuroLLM with the #LREC2026 crowd. Let's talk data, infra, evals and open multilingual LLM models together! Come to booth #5 at the poster area 1, Elyxir Building.
#multingualLLMs #openLLMs #diverseLLMs #safeLLMs
Quite a nice "representation" of the OpenEuroLLM crowd will be at the International Conference on Learning Representations (ICLR) this week.
On Friday 24, come to poster "OpenThoughts: Data Recipes for Reasoning Models", work partially supported by our project, and meet us!👋
Experimenting with model-based annotation for better data selection? A candidate to consider is propella-1, a multi-property annotator partially funded by #OpenEuroLLM which is fully open-source.
🔓Code, annotations and paper available! arxiv.org/pdf/2602.12414
We released propella-1, a small model for advanced pre-training data annotation 🙃.
Work led by @maxidahl within the @OpenEuroLLM project. Link to model + annotations for important pre-training datasets below 👇
🎉 One year of OpenEuroLLM!
🇪🇺We’re building Europe’s next-gen open-source LLMs to boost digital sovereignty.
More about our achievements and next steps for infrastructure, data, models and evaluation at openeurollm.eu/blog/first-ye….
Year 2 = full speed ahead. 🚀
Go #OpenEuroLLM
First OpenEuroLLM Winter School in collaboration with the @CircleU_eu Alliance 🧑🎓and the Nordic Language Processing Laboratory 🧑💻
Focus on Multilinguality in LLM Development and Evaluation with speakers from world organisations, academia and industry.
wiki.nlpl.eu/Community/train…
Strategic access to EuroHPC resources granted to OpenEuroLLM!!!
-first AI project granted strategic access across multiple EuroHPC centres
-for over 10 million GPU hours
Thanks @EUComission and @EuroHPC_JU!
OpenEuroLLM retweeted
Proud to present the @OpenEuroLLM project and its results so far, with Sampo Pyysalo (@UniTurku) at the 1st Workshop on Open Source Sovereign LLMs in Berlin osfm.info/ Great opportunity to talk to many OS LLM developers! @CharlesUniPRG @hplt_eu – at Berlin, Germany
We strongly agree! Let's make it happen! Thanks @EU_Budget & STEP for the support.
Future-proof AI in all EU languages isn’t a dream, it’s OpenEuroLLM 🗣️💬
9 countries, the EU budget & STEP join forces to build transparent, AI Act-compliant tech for Europe’s innovators.
Find out how we will turn ambition into action for 2028-2034: europa.eu/!w77nKY
Well done #HPLT, we surely need more high performance datasets for the present and future landscape of multilingual LLMs. See you at #emnlp2025!
The #HPLT crowd is at #EMNLP2025!!!
If you are around, please visit our booth to discuss:
- multilingual datasets 🌏
- dataset insights and stats 📊
- dataset performance 🔝
- efficient MT models ⏱️
- and the future of multilingual LLMs 💡
We don't want to miss U!
OpenEuroLLM completing 2 days of sharing progress and next steps pursuing the goal of developing strong multilingual foundation models aligned with European strategic vision & standards.
Gathering at BSC nearby MareNostrum 5 supercomputer made us feel home.
#Barcelona #NLProc