Open and responsible research and development of large language models for code. #BigCodeProject run by @huggingface + @ServiceNowRSRCH
Joined August 2022
- Tweets271
- Following3
- Followers9.1K
- Likes226
- For more details, please check out the blog: huggingface.co/blog/bigcode/…
- Try recent LLMs (e.g., Qwen3 series and DeepSeek-V3.2) on BigCodeArena now: huggingface.co/spaces/bigcod…
- Paper Link: drive.google.com/file/d/1gt5…
- GitHub: github.com/bigcode-project/b…
BigCodeArena cannot be built without the support of the BigCode community. We are grateful for the huge credits provided by the @e2b team. We thank @hyperbolic_labs, @nvidia, and @Alibaba_Qwen for providing the model inference endpoints.
Beyond human votes, we release two new benchmarks:
- BigCodeReward: tests reward models on 4.7K human preference votes. Execution feedback improves judgment accuracy.
- AutoCodeArena: automated evaluation of 20+ LLMs. GPT-5 leads, followed by Claude-Opus-4 and Claude-Sonnet-4.
In 5 months, BigCodeArena collected 14K conversations & 4.7K preference votes across 10 frontier LLMs.
Findings:
1. o3-mini & o1-mini consistently top Elo rankings
2. Claude-3.5-Sonnet excels in matched-language settings
3. Previous open models still lag
Why does this matter?
Benchmarks like HumanEval only scratch the surface. Reading code “by eye” is error-prone. True quality emerges when you actually run it: web apps render, games play, edge cases break.
BigCodeArena makes execution feedback the default.
Introducing BigCodeArena, a human-in-the-loop platform for evaluating code through execution.
Unlike current open evaluation platforms that collect human preferences on text, it enables interaction with runnable code to assess functionality and quality across any language.
Releasing BigCodeBench-Hard: a subset of more challenging and user-facing tasks.
BigCodeBench-Hard provides more accurate model performance evaluations and we also investigate some recent model updates.
Read more: huggingface.co/blog/terryyz/…
Leaderboard: huggingface.co/spaces/bigcod…
We release leaderboard, dataset, code, and paper:
- 🤓 Blog: hf.co/blog/leaderboard-bigco…
- 🌐 Website: bigcode-bench.github.io/
- 🏆 Leaderboard: huggingface.co/spaces/bigcod…
- 📚 Dataset: huggingface.co/datasets/bigc…
- 🛠️ Code: github.com/bigcode-project/b…
- 📄 Paper: github.com/bigcode-bench/big…
BigCodeBench contains 1,140 function-level tasks to challenge LLMs to follow instructions and compose multiple function calls as tools from 139 Python libraries. To evaluate LLMs rigorously, each programming task encompasses 5.6 test cases with an average branch coverage of 99%.
Introducing 🌸BigCodeBench: Benchmarking Large Language Models on Solving Practical and Challenging Programming Tasks!
BigCodeBench goes beyond simple evals like HumanEval and MBPP and tests LLMs on more realistic and challenging coding tasks.
We release all the code, datasets, and models with a permissive license:
🤖Model: huggingface.co/bigcode/starc…
⚙️Code: github.com/bigcode-project/s…
📚Dataset: huggingface.co/datasets/bigc…
Releasing StarCoder2 Instruct! 🚀
Achieves 72% HumanEval score using only self-generated content without any GPT-3.5/4 data. This work demonstrates that self-instruct works already well at the 15B scale without data from proprietary models!
Read more: huggingface.co/blog/sc2-inst…
This is the result of hard work by the BigCode community and supported by @ServiceNowRSRCH, @huggingface and @nvidia to train the 3B, 7B and 15B models!
The Stack v2 was built with @SoftwareHeritage and the full processed training data is coming soon!
huggingface.co/datasets/bigc…
StarCoder2 performs well on a wide range of coding and math tasks. The 15B model is best in its class, while the 3B model is at the performance of StarCoder1-15B.
This makes StarCoder2 models more efficient and performant!
Read the full report: hf.co/bigcode/report
Introducing: StarCoder2 and The Stack v2 ⭐️
StarCoder2 is trained with a 16k token context and repo-level information for 4T+ tokens. All built on The Stack v2 - the largest code dataset with 900B+ tokens.
All code, data and models are fully open!
hf.co/bigcode/starcoder2-15b
Exciting times: we are working on the next generation of StarCoder trained on a new dataset! 🚀
If you would like to have your code excluded from the training run you can check if your data is in the dataset and follow the link to opt-out:
huggingface.co/spaces/bigcod…
✨ For more information on fine-tuning and deploying these models check the StarCoder and starcoder.cpp GitHub repos:
github.com/bigcode-project/s…
github.com/bigcode-project/s…
@nomic_ai team already added support for StarCoderBase-3B in their GPT4ALL local models.
Download the model at: gpt4all.io/models/starcoderb…
& follow the docs: docs.gpt4all.io/gpt4all_pyth…
Stay tuned for the 7B model integration!