Open and responsible research and development of large language models for code. #BigCodeProject run by @huggingface + @ServiceNowRSRCH
Joined August 2022
- Tweets271
- Following3
- Followers9.1K
- Likes226
Beyond human votes, we release two new benchmarks:
- BigCodeReward: tests reward models on 4.7K human preference votes. Execution feedback improves judgment accuracy.
- AutoCodeArena: automated evaluation of 20+ LLMs. GPT-5 leads, followed by Claude-Opus-4 and Claude-Sonnet-4.
In 5 months, BigCodeArena collected 14K conversations & 4.7K preference votes across 10 frontier LLMs.
Findings:
1. o3-mini & o1-mini consistently top Elo rankings
2. Claude-3.5-Sonnet excels in matched-language settings
3. Previous open models still lag
Why does this matter?
Benchmarks like HumanEval only scratch the surface. Reading code “by eye” is error-prone. True quality emerges when you actually run it: web apps render, games play, edge cases break.
BigCodeArena makes execution feedback the default.
Introducing BigCodeArena, a human-in-the-loop platform for evaluating code through execution.
Unlike current open evaluation platforms that collect human preferences on text, it enables interaction with runnable code to assess functionality and quality across any language.
Releasing BigCodeBench-Hard: a subset of more challenging and user-facing tasks.
BigCodeBench-Hard provides more accurate model performance evaluations and we also investigate some recent model updates.
Read more: huggingface.co/blog/terryyz/…
Leaderboard: huggingface.co/spaces/bigcod…
Introducing 🌸BigCodeBench: Benchmarking Large Language Models on Solving Practical and Challenging Programming Tasks!
BigCodeBench goes beyond simple evals like HumanEval and MBPP and tests LLMs on more realistic and challenging coding tasks.
Releasing StarCoder2 Instruct! 🚀
Achieves 72% HumanEval score using only self-generated content without any GPT-3.5/4 data. This work demonstrates that self-instruct works already well at the 15B scale without data from proprietary models!
Read more: huggingface.co/blog/sc2-inst…
This is the result of hard work by the BigCode community and supported by @ServiceNowRSRCH, @huggingface and @nvidia to train the 3B, 7B and 15B models!
The Stack v2 was built with @SoftwareHeritage and the full processed training data is coming soon!
huggingface.co/datasets/bigc…
StarCoder2 performs well on a wide range of coding and math tasks. The 15B model is best in its class, while the 3B model is at the performance of StarCoder1-15B.
This makes StarCoder2 models more efficient and performant!
Read the full report: hf.co/bigcode/report
Introducing: StarCoder2 and The Stack v2 ⭐️
StarCoder2 is trained with a 16k token context and repo-level information for 4T+ tokens. All built on The Stack v2 - the largest code dataset with 900B+ tokens.
All code, data and models are fully open!
hf.co/bigcode/starcoder2-15b