Open and responsible research and development of large language models for code. #BigCodeProject run by @huggingface + @ServiceNowRSRCH

Joined August 2022
BigCode@BigCodeProject
8 Oct 2025
Beyond human votes, we release two new benchmarks: - BigCodeReward: tests reward models on 4.7K human preference votes. Execution feedback improves judgment accuracy. - AutoCodeArena: automated evaluation of 20+ LLMs. GPT-5 leads, followed by Claude-Opus-4 and Claude-Sonnet-4.
1
7
396
BigCode@BigCodeProject
8 Oct 2025
In 5 months, BigCodeArena collected 14K conversations & 4.7K preference votes across 10 frontier LLMs. Findings: 1. o3-mini & o1-mini consistently top Elo rankings 2. Claude-3.5-Sonnet excels in matched-language settings 3. Previous open models still lag
1
5
429
BigCode@BigCodeProject
8 Oct 2025
Why does this matter? Benchmarks like HumanEval only scratch the surface. Reading code “by eye” is error-prone. True quality emerges when you actually run it: web apps render, games play, edge cases break. BigCodeArena makes execution feedback the default.
1
6
532
BigCode@BigCodeProject
8 Oct 2025
Introducing BigCodeArena, a human-in-the-loop platform for evaluating code through execution. Unlike current open evaluation platforms that collect human preferences on text, it enables interaction with runnable code to assess functionality and quality across any language.
4
29
3
82
45,092
BigCode@BigCodeProject
17 Jul 2024
Releasing BigCodeBench-Hard: a subset of more challenging and user-facing tasks. BigCodeBench-Hard provides more accurate model performance evaluations and we also investigate some recent model updates. Read more: huggingface.co/blog/terryyz/… Leaderboard: huggingface.co/spaces/bigcod…
23
6
96
35,652
BigCode@BigCodeProject
18 Jun 2024
Introducing 🌸BigCodeBench: Benchmarking Large Language Models on Solving Practical and Challenging Programming Tasks! BigCodeBench goes beyond simple evals like HumanEval and MBPP and tests LLMs on more realistic and challenging coding tasks.
9
61
15
210
102,314
BigCode@BigCodeProject
29 Apr 2024
Releasing StarCoder2 Instruct! 🚀 Achieves 72% HumanEval score using only self-generated content without any GPT-3.5/4 data. This work demonstrates that self-instruct works already well at the 15B scale without data from proprietary models! Read more: huggingface.co/blog/sc2-inst…
4
73
5
283
38,707
BigCode@BigCodeProject
28 Feb 2024
This is the result of hard work by the BigCode community and supported by @ServiceNowRSRCH, @huggingface and @nvidia to train the 3B, 7B and 15B models! The Stack v2 was built with @SoftwareHeritage and the full processed training data is coming soon! huggingface.co/datasets/bigc…
4
1
35
4,904
BigCode@BigCodeProject
28 Feb 2024
StarCoder2 performs well on a wide range of coding and math tasks. The 15B model is best in its class, while the 3B model is at the performance of StarCoder1-15B. This makes StarCoder2 models more efficient and performant! Read the full report: hf.co/bigcode/report
2
3
51
4,597
BigCode@BigCodeProject
28 Feb 2024
Introducing: StarCoder2 and The Stack v2 ⭐️ StarCoder2 is trained with a 16k token context and repo-level information for 4T+ tokens. All built on The Stack v2 - the largest code dataset with 900B+ tokens. All code, data and models are fully open! hf.co/bigcode/starcoder2-15b
12
202
46
660
223,121