@datacurve

Research and data to advance frontier models.

San Francisco
Joined February 2024
Gemini 3.8 Flash achieves 73.7% on DeepSWE. It pushes the frontier with an impressive 8.2% increase over Gemini 3.7 Flash, at the same cost but with more steps and output tokens.
42
27
16
476
67,329
Gemini 3.8 Flash on DeepSWE 1.1, scores 73.7%!
292
133
128
3,642
541,816
Gemini 3.7 Flash debuts at 65.5% on DeepSWE. It delivers substantial improvements over 3.6 Flash, scoring +18.8% higher while costing less than half as much per task.
47
53
30
804
222,859
Claude Opus 5's results are now on DeepSWE With a score of 74%, it's the best long-horizon coding model we've seen.
104
56
42
1,384
245,252
Opus 5 costs 45% less per task on average compared to Fable 5.
1
4
98
8,048
Gemini 3.6 Flash is now on DeepSWE at 49%, roughly on par with Opus 4.8 at medium reasoning.
51
51
15
1,323
249,213
Gemini 3.6 Flash is significantly more efficient while scoring higher. It uses 65% fewer output tokens per task and costs 52% less per task compared to its predecessor.
2
2
95
9,465
Kimi K3 debuts at #3 on DeepSWE. It's the first open-weights model that delivers frontier-level performance, achieving results similar to Claude Fable and GPT-5.6 Sol.
107
342
127
3,664
669,231
GPT-5.6 tops the DeepSWE leaderboard at 73%. Sol, Terra, and Luna results are now available.
71
168
88
2,038
346,860
Datacurve is at @aiDotEngineer! Come by our booth in the expo hall to learn more about DeepSWE (and snag some merch). Plus, join us on Thursday @ 10:30am for our talk.
1
3
1
43
7,372
GLM 5.2 is now on DeepSWE as the top open-source model on our leaderboard. With a pass@1 score of 44% at max effort, GLM 5.2 is indisputable #1 open-source model besting Kimi K2.7 Code by 17%.
104
239
119
2,787
586,062
Claude Fable 5 debuts at #1 on DeepSWE. It outscores the previous best by 3% and sets a new state-of-the-art on our long-horizon coding benchmark.
108
84
55
1,621
474,525
Fable 5 scores 70% pass@1 and tracks GPT-5.5 on cost-performance at the default high effort. Kimi K2.7 also joins the leaderboard with a score of 31%.
9
8
4
136
26,072