@ValsAI

Public LLM Evaluation // https://nitter.cf/t.co/rNS39rhVQ9 @8vc @BloombergBeta @pearvc @a16z @HRTVentures @nextladder

San Francisco, CA
Joined March 2024
Existing coding benchmarks stop at the first working version. We are releasing Vibe Code Bench 1-100 today, to measure what comes next. This benchmark asks if models can handle a large number of modifications to a product, without breaking what already works.
13
8
4
111
12,371
Vals AI retweeted
If you're interested in making coding benchmarks for your internal repos and workflows, check out vals.ai/vals-smith. It allows you to turn your codebase into a customized benchmark.
if you are building a product using AI, you should be spending >25% of your time making benchmarks and trying to get the model labs to care about said benchmarks easiest path to accelerate your progress as a company
1
1
1
20
3,390
We evaluated Grok 4.7 across the Vals benchmark suite. It ranks #24 on the Vals Index at 54.2%, down 5.0 points from Grok 4.6 (#14, 59.2%), but still ahead of Grok 4.5 (#30, 51.5%). Grok 4.7 improves the most on legal and medical work.
44
28
40
564
315,430
Post-launch, xAI has updated its SDK, significantly improving its performance. It is now #10 on the Vals Index.
1
1,725
Post-launch, xAI has updated its SDK, significantly improving its performance. It is now #10 on the Vals Index.
11
4
9
194
39,870
Previously, the xAI Responses API returned encrypted_reasoning_content on each turn. This acts as a pointer for future turns to reasoning the model has already done. However, this was not returned by default when using the native xAI SDK, meaning that reasoning was being lost between model turns. xAI is rolling out a change to make encrypted_reasoning_content returned by default across all APIs.
3
3
4
94
26,309
All results currently on Vals AI reflect the updated configuration. Additional benchmarks will be published to the website as they finish.
33
3,165
Grok 4.7 also costs slightly more per Vals Index test, at $4.78 versus $4.34 for Grok 4.6. It uses roughly 15% more input tokens, but far fewer reasoning tokens.
1
23
6,850
This model has 500k max output tokens and 500k context window. It was run on SpaceXai’s default provider setings: Temperature=0.7, Top P= 0.95, Top-K= default, and run on xhigh reasoning effort. For full results, visit: vals.ai/models/grok_grok-4.7
1
27
6,518
AI models are advancing faster than legacy benchmarks can keep up. @TechCrunch @LucasRopek1 visited us to see how we’re building independent evaluations grounded in real-world work and designed to measure both capability and risk
5
1
41
5,789
Grok Voice Transcribe 2.0 is now #2 on Voice Code Bench—up 13pp and 12 spots from its predecessor, Grok Voice Transcribe 1.0.
2
1
23
3,442
There was an incredible improvement from the previous model, Grok Voice Transcribe 1.0 (52.0% TSR, 84.7% CTEM). Almost all of it came from difficult categories: IP addresses 48.0 to 84.0%, postal 50.0 to 70.0%, email 63.1 to 80.0%, at the same per-task price.
1
577
GPT-6 Astra had gotten further than any AI system had ever gone in Minecraft. It was able to set up a semi-automatic blaze farm, allowing it to collect 6 blaze rods. It then located a warped forest, where it killed 6+ endermen and collected 3 pearls. As thousands of viewers watched the stream, Astra put all of its valuable items in a chest. Then a creeper showed up and blew up both the chest and Astra’s bed, wiping out all of its progress. Astra seemed to get extremely frustrated after this happened. Astra said to itself, "ALWAYS CARRY CRITICALITEMS withkeepInventory; don'tstoreinunguardedchesteveragain."
127
496
191
7,930
967,686
What tasks are left that humans find easy but today's models still find hard? Two such tasks are computer use and games. We’re launching CUA-Bench, a benchmark testing how well AI can use a keyboard and mouse across 6 games (3 kept private) in real time. To saturate it, models will need to output real-time actions and learn continuously from video, not just text.
4
2,232
What tasks are left that humans find easy but today's models still find hard? Two such tasks are computer use and games. We’re launching CUA-Bench, a benchmark testing how well AI can use a keyboard and mouse across 6 games (3 kept private) in real time. To saturate it, models will need to output real-time actions and learn continuously from video, not just text.
22
18
14
295
29,237
This eval is extremely difficult, as all of the frontier models score below 20%, with saturation implying another breakthrough in the field.
1
1
17
1,126