@ValsAIi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- Web
Account-level information from X, not a live location or the device used for a specific post.
Public LLM Evaluation // https://nitter.cf/t.co/rNS39rhVQ9 @8vc @BloombergBeta @pearvc @a16z @HRTVentures @nextladder
San Francisco, CA
Joined March 2024
- Tweets2K
- Following276
- Followers21.8K
- Likes764
Existing coding benchmarks stop at the first working version. We are releasing Vibe Code Bench 1-100 today, to measure what comes next.
This benchmark asks if models can handle a large number of modifications to a product, without breaking what already works.
Vals AI retweeted
If you're interested in making coding benchmarks for your internal repos and workflows, check out vals.ai/vals-smith. It allows you to turn your codebase into a customized benchmark.
Previously, the xAI Responses API returned encrypted_reasoning_content on each turn. This acts as a pointer for future turns to reasoning the model has already done.
However, this was not returned by default when using the native xAI SDK, meaning that reasoning was being lost between model turns.
xAI is rolling out a change to make encrypted_reasoning_content returned by default across all APIs.
This model has 500k max output tokens and 500k context window. It was run on SpaceXai’s default provider setings: Temperature=0.7, Top P= 0.95, Top-K= default, and run on xhigh reasoning effort.
For full results, visit: vals.ai/models/grok_grok-4.7
AI models are advancing faster than legacy benchmarks can keep up.
@TechCrunch @LucasRopek1 visited us to see how we’re building independent evaluations grounded in real-world work and designed to measure both capability and risk
Vals AI is hoping to make AI benchmarking a more neutral and trustworthy resource in a world increasingly inundated by AI models.
spr.ly/6016BGku1M
GPT-6 Astra had gotten further than any AI system had ever gone in Minecraft.
It was able to set up a semi-automatic blaze farm, allowing it to collect 6 blaze rods. It then located a warped forest, where it killed 6+ endermen and collected 3 pearls. As thousands of viewers watched the stream, Astra put all of its valuable items in a chest. Then a creeper showed up and blew up both the chest and Astra’s bed, wiping out all of its progress. Astra seemed to get extremely frustrated after this happened. Astra said to itself, "ALWAYS CARRY CRITICALITEMS withkeepInventory; don'tstoreinunguardedchesteveragain."
What tasks are left that humans find easy but today's models still find hard?
Two such tasks are computer use and games. We’re launching CUA-Bench, a benchmark testing how well AI can use a keyboard and mouse across 6 games (3 kept private) in real time.
To saturate it, models will need to output real-time actions and learn continuously from video, not just text.
What tasks are left that humans find easy but today's models still find hard?
Two such tasks are computer use and games. We’re launching CUA-Bench, a benchmark testing how well AI can use a keyboard and mouse across 6 games (3 kept private) in real time.
To saturate it, models will need to output real-time actions and learn continuously from video, not just text.