@TimTheSloth

Brand and Vibes @scale_ai. Unashamed generalist.

Austin, TX
Joined April 2009
how do you win consumer? mascots with butts
2
2
21
1,953
Tim Bauer retweeted
We're releasing HLE-Diamond: a refined version of Humanity's Last Exam, built with @CAIS. A year of review and community feedback went into refining this subset to make it more reliable for measuring frontier models. Top model tested 60.6% overall. We expect HLE-Diamond to carry signal for the next 6-12 months.
4
15
5
96
7,036
Tim Bauer retweeted
there is one (1) major firm that’s friendly with the gov + F500 and has done embedded eval consulting for years and has a real track record of ai safety expertise but is now largely divested from commercial conflicts with the leading labs why is nobody talking about scale ai
15
11
2
315
36,514
The future is here – at Austin, TX
7
495
Tim Bauer retweeted
we're building a world class design team at @modal here in new york to build the future of ai infrastructure. we're looking for a sr/staff brand designer to join our team: folks who want to build new systems and tools, and rethink how brands are made from the ground up. DMs open, listing in the thread.
7
8
3
73
6,322
I'm really proud of this work from the team. Massive thanks to Travis Britton for leading the identity redesign with contributions from @nocellcoverage, @lin_zagorski, @SnakeFooc, and @rickyrauch. @active_theory was an amazing partner on the scale.com refresh. They nailed the strategy and execution on a very tight timeline. More to come!
Still shipping nonstop. Here’s our new look
2
3
24
1,346
Tim Bauer retweeted
gm
39
64
3
1,240
52,457
I have massive respect for every designer who didn't weigh in on the Instagram wordmark.
1
8
448
Tim Bauer retweeted
📣Call for contributions + co-authorship! RSI Bench is our ongoing effort to evaluate whether AI agents can develop the capabilities required to advance AI R&D. Got a frontier AI research problem on your mind? Submit it to RSI Bench, and set the standard the whole industry uses to measure progress! Contributors receive: → Co-authorship on the RSI Bench research paper → $2,000 per accepted task for the initial 50 tasks → Modal compute credits to build + iterate → Access to the RSI Bench research community Register below for full requirements.
Launching rsi-benchmark.com: The work of AI R&D has always belonged to humans. For the first time, though, it no longer seems certain that it always will. Recursive self-improvement is within a line of sight. It may still be far, but it is close enough that we should start measuring it.
13
30
11
297
72,609
It's hard not to not be nostalgic for the dying embers of American monoculture. – at San Francisco, CA
puddle of mudd blurry mtv spring break 2002 this video is a perfect encapsulation of everything that's been stolen from us
3
573
Welcome @fdesouza to @scale_AI as CEO! I started Scale a decade ago at 19, and it has grown to be the backbone of the AI world, powering frontier labs, Fortune 500 companies, and the US Government. Francis is an exceptional leader and the right steward. After spending a lot of time together, I have no doubts he’ll take Scale to greater heights. Thank you @jdroege for your leadership as interim CEO 🙏
Big News: I’m joining @scale_AI as CEO, starting August 10. Scale sits at a rare intersection, working with the top AI labs to push the frontier while helping enterprises and governments actually deploy AI they can verify and trust. I’m excited to lead a company with a mission to develop reliable AI systems for the most important decisions. More to come.
59
55
8
1,121
222,912
Every researcher in the world has to press a red or blue button. If more than 50% press the blue button, all frontier models go through mandatory safety testing. If less than 50% press blue, only models from researchers who pressed red are tested. Which button would you press? – at Hoboken, NJ
100%blue
0%red
4 votes • Final results
1
1
191
Uh, thanks I guess? – at San Francisco, CA
10
1,170
This is where I think the market has @scale_AI wrong. People still see us as a "data company". We are so much more! We're the eval and reward layer behind most of the frontier labs. We're the ones who work out what "good" even means for a model. That's exactly what enterprises need now, except it has to run inside their own walls!
1
1
5
368
SBF being a bottom 20% League of Legends player was a clear red flag in hindsight. – at San Francisco, CA
1
6
1,550
In SF for the next month. Let's hold hands on the cable cars. – at San Francisco, CA
4
17
1,300
Heard you talking shit and... actually you make some good points. – at San Francisco, CA
We appreciate the community's feedback on SWE-Bench Pro. Much of it maps to changes already underway in v1.1, which we've been building for a while. Keeping evals current with frontier models is hard, and we're always iterating. SWE-Bench Pro Verified coming soon. 👀
14
4,430
My first vibe coded project is live: brand.scale.com/ I always wanted to build a brand guide with some quality-of-life features that were accessible to partners/vendors/rank and file at the company. Hopefully this leads to fewer DMs asking where the logo files are. 😅
4
2
2
21
2,398
– at San Francisco, CA
1
11
798
Austin may not have any frontier AI labs, but it does have a place where you can have coffee delivered to you by a fuzzy hand reaching through a hole in a cement wall. – at Austin, TX
3
5
391