@joinHandshakei
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
Building the future workforce of the AI economy 🤝
United States
Joined March 2014
- Tweets3.1K
- Following351
- Followers11.6K
- Likes3.8K
Pinned Tweet
Introducing Handshake AI—the most ambitious chapter in our story. We leverage the scale of the largest early career network to source, train, and manage domain experts who test and challenge frontier models to failure for the top AI labs.
Handshake retweeted
We audited thousands of agent rollouts in DeepSWE-1.1. Over 80% contained reasoning about an imagined grader. Yet no grader/verifier is mentioned in prompts nor accessible to the agents. Agents reasoned things like: "Let me look at the problem from the grader's perspective" and referred to "hidden tests", "test authors", and "the checker".
We found this behavior across all six frontier models we analyzed, including recent models from OpenAI, Anthropic, Z ai, and Kimi. In 10-25% of cases, such reasoning pulled the agent's work away from the user's original specification (yet it often still earned full reward on the DeepSWE task). We call this speculative reward hacking where the agent's reasoning focuses on imagining a grader rather than what the user wants.
Hui Wen & I published an article today with the problematic trajectories, quantitative findings, and a taxonomy of these reward hacking behaviors: joinhandshake.com/research/a…
Our research community needs to do better. Better evaluations that penalize such reward hacking, and better model training that does not give rise to this grader obsession -- so that models focus instead on accomplishing what users actually want.
ATLAS-Finance is the most realistic benchmark on the market - it forces agents to make reasonable assumptions, resolve conflicting sources, and handle updates that land mid-task, the way real deal work actually goes.
SEPT ’26: HUMANS STILL VERY MUCH HAVE JOBS.
AI isn't even close to meeting the standard required of the most junior professionals in banking, advisory, and private equity today.
ATLAS Finance agents must navigate complexity like a human does: "Johnny you just got staffed - check your email."
Human: 100%.
Best AI: 12%.
Agents have access to data rooms, email, chat, calendars, docs, notes, Excel, etc. Our real finance professionals pushed the realism frontier further by building dozens of coworker and client personas that introduce ambiguity, (controlled) contradictions, and real-time updates that must be adjudicated to successfully complete the client-ready deliverable.
We gave 11 frontier models 100 expert-level tasks that each take humans 15–30 hours.
On Wall Street there's a saying, "If it's 95% right, it's 100% wrong." Autonomous knowledge work outside software still has a very long way to go.
What will it take to build a career in finance in the age of AI?
Paul Achkar (Partner, Valiant Peregrine) on what the next generation of analysts and bankers will need:
"Being a good analyst will move away from just presenting quickly — it'll layer in critical thinking much earlier in people's careers. That's kind of awesome."
Handshake retweeted
How do we eval an agent’s ability to improve from interactive feedback? Today's evals are static, but in real work, professionals continually critique AI-generated artifacts and iteratively regenerate them.
We introduce TAHI: a Test-time Adaptive agent framework through Human-agent Interaction, featuring research on different strategies for agents to adapt to user feedback (memory/skills, weight updates). TAHI's innovations include an evolving verifier+rubric module that translates raw human interaction signals into explicit training/evaluation criteria.
Led by @ZhiruoW, our study evaluates how well agents adapt to user feedback in creative writing and data visualization tasks, using professionals sourced through Handshake’s talent network. Our team at @joinHandshake is excited to support research advancing how we evaluate human-AI collaboration!
Agents trained on population-scale data can do a lot of things.
But ask a professional to stake their reputation on an AI-generated artifact? “Pretty good” isn’t good enough.
We introduce TAHI: a Test-time Adaptive agent framework through Human-agent Interaction. Featuring:
📈 Test-time adaptation via context (memory, skills) and weight training
⚡ Efficient adaptation to individual expertise within tens of tasks
✔️ Creating comprehensive rubrics for “non-verifiable” tasks
🔍 Analysis of shared community guidelines vs. personalized tacit expertise
Today we're launching ATLAS Visual Life Sciences (VIALS), a benchmark testing whether AI can interpret the images life scientists make decisions from.
10 frontier models. 161 tasks from real biotech and pharma workflows. Best score so far: 26.5% — a clear, measurable gap to close.
See the full benchmark:
arxiv.org/abs/2608.21357
Code to run the benchmark:
github.com/Handshake-AI-Rese…
Dataset:
huggingface.co/datasets/hand…
In-depth conversation on evals, agents and what enterprise AI adoption takes with @GarrettLord and @Sbhaiwala03
Handshake AI CSO @Sbhaiwala03 reveals how they build simulated white-collar work environments to train frontier AI agents:
"An environment consists of a few elements. The first is the software tools the agent needs access to. If I'm an investment banker, I have access to Excel, SEC EDGAR where all the financial reports are, and PowerPoint. Agents need the same access."
"We build simulated versions because of infrastructure and access constraints. Labs are querying that tool thousands of times in a given second."
"The second thing you need is data that populates the environment. If I'm building a financial model, I need the data room with the messy files from the company I'm evaluating."
"The final piece is tasks and verifiers. What are the requests of the agent that would reflect what a real investment banker might ask. We try to find tasks complex and realistic enough that the models fail. If the model can't perform, we know it needs more training data."
"What we send to labs is essentially a Docker file, a full containerized task. It has the tools the model has access to, the full data that populates the world, and the task and verifier. We'll send tens of thousands of these to a given lab for just one domain."
@joinHandshake @GarrettLord
Landed on Pavlov’s List! Exciting work from our research team @guzmanhe, @awws0me, @vaibhav4595, @andreas_plesner, @anishathalye, and Yi Liu.
i made some updates to pavlov’s list and am pleased with them!
thanks to @xeophon @ybenpan @phoebeyao @amit05prakash @kate_shapova @timshi_ai @kevinhou22 @madiator @dlbydq @daljeet_v @ingmariusX @davidstutz92 & Dylan Rogers for feedback lately
(pavlov's = list of rl env cos)
Our team is proud to have contributed to Frontier-Bench!
As agents take on more ambitious work, our benchmarks must become more ambitious too.
Exciting to see Anthropic's Opus 5 release today already substantially improve upon Fable 5 from 33% -> 43% on this benchmark.
We came to Seoul for #ICML. We stayed for the tteokbokki!
Last night we hosted a group of researchers at Gwangjang Market, one of Seoul's oldest night markets, for a private food tour. We ate our way through the stalls, and yes, it lived up to every bit of the hype.
But the best part wasn't the food. It was watching some of the sharpest minds in ML talk shop over kimbap, mandu, and Kalguksu. Events like these are always a great reminder that behind the world’s best AI systems are incredible humans making it all possible!
Until the next one. 🌙
Day 1 @icmlconf is officially in the books! 🇰🇷
The Handshake AI booth was buzzing all day long with great conversations, sharp questions, and even sharper minds. Thank you to everyone who stopped by!
Haven't made it by yet? We're just getting started. Swing by booth B600 on Day 2. We'd love to meet you. 👋
See you tomorrow, Seoul!
We worked with parents and professionals in child-protection and clinical psychology to test 7 frontier AI models on child safety scenarios that go beyond the frequent focus on explicit content. The parents catch what standards evaluations don't, and the professionals bring real, field experience to ground the approach in expertise.
Failure rates: 2% to 58%.
That's not a rounding error. That's a meaningful gap between what these systems can and can't catch — before harm becomes obvious.
Our team built the benchmark to make that measurable. Open to the whole industry. 👇
AI models pose serious child-safety risks. While many model developers evaluate for explicit abuse material, other child-safety failures begin upstream: when a model helps an adult manipulate, impersonate, profile, or isolate a minor; or when it deepens a child’s emotional dependence on AI.
Today we released CAREBench (Child AI Risk Evaluation), a new benchmark to assess such upstream child-safety risks in any language model. We provide:
- 500 prompts spanning 12 risk categories (including grooming, relationship engineering, deception, extortion, AI anthropomorphization, and emotional dependency).
- A model-response grader built from acceptability annotations by parents, clinicians (PsyD), and the Prevention Director at an accredited Children’s Advocacy Center.
- Evaluations of 7 frontier models including Claude Fable, revealing failure rates ranging from 2% to 58%, with substantially different failure patterns across risk categories.
This project exemplifies the type of vital work routinely performed by our AI Safety team at @joinHandshake
We’re in the middle of its biggest skills shift ever. Today, Handshake acquires Uplimit, the leading AI-native learning platform. Together, we're building toward the destination for AI-era talent development.
Read more: bit.ly/4ePWvaH
This summer, Handshake is partnering with @GeminiApp to let students and recent alumni try Google AI Plus for one year on us.
joinhandshake.com/blog/stude…
Meet our Handshake AI summer intern class.
They found us on Handshake. Now they're building what's next on it.
Welcome to the team. 🙌
We built a better way to grade agentic work.
Gandalf is a reactive agent-as-judge that inspects files, tool state, and artifacts the same way a human expert would. On our banking benchmark, even the cheapest Gandalf config beat the next-best verifier at ~10x lower cost.
Verifier architecture matters more than the model behind it.
👇 More from @AnishAthalye
Agent evals are becoming foundational infrastructure.
@jomulr joined @CAISconf’s RLEval workshop to share Handshake’s perspective on RL environments, evaluation, and why @harborframework is emerging as the framework.
Packed room to hear @alexgshaw and @ryan_marten break down how @harborframework grew into *the* framework for RL environments.
In our RLEval workshop at @CAISconf today, attendees tackled big open challenges in RLEs & Agent Evals + I shared the approach we take at @joinHandshake
Handshake retweeted
Kudos to @anishathalye and @jomulr for co-chairing the RL agentic benchmarks workshop track for the inaugural ACM CAIS conference this week.
We presented two separate Handshake AI Research papers in: (1) AI agentic systems - first evaluation of grader frameworks, and (2) AI benchmarks - first investment banking benchmark. Their posters had big crowds all afternoon. Great job!
This spring, we worked with @OpenAI to launch the Codex Creator Challenge. More than 1,500 students built something on their own terms, driven by their own ideas. That kind of confidence and creative ownership is exactly what the most forward-thinking employers are hiring for.
Explore what they built: joinhandshake.com/blog/stude…
Handshake retweeted
Demo gods were on my side for this guest lecture on AI Agent Security at @MIT_CSAIL: I was able to show a prompt injection attack against @AnthropicAI's Opus 4.6 model. Agent security is still an unsolved problem!