@trybasisi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- Web
Account-level information from X, not a live location or the device used for a specific post.
Agents you can actually rely on. Built specifically for accountants. Join us https://nitter.cf/t.co/opwj4XyvS9
New York
Joined January 2023
- Tweets90
- Following6
- Followers3.1K
- Likes140
Read about how we build agents with @cursor_ai !
Maybe counterintuitive but I think IDE's will make a comeback.
Agents allow you to build far more complex systems that are NOT provable at runtime. So if you want to be a good agent engineer you better read the damn english
cursor.com/blog/basis
Every person at Basis, no matter what team they're on, has agents doing work alongside them within minutes of their first day.
We’ve been early to integrate AI into every process in our company - in part because we invested in building a team dedicated to making this happen.
If that sounds interesting, you should come join the Atlas team!
Read more from OpenAI about how we’re using Codex to make onboarding teachable with agents:
openai.com/index/ai-native-c…
Last week, @mitch_troy joined @AI_in_the_AM to discuss accounting as “an intelligence over the economy,” how AI can help firms create capacity for growth, and how Behavior Specs can help supervise long-running agents.
Watch the full conversation:
Accounting firms can increase their revenue 50% year on year with Basis
nitter.cf/i/broadcasts/1nKOLQbAk…
Basis retweeted
☀️ AI:AM is live.
Today: @mitch_troy, co-founder of @trybasis, on running long-horizon agents on real money, 9:30a PT.
Then Jay Dawani, co-founder & CEO of Lemurian Labs, on whether CUDA's moat is really a software problem, 10:15a PT.
nitter.cf/i/broadcasts/1aKbdEopl…
“It’s just doing the work and I’m just reviewing it.” — Todd Whitcomb, Managing Partner at Whitcomb & Sheridan.
For the past few months, we’ve been building Basis for Small Firms with a community of small firm owners across tax, CAS, and audit.
Accountants in our community beta are already getting 10–20 hours back each week.
Today, we’re inviting more firms in.
Firm leaders can join the beta here: getbasis.ai/solutions/small-…
“Basis is the end all be all,” one of our customers told us.
It’s only possible because our customers built Basis with us.
At our first customer conference, accounting leaders kept returning to three ideas:
1. Capacity no longer has to constrain growth.
2. Change management makes or breaks a transformation.
3. Firms have to redesign how they train and upskill their teams.
Thank you to every customer who shared their perspective and challenged our thinking.
You are building the future of accounting. We are grateful to get to do it with you.
Listen to our cofounder @mitch_troy talk about the specifics of how we build agents to be reliable and trustworthy not just smart
How to build long-horizon AI agents: behavior specs, ontologies, process supervision - my conversation with @mitch_troy, co-founder of @trybasis
01:09 Why Everyone at Basis Was Whispering to AI when @steph_palazzolo walked in
04:12 Accounting as "an Intelligence Over the Economy"
06:11 What Makes an Agent Truly Long-Horizon
08:24 Inside an Autonomous, Multi-Day Tax Return
10:19 Agents That Hand Off Like Senior Engineers
11:17 A Brief History of Agents: From ReAct to Today
12:33 Why LLMs Have No Long-Term Memory
14:13 Why AutoGPT Didn't Live Up to Its Promise
15:51 The Three Breakthroughs: Opus 3, o1, o3
17:07 Why Reasoning Models Unlocked Agents
18:23 "Let's Verify Step by Step": The Road Not Taken
20:32 Pushing Back on the METR Chart
22:09 Why Coding Agents Won First
25:14 Why Real-World Agents Are Harder
26:55 How Accountants Verify Non-Deterministic Work
29:18 You Can't Scale Tax Returns Like Math
33:16 100 Evals Pass - So What?
35:53 Right Answer, Wrong Process
36:37 Behavior Specs, Explained
39:58 How Specific Should Behaviors Be?
42:18 Context Is Runtime Training Data
44:21 Who Judges the Judge?
46:45 The Move 37 Objection
50:02 The Magic Box Mental Model
52:41 "Nothing Has Changed Since o3"
54:56 Open-Sourcing Behavior Specs with @ankrgyl @braintrust
59:45 Ontologies: A World for Agents to Live In
01:04:20 Documentation as Codebase
01:06:33 Why the Founding Fathers Were Context Engineers
01:09:05 Onboarding 300 Brilliant Alien Employees
01:11:10 Self-Improving Agent Systems
01:12:50 The Context Mistake Agent Builders Make
01:14:29 RL on Behavior Adherence
01:17:01 Will the Bitter Lesson Swallow the Harness
01:18:46 "Technical Moats Are Not Real Moats"
01:21:03 Advice for AI Builders
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
Basis retweeted
Excited to announce the Basis End-to-End Tax Platform, the first production deployment I’m aware of built around truly proactive agents.
By proactive agents, I mean agents that operate against goals and jobs without needing a human to constantly give them work.
I think this is the next logical evolution of long-horizon agents.
Basis agents take responsibility for the entire first pass. They gather the required information and evidence from clients, prepare and review the return, and bring in accountants when an issue requires judgment.
The agents keep the work moving. Accountants can inspect every step and control every decision.
The future of tax is end-to-end intelligence.
Today at Basis Tax Day, we hosted leaders from 30 of the Top 100 accounting firms to explore what changes when Basis takes the first pass on complex returns.
We launched Basis End-to-End Tax for that future. So accountants get out of data entry and start with review.
The 2027 busy season will feel very different with Basis.
The future of tax is end-to-end intelligence.
Today at Basis Tax Day, we hosted leaders from 30 of the Top 100 accounting firms to explore what changes when Basis takes the first pass on complex returns.
We launched Basis End-to-End Tax for that future. So accountants get out of data entry and start with review.
The 2027 busy season will feel very different with Basis.
Basis retweeted
Nobody wants to be an accountant anymore. Basis CEO Matt Harpe thinks he can help.
Founded in 2023 by Harpe and @mitch_troy, Basis builds AI agents for accounting firms, automating hours of number crunching and copy-paste into a review process that takes minutes.
@trybasis already works with 25% of the top 150 firms in the U.S. and reached a valuation of $1.15B in February.
And far from eliminating jobs, Harpe believes his startup will help human accountants get paid.
"It will increase how valuable accountants are, because they will be able to accomplish so much more," he says.
On The Upstarts Podcast, Harpe shares:
▪️how Basis can stay ahead of the big labs like Anthropic and OpenAI in its vertical
▪️why outcome-based pricing is nice work, if you can get it
▪️how he pitches top AI talent in New York on building accounting tools, of all things.
Plus, he shares his Upstart Moment: resisting calls to chase short-term revenue until his product felt ready for primetime.
TIMESTAMPS
00:00 Introduction
01:47 The basis for Basis
07:25 Why no one wants to be an accountant any more 14:37 CEO Matt Harpe's founder journey
19:46 "The home for applied AI talent in New York"
22:58 Winning over skeptical accountants as customers
28:41 A big bet on agents from the start
31:31 Outcome-based pricing and its challenges
35:52 Competing with the big AI labs like OpenAI
41:11 Making human accountants more valuable
A big thanks to our season sponsor, @Rippling, for supporting founder stories like this 🫡
Enjoy!
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
Basis retweeted
I have long felt the diminishing effectiveness of LLM-as-a-judge for agents, and one day over coffee @mitch_troy gave me a rant that helped clarify why. The ground truth for an agent isn't its output, it is how the agent should behave.
This is important for two reasons:
(1) it is very difficult to generate good ground truth values
(2) to effectively debug why an agent produced invalid outputs, you need to introspect its behavior.
Mitch then walked through why this is so hard for their team's tax agents. The behaviors themselves are subtle, hard-earned lessons from analyzing lots of failures in their traces in Braintrust, and to effectively flag them, you must effectively document them.
That discussion led us to build behavior specs: a new open standard we're releasing with @trybasis that documents how agents should behave. The repo includes a definition of the spec along with open examples for how to write good specs, evaluate agents using it, and more.
Clear, human-articulated prose is the highest leverage way to drive agentic systems to produce great outcomes. Behavior specs provide a framework to do that with evals. I'm super excited to work on this in the open, and would love to get feedback from others on how we can make this spec maximally useful.
Please try it out, and share your thoughts!
agentbehavior.dev/
Out of the box, long-horizon agents struggle to accurately perform end to end work in the real economy (outside of coding) because those tasks are not easily verifiable, the data is hard to scale, and going from inputs to real outcomes can actually take many days.
Even if you had a reliable way to verify outcomes at scale (and weren’t bothered by the multi-hour iteration loops), the sheer volume of decisions by the agent that occur in a multi-hour job makes it hard to know whether performing well will generalize to production.
Over the last two years at @trybasis, we've been solving this problem by supervising the process our agents take to get to outcomes, rather than just looking at whether the outcome itself is correct.
We think this is the key to building production agents at scale.
It's what has allowed us to run agents in production that operate for hours, sometimes days, and reliably perform tasks like entire complex tax returns end to end.
Today, alongside @braintrust, we're open sourcing a standard for defining, evaluating, and eventually rewarding agent behaviors.
Thread below with all the details on how we’re scaling behaviors to close the loop for long-horizon agents.
To trust any long-horizon agent, you have to supervise how it works over a full trajectory, rather than just looking at its final answer.
Behavior specs are now an open standard, built with @Braintrust and based on how we actually evaluate our production agents: agentbehavior.dev
Out of the box, long-horizon agents struggle to accurately perform end to end work in the real economy (outside of coding) because those tasks are not easily verifiable, the data is hard to scale, and going from inputs to real outcomes can actually take many days.
Even if you had a reliable way to verify outcomes at scale (and weren’t bothered by the multi-hour iteration loops), the sheer volume of decisions by the agent that occur in a multi-hour job makes it hard to know whether performing well will generalize to production.
Over the last two years at @trybasis, we've been solving this problem by supervising the process our agents take to get to outcomes, rather than just looking at whether the outcome itself is correct.
We think this is the key to building production agents at scale.
It's what has allowed us to run agents in production that operate for hours, sometimes days, and reliably perform tasks like entire complex tax returns end to end.
Today, alongside @braintrust, we're open sourcing a standard for defining, evaluating, and eventually rewarding agent behaviors.
Thread below with all the details on how we’re scaling behaviors to close the loop for long-horizon agents.
Check out this great demo from the Braintrust team!
nitter.cf/braintrust/status/2082…
To build long-horizon agents you can trust, you have to supervise the process, not just the final output. Agents that make hundreds of decisions behave in ways that can't be reduced to a single outcome metric.
Behavior specs are an open standard for defining and evaluating how your agent behaves across a whole trajectory. Codify how your agent should work, then turn each spec into a standing eval.
Comes with example behavior specs, authoring skills, and judge prompts. Built in collaboration with @trybasis.
Everything is open source at agentbehavior.dev.
Read more → braintrustdata.link/behavior…
Excited to share our "Ode to Accounting" short film.
It's a tribute to the profession, tracing accounting's 10,000-year arc from clay tokens to double-entry to AI agents.
Accounting's next chapter is being written now, by the people doing the work. It's their era to shape.
An Ode to Accounting
A short film tribute to the accounting profession, from the team at Basis.
Alongside the short film is an accompanying essay, and love letters from practitioners and leaders across the industry about what accounting means to them.
Check it out: odetoaccounting.com/