@Agentese_AIi
iAccount based inSoutheast Asia
About this account
- Account based in
- Southeast Asia
- Connected via
- Southeast Asia App Store
Account-level information from X, not a live location or the device used for a specific post.
One dashboard for multiple AI agents. AI that acts, not just responds. https://nitter.cf/t.co/f8R5RQSrM6
Joined April 2026
- Tweets757
- Following54
- Followers276
- Likes645
Pinned Tweet
Building a company used to mean hiring a team.
Now it means deploying one.
We stopped building disconnected AI “tools” and started hiring AI “employees.”
Meet the Agentese workforce for your company:
🎙️ Meeting Copilot - Live meeting assistant
Answers hard questions out loud, in real time, during your meetings. No prep and no recaps, just the answer the moment the room needs one.
meeting-copilot.agentese.ai
🦞 LifeClaw - Personal call & admin agent
Makes your phone calls, bookings, and disputes for you: the cancellation, the appointment, the hold music, all handled.
lifeclaw.agentese.ai
🥔 Little Pink Potato - Automated job application agent
Applies to jobs for you at a scale no human could, screening thousands of roles, scoring the real matches, applying and negotiating based on your preference.
littlepinkpotato.com
☕ Kopi - Approval assistant for Lark
Clears the routine approvals piling up in your Lark queue, auto-handling the obvious yeses and flagging only what needs your judgment.
hellokopi.com
🔒 CodeAutrix - Smart contract & AI skill auditor
Audits and stress-tests smart contracts and AI skills in minutes, surfacing the critical flaw before you ship.
codeautrix.com
🍀 Yarrow - Multi-agent simulation platform
Ask a question, get a calibrated probability backed by a full chain of evidence and reasoning. Predictions you can grade, not opinions.
yarrowlab.ai
Stop managing tools, start delegating. With Agentese. 🚀
Six months later, has your view on D/ACC vs E/ACC changed?
The hardest part of D/ACC, to us, is the incentive to race.
A team can believe AI needs safeguards and still accelerate because it expects competitors to move faster. Concern about risks coexists with fear of being left behind.
This makes voluntary restraint fragile. Even if everyone prefers mutual caution, nobody wants to be the only one exercising it.
D/ACC’s emphasis on accelerating defensive technology offers a path that does not require everyone to slow down together. People have reasons to invest in protection regardless of what competitors believe.
But it leaves a difficult question:
Can defenses keep pace with the capabilities they are meant to contain?
Where restraint is necessary, cooperation must be verifiable, and breaking agreements must carry a cost. Otherwise, the advantage goes to whoever ignores them.
We find D/ACC compelling as a direction. We are less convinced it holds under competitive pressure without changing the incentives that drive the race.
What’s your take on this?
E/ACC vs. D/ACC: THE DEBATE
@VitalikButerin thinks slowing down AGI by four years is worth it. @beffjezos thinks that's exponential opportunity cost. They debated it live, moderated by @eddylazzarin and @shawmakesmagic.
00:00 Opening
07:02 Thermodynamics and first principles
16:04 Acceleration, entropy, and civilization
28:29 The core disagreement
32:42 Comparing and contrasting e/acc and d/acc
36:20 Open source, open hardware, and local intelligence
54:18 Should AI be slowed down?
1:02:35 Autonomous agents and artificial life
1:21:07 Crypto as the trust layer between humans and AI
1:35:37 Closing arguments
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
If Agentese products have their own apps, which one would be your favorite?
If you tell me your AI can:
- Sit on hold with the gym for 40 minutes to cancel my membership
- Negotiate my salary so I do not have to be polite
- Clear my 400 Lark approvals at 9 PM on a Sunday
- Audit my vibe-coded app before I push to production
- Forecast the next policy shock and show me its actual track record
Then:
Agentese AI retweeted
Say you want to know: will the Fed raise rates next month?
If you're a regular person, you Google it. You find three headlines — one says yes, one says no, one says maybe. You pick the one that sounds most convincing, or the one from the source you trust most. That's not a forecast. That's a vibe check.
If you ask ChatGPT, it gives you a confident paragraph. "Based on current economic indicators, it appears likely that..." It sounds smart. But it has no track record, no accuracy score, and if you ask the same question tomorrow with slightly different wording, you might get a different answer. It doesn't know how often it's been right before — because nobody's counting.
If you check Polymarket, you get a number — say 55%. That's better. Real money is behind it. But it's just a price. There's no explanation for why it's 55% and not 40%. No evidence trail. No reasoning you can walk through. And if the number changes overnight, you have no idea what drove it.
☘️ Here's what happens inside Yarrow.
The same question goes to many independent analysts. They're running on three different AI models from two different providers. None of them can see the market price. None of them can see each other's work.
Each one gets the same frozen evidence pack — official data, dated sources, verified statistics. No live search. No Twitter. No vibes. They each have to produce a complete chain: what mechanism drives the outcome → what the data says → what probability they assign.
Before any of that ships, the evidence pack runs through five checks: are the numbers internally consistent? Is anything missing that should be there? Is there any information that could only be known after the fact? Has an independent red team tried to break the argument?
The final number is the median of the valid analysts — not a vote, not a discussion, not a compromise. Just the middle of six independent judgments.
Then it gets frozen. Locked in git. Published. And when the outcome is known, scored — right next to every other prediction we've ever made.
Google gives you headlines. ChatGPT gives you confidence. Polymarket gives you a price. Yarrow gives you a probability with the evidence behind it, the reasoning you can audit, and a track record you can check.
That's the difference between looking for an answer and building one.
Agentese AI retweeted
On June 23, 2016, the UK held a referendum: remain in the EU, or leave.
Betting markets gave Remain an 85–90% probability. Pollsters predicted Remain.
The UK government had no contingency plan for Brexit, and the prime minister had effectively staked his political career on an outcome he himself could not predict with certainty.
Brexit won. 51.9% to 48.1%.
The pound plunged 10% overnight, its largest single-day drop since 1985. David Cameron resigned the following morning. Global markets lost $2 trillion in value within two days. And Britain spent the next four years trying to figure out what “Brexit” actually meant.
The data telling a different story was already there.
Online polls had consistently been closer to the truth than telephone polls. The turnout models had never been properly validated against a nationwide referendum. And anti-immigration sentiment in northern England had been systematically underestimated.
Betting odds were anchored to the same telephone polls that were later shown to be wrong.
No one stress-tested the consensus.
An 85% probability became a “fact,” and every plan was built around it.
☘️ Now imagine if the UK government had used a system like Yarrow before the vote.
The system receives one question: “Will the UK vote to leave the EU?”
Yarrow has independent analysts approach the question from different bodies of evidence: polling data using two different methodologies, regional turnout models, historical precedents from referendums versus elections, and sentiment data from regions that conventional polling struggled to reach.
Three analysts flag online polls showing Leave ahead.
Two build turnout scenarios in which older, non-urban voters who leaned toward Leave turn out at higher rates than assumed by the polling models.
One points out that betting markets are pricing the same telephone polls—not independent information.
The system ultimately returns:
Remain: 58%, Leave: 42%
And it adds: The 85% consensus probability for Remain depends heavily on a telephone-polling methodology that had never been validated in a nationwide referendum.
Cameron might still have held the referendum.But the Treasury might have had contingency plans ready. The Bank of England could have positioned itself in advance. And perhaps $2 trillion would not have vanished from global markets within 48 hours simply because everyone was blindsided by an outcome that, if someone had looked at the full evidence base, was clearly possible.
The most expensive predictions aren't the ones that turn out to be wrong.
They're the ones that are wrong—and that nobody prepared for.
Give LifeClaw the details. Get on with your day.
You need to move a reservation, but “any time on Friday” actually means after work, before 8, and only if there’s no extra fee.
Tell LifeClaw those details upfront so it can call the business with a clear request. Here’s what to include before handing over the call.
1. Your personal details
Provide your name and the contact or booking information needed for the call.
2. Your saved preferences
Share your usual dining times, dietary requirements, seating preferences, and budget so you do not have to repeat them next time.
3. The action you want
Specify whether you want LifeClaw to make a reservation, change an existing booking, cancel a booking, or ask about availability.
4. The business details
Provide the business name, specific location, and phone number, if available. If not, just let LifeClaw research for you.
5. Any additional instructions
Note acceptable alternatives, fees that require your approval, and what LifeClaw should do if your first choice is unavailable.
For example:
“Call Sakura in District 1 to book a table for two this Saturday. Prefer 7 p.m., but anything between 6 and 8 works. Ask me before accepting a deposit.”
Give LifeClaw the details, let it make the call, and then check the result.
In The Devil Wears Prada, Miranda’s cancelled flight takes over Andy’s dinner with her dad.
We gave the scene a different ending.
Andy opens @LifeClaw_AI, hands over the airline call, and goes back to dinner.
LifeClaw is an AI assistant that calls businesses on your behalf to handle bookings and cancellations.
You probably have a less dramatic call you’ve been putting off.
Agentese AI retweeted
Little Pink Potato giúp người dùng tìm việc và chuẩn bị cho các buổi phỏng vấn với các công cụ được hỗ trợ bởi AI.
Đây là một trong số ngày càng nhiều sản phẩm AI thuộc hệ sinh thái @Agentese_AI
Your best interview answer shouldn’t arrive after the interview.
The interview ends. You close your laptop.
And suddenly, you have the perfect answer.
You remember the project that would have shown exactly why you’re right for the role. During the interview, you spent 4 minutes explaining something barely relevant.
Now you’re wondering whether a follow-up email can somehow fit an entirely new interview inside it.
Give yourself a practice round before you get to that point.
Little Pink Potato’s mock interview feature gives you a place to rehearse before the real conversation.
🥔 Here’s how to make that practice useful:
1. Pick the role you actually want.
Keep the job description beside you. Practise explaining how your experience connects to what they need.
2. Do the first round before you feel ready.
Answer without checking your notes every few seconds. Find out which parts you can explain and where you start rambling.
3. Rework the answer that gave you trouble.
If you talked about what “we” achieved, explain your own contribution. If you made a broad claim, bring a concrete example.
4. Try again without memorising every word.
You want to be able to explain your experience even when the question comes in a form you didn’t expect.
Have the awkward first attempt while it’s still practice.
Agentese AI retweeted
What should a prediction actually look like?
Most AI tools give you a number. "There's a 70% chance this happens." Maybe a paragraph of explanation. That's it. No source trail, no what-if analysis, no way to check later whether 70% was the right call. 👀
Here's what we think a prediction should contain — at minimum — before anyone should take it seriously:
1️⃣ a probability that means what it says. When the system outputs 70%, events like that should happen about 70% of the time. Not "I'm pretty sure" dressed up as a number. Actual calibration, measured over hundreds of questions.
2️⃣ a reasoning chain you can walk through. Not "based on our analysis" — the actual evidence, the actual logic, step by step. Which data points support the conclusion? Which ones cut against it? If you can't trace from evidence to number, it's not a forecast — it's a vibe.
3️⃣ the conditions under which the prediction breaks. Every good forecast comes with its own kill switch. "We'd revise if X happens." "This falls apart if Y turns out to be true." If a prediction doesn't tell you what would prove it wrong, it's not brave — it's unfalsifiable.
4️⃣ a frozen record before the outcome. Anyone can claim they predicted something after the fact. The test is whether the prediction was locked in — with all its reasoning — before the answer was known. No edits, no revisions, no "well what I actually meant was."
5️⃣ a public track record. Not just the wins. The misses too. Side by side. Scored by the same rules. Because a system that only shows you its highlights isn't building trust — it's building a marketing deck.
☘️ This is the standard Yarrow holds itself to. Every prediction frozen in git. Every outcome scored. Every miss published alongside every hit.
Not because we're always right. But because a prediction you can't audit is just an opinion with better formatting.
Agentese AI retweeted
Little Pink Potato допомагає користувачам шукати роботу та готуватися до співбесід за допомогою інструментів штучного інтелекту. Це один із дедалі більшої кількості продуктів на базі штучного інтелекту в екосистемі Agentese_AI.
nitter.cf/Agentese_AI/status/210…
Your best interview answer shouldn’t arrive after the interview.
The interview ends. You close your laptop.
And suddenly, you have the perfect answer.
You remember the project that would have shown exactly why you’re right for the role. During the interview, you spent 4 minutes explaining something barely relevant.
Now you’re wondering whether a follow-up email can somehow fit an entirely new interview inside it.
Give yourself a practice round before you get to that point.
Little Pink Potato’s mock interview feature gives you a place to rehearse before the real conversation.
🥔 Here’s how to make that practice useful:
1. Pick the role you actually want.
Keep the job description beside you. Practise explaining how your experience connects to what they need.
2. Do the first round before you feel ready.
Answer without checking your notes every few seconds. Find out which parts you can explain and where you start rambling.
3. Rework the answer that gave you trouble.
If you talked about what “we” achieved, explain your own contribution. If you made a broad claim, bring a concrete example.
4. Try again without memorising every word.
You want to be able to explain your experience even when the question comes in a form you didn’t expect.
Have the awkward first attempt while it’s still practice.
Agentese AI retweeted
🔹 Little Pink Potato
Little Pink Potato membantu pengguna mencari pekerjaan dan mempersiapkan wawancara dengan berbagai tools berbasis AI.
Produk ini menjadi salah satu dari semakin banyak aplikasi AI yang hadir di ekosistem @Agentese_AI 👇
nitter.cf/Agentese_AI/status/210…
Your best interview answer shouldn’t arrive after the interview.
The interview ends. You close your laptop.
And suddenly, you have the perfect answer.
You remember the project that would have shown exactly why you’re right for the role. During the interview, you spent 4 minutes explaining something barely relevant.
Now you’re wondering whether a follow-up email can somehow fit an entirely new interview inside it.
Give yourself a practice round before you get to that point.
Little Pink Potato’s mock interview feature gives you a place to rehearse before the real conversation.
🥔 Here’s how to make that practice useful:
1. Pick the role you actually want.
Keep the job description beside you. Practise explaining how your experience connects to what they need.
2. Do the first round before you feel ready.
Answer without checking your notes every few seconds. Find out which parts you can explain and where you start rambling.
3. Rework the answer that gave you trouble.
If you talked about what “we” achieved, explain your own contribution. If you made a broad claim, bring a concrete example.
4. Try again without memorising every word.
You want to be able to explain your experience even when the question comes in a form you didn’t expect.
Have the awkward first attempt while it’s still practice.
Agentese AI retweeted
Little Pink Potato помогает пользователям искать работу и готовиться к собеседованиям с помощью AI-инструментов. Это один из растущего числа AI-продуктов в экосистеме Agentese_AI.
nitter.cf/Agentese_AI/status/210…
Your best interview answer shouldn’t arrive after the interview.
The interview ends. You close your laptop.
And suddenly, you have the perfect answer.
You remember the project that would have shown exactly why you’re right for the role. During the interview, you spent 4 minutes explaining something barely relevant.
Now you’re wondering whether a follow-up email can somehow fit an entirely new interview inside it.
Give yourself a practice round before you get to that point.
Little Pink Potato’s mock interview feature gives you a place to rehearse before the real conversation.
🥔 Here’s how to make that practice useful:
1. Pick the role you actually want.
Keep the job description beside you. Practise explaining how your experience connects to what they need.
2. Do the first round before you feel ready.
Answer without checking your notes every few seconds. Find out which parts you can explain and where you start rambling.
3. Rework the answer that gave you trouble.
If you talked about what “we” achieved, explain your own contribution. If you made a broad claim, bring a concrete example.
4. Try again without memorising every word.
You want to be able to explain your experience even when the question comes in a form you didn’t expect.
Have the awkward first attempt while it’s still practice.
Agentese AI retweeted
Little Pink Potato helps users find jobs and prepare for interviews with AI-powered tools. It’s one of a growing range of AI products across the @Agentese_AI ecosystem.
nitter.cf/Agentese_AI/status/210…
Your best interview answer shouldn’t arrive after the interview.
The interview ends. You close your laptop.
And suddenly, you have the perfect answer.
You remember the project that would have shown exactly why you’re right for the role. During the interview, you spent 4 minutes explaining something barely relevant.
Now you’re wondering whether a follow-up email can somehow fit an entirely new interview inside it.
Give yourself a practice round before you get to that point.
Little Pink Potato’s mock interview feature gives you a place to rehearse before the real conversation.
🥔 Here’s how to make that practice useful:
1. Pick the role you actually want.
Keep the job description beside you. Practise explaining how your experience connects to what they need.
2. Do the first round before you feel ready.
Answer without checking your notes every few seconds. Find out which parts you can explain and where you start rambling.
3. Rework the answer that gave you trouble.
If you talked about what “we” achieved, explain your own contribution. If you made a broad claim, bring a concrete example.
4. Try again without memorising every word.
You want to be able to explain your experience even when the question comes in a form you didn’t expect.
Have the awkward first attempt while it’s still practice.
Your best interview answer shouldn’t arrive after the interview.
The interview ends. You close your laptop.
And suddenly, you have the perfect answer.
You remember the project that would have shown exactly why you’re right for the role. During the interview, you spent 4 minutes explaining something barely relevant.
Now you’re wondering whether a follow-up email can somehow fit an entirely new interview inside it.
Give yourself a practice round before you get to that point.
Little Pink Potato’s mock interview feature gives you a place to rehearse before the real conversation.
🥔 Here’s how to make that practice useful:
1. Pick the role you actually want.
Keep the job description beside you. Practise explaining how your experience connects to what they need.
2. Do the first round before you feel ready.
Answer without checking your notes every few seconds. Find out which parts you can explain and where you start rambling.
3. Rework the answer that gave you trouble.
If you talked about what “we” achieved, explain your own contribution. If you made a broad claim, bring a concrete example.
4. Try again without memorising every word.
You want to be able to explain your experience even when the question comes in a form you didn’t expect.
Have the awkward first attempt while it’s still practice.
Remember when @RoundtableSpace asked if this is the future of job application?
From searching, applying to negotiating, all are handled by Little Pink Potato.
Agentese AI retweeted
We actually tested how Jev performs in a prediction workflow.
Yarrow used 48 real-world cases across 31 independent documents.
Total model cost: less than $0.10.
Yes, we have a cost advantage too. ☘️
Here’s what we found — what worked, what didn’t, and what remains inconclusive.
Where Jev did well:
Jev can read complex corporate announcements and turn them into structured, verifiable event records.
48/48 cases — it correctly identified the current status.
It also caught distinctions that simple keyword matching would miss.
For example, one announcement was titled “Withdrawal of Bank Regulatory Application”, while the body also said the company “reaffirmed its full-year guidance.”
A keyword-based system could mistakenly interpret this as a withdrawal of the guidance.
Jev correctly distinguished the two.
Where it failed:
Understanding the current status is not the same as determining whether a contract has already resolved.
In 10 of 48 cases, Jev correctly recognized that the available material did not mention the outcome — and then incorrectly concluded that the outcome was “No.”
For anyone trying to use it in production, this is a critical issue.
Overall judgment accuracy:Earnings guidance: 22/25
HBM supply: 16/23
Neither cleared our 90% threshold.
Stability is still an issue.
We ran the exact same inputs repeatedly. In 8 out of 48 cases, at least one field changed.
We also changed only the order of the input paragraphs. In 20 out of 48 cases, the result changed.
That means it is not yet suitable for automated resolution or fully unattended workflows.
Can Jev predict?
We’ve already frozen 12 forward-looking questions — covering Fed meetings, weather dates, and geopolitical events — along with 6 corporate disclosure predictions.
The outcomes are still pending.
But there is already one signal worth investigating:
For companies with regular quarterly disclosure patterns, Jev assigned 93–99% probability to the statement “the disclosure has not been released.”
We’re watching this closely.
There’s another interesting finding:
When we asked about the same event in two different formats — binary Yes/No versus multiple-choice — the probabilities shifted by an average of 12–13 percentage points.
The probability depends on how you ask the question.
Our interim conclusion
Jev is genuinely useful for one thing:
Turning messy corporate announcements into structured, checkable records.
It’s fast.
It’s cheap.
And it reads surprisingly well.
But reading accurately is not the same as judging accurately.
Understanding is not prediction.
We’ll keep experimenting and keep testing.
And if you’re currently researching Jev — especially its applications in prediction — we’d be happy to compare notes and discuss the results together. 👀 ☘️
Everyone is talking about Jev this week.
From the perspective of a prediction system, here’s how we see it.
Jev is not another chatbot. It’s a “System 1” model — essentially a classification engine that returns calibrated probabilities rather than generating text.
We noticed something interesting: Jev achieves accuracy comparable to GPT-5.6 Terra, while being 25× faster and 75% cheaper. Zero hallucinations — not because the model is smarter, but because its output space is predefined. It physically cannot generate content it wasn’t asked to produce.
We’ve also seen some impressive demos recently: searching for flights in 7 seconds for $0.004, running 50 simultaneous games of Subway Surfers for less than a cent, and rebuilding a Tesla FSD prototype in an hour.
Jev’s core training method is called RLCD — reinforcement learning for calibrated decision-making. Its optimization target is simple: when the model says 70%, the event should actually happen roughly 70% of the time.
That’s calibration — and it’s exactly the same core objective Yarrow is built around. ☘️
The difference is scope.
🤖 Jev answers: “Which button should I click right now?” — in milliseconds.
☘️ Yarrow answers: “If this policy passes, how will the corresponding market price change?”
Jev is System 1 — fast, intuitive, cheap.
Yarrow is System 2 — slower, more deliberate, evidence-dense.
Here’s the interesting possibility:
What happens if we combine Jev and Yarrow? 👀
Use Jev for thousands of micro-decisions throughout the prediction pipeline — scoring evidence relevance, classifying sources, verifying claims — cutting costs by 100×, while keeping the full reasoning engine for the final probability judgment and scenario distribution.
Small decisions get fast judgment.
Big decisions get deep simulation.
Jev and Yarrow are different — but as a technology stack, they could be highly complementary.
Stay tuned for Yarrow’s upcoming research. 👀 ☘️
Agentese AI retweeted
☎️ New Agent live on Finch: LifeClaw, made by @Agentese_AI
Tell it which local business you need and what you want done — it picks up the phone for you: booking a table, cancelling, rescheduling, or just asking what time they close.
📞 Calls on your behalf — the AI speaks the local language on the other end while you brief it in your own; works across most countries and regions (mainland China / +86 not supported yet)
🔍 Finds the business for you — restaurants, salons, clinics, dentists, hotels — and returns ratings, address, opening hours and phone number
💬 Fair billing — free if the call never connects, no extra charge for the back-and-forth to nail down details, 2 minutes max per call
🔒 Transparent by default — it says it's an AI at the start of every call, never impersonates you, and won't handle on-the-spot payments or identity checks
No more dreading a phone call in a language you don't speak. Free to acquire.
finchtech.ai/market/chips/ag…
The core insight here is calibration.
When a model says "70%," the event should happen 70% of the time.
Most AI systems cannot do this. They generate confident text, but the confidence is not calibrated to reality.
Jev achieves this for fast classification. Yarrow achieves this for slow forecasting.
This is the difference between "AI that sounds smart" and "AI that you can trust."
Everyone is talking about Jev this week.
From the perspective of a prediction system, here’s how we see it.
Jev is not another chatbot. It’s a “System 1” model — essentially a classification engine that returns calibrated probabilities rather than generating text.
We noticed something interesting: Jev achieves accuracy comparable to GPT-5.6 Terra, while being 25× faster and 75% cheaper. Zero hallucinations — not because the model is smarter, but because its output space is predefined. It physically cannot generate content it wasn’t asked to produce.
We’ve also seen some impressive demos recently: searching for flights in 7 seconds for $0.004, running 50 simultaneous games of Subway Surfers for less than a cent, and rebuilding a Tesla FSD prototype in an hour.
Jev’s core training method is called RLCD — reinforcement learning for calibrated decision-making. Its optimization target is simple: when the model says 70%, the event should actually happen roughly 70% of the time.
That’s calibration — and it’s exactly the same core objective Yarrow is built around. ☘️
The difference is scope.
🤖 Jev answers: “Which button should I click right now?” — in milliseconds.
☘️ Yarrow answers: “If this policy passes, how will the corresponding market price change?”
Jev is System 1 — fast, intuitive, cheap.
Yarrow is System 2 — slower, more deliberate, evidence-dense.
Here’s the interesting possibility:
What happens if we combine Jev and Yarrow? 👀
Use Jev for thousands of micro-decisions throughout the prediction pipeline — scoring evidence relevance, classifying sources, verifying claims — cutting costs by 100×, while keeping the full reasoning engine for the final probability judgment and scenario distribution.
Small decisions get fast judgment.
Big decisions get deep simulation.
Jev and Yarrow are different — but as a technology stack, they could be highly complementary.
Stay tuned for Yarrow’s upcoming research. 👀 ☘️
Agentese AI retweeted
We said AI demos should show the failed attempt. Yarrow takes that transparency a step further: showing what caused the failure.
In its test, one fabricated claim in the evidence pack worsened forecasting results across all eight runs.
When an agent gets something wrong, the model is only part of what we need to inspect. What information did it receive? Why was that information trusted?
A strong model still needs good evidence it can rely on. nitter.cf/Yarrow_ai/status/21008…
The biggest threat to AI forecasting isn't that the model isn't smart enough.
It's that the evidence isn't clean enough.
We tested this ourselves.
We planted one piece of information into the evidence pack that looked completely real—but was entirely fabricated.
We ran it 8 times.All 8 times, the Brier score deteriorated from 0.14 to 0.52—almost no better than a coin flip.
One false statement.
Total contamination.
Use a smarter model? It doesn't help.
Give it a larger context window? Still doesn't help.
Because we deliberately made the false information sound plausible, every model treated it as true.
The solution isn't intelligence.
It's clean evidence.
☘️ In Yarrow's evidence packs, every data point must cite an official source or be explicitly labeled as an estimate.
Every conditional conclusion must be calculated under the condition that its premise actually holds.
Before any forecast is published, the evidence pack goes through five layers of audit—including a red team whose only job is to find what everyone else missed.
Clean evidence in.
Good forecasts out.
Garbage in, garbage out.
What model sits in the middle? --- It doesn't matter.
An AI demo should show the second attempt.
The first attempt tells me the agent can complete a task when everything follows the expected path.
I want to see what happens when the booking requires clarification, the policy blocks the obvious option, or one step fails halfway through.
Does the agent recover? Does it ask a useful question? Or does the task quietly become my problem again?
A recent benchmark called Thinkingbox tested 17 AI models on 507 stateful business workflows, including bookings, hospitality, retail, and customer support. Each task was attempted 20 times.
One model eventually succeeded on 89.35% of tasks when given up to 20 attempts. But it completed the task correctly every single time for only 7.5%.
That difference matters.
An agent that can finish your booking once makes a good demo. An agent that can handle the next booking, with different constraints and complications, makes a useful product.
If I have to watch every step because I’m afraid it will fail, I’ve given myself a supervisory job.
Our view at Agentese: show the failed attempt. Show the recovery. Show when the user had to intervene. Then show what actually got finished.
One successful run proves capability.
Consistency earns delegation.
We said AI demos should show the failed attempt. Yarrow takes that transparency a step further: showing what caused the failure.
In its test, one fabricated claim in the evidence pack worsened forecasting results across all eight runs.
When an agent gets something wrong, the model is only part of what we need to inspect. What information did it receive? Why was that information trusted?
A strong model still needs good evidence it can rely on. nitter.cf/Yarrow_ai/status/21008…
The biggest threat to AI forecasting isn't that the model isn't smart enough.
It's that the evidence isn't clean enough.
We tested this ourselves.
We planted one piece of information into the evidence pack that looked completely real—but was entirely fabricated.
We ran it 8 times.All 8 times, the Brier score deteriorated from 0.14 to 0.52—almost no better than a coin flip.
One false statement.
Total contamination.
Use a smarter model? It doesn't help.
Give it a larger context window? Still doesn't help.
Because we deliberately made the false information sound plausible, every model treated it as true.
The solution isn't intelligence.
It's clean evidence.
☘️ In Yarrow's evidence packs, every data point must cite an official source or be explicitly labeled as an estimate.
Every conditional conclusion must be calculated under the condition that its premise actually holds.
Before any forecast is published, the evidence pack goes through five layers of audit—including a red team whose only job is to find what everyone else missed.
Clean evidence in.
Good forecasts out.
Garbage in, garbage out.
What model sits in the middle? --- It doesn't matter.