@backnotpropi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
Cofounder, AI @EQTYLab / prev dc - complex systems / veteran / For fun: @plannotator
CA
Joined March 2024
- Tweets1.8K
- Following918
- Followers1.8K
- Likes4.4K
Pinned Tweet
Anthropic code review this, clanker review that ... why don't you shut up and review+annotate your own code.... (yes im a loser who still manually reviews code)
Originally inspired by a bunch of feature requests and then seeing @dillon_mulroy tweet a similar cool ux.
@plannotator for reviewing plans (primary focus) and code, fully oss.
OpenCode, @badlogicgames 's pi.dev, and Claude Code and other clankers
I don’t think we are talking/arguing about this enough yet. I think we will and I don’t think this is the same type of industrial/economic advantage history offered to companies that spend more. It feels much worse. Could be wrong.
nitter.cf/backnotprop/status/210…
Sharing some prompt injection "jailbreak" benchmarks I've been running with Jev, competing against a bunch of standard guardrail classifiers including Meta's.
Part of it is a standard open suite that's getting dated. Jev wins that one outright, and wins the newest attack set too. It loses on the older sets and can't run at a tight false-alarm budget.
It failed some of our internal benchmarks when something more nefarious is going on & is obscured, but there's really not any good off the shelf classifier that could be used for those type of scenarios. You can see a glimpse of this with how every model performed poorly at allenai's WildJailbreak.
Since the poker eval I've been trying to pin down where Jev fits. My hunch is perfect-information calls with little strategy in them. Still testing it.
We are in the age of orchestration. It is very effective to work through 1 agent per project, while it delegates tasks to durable subagents. Hence Cursor/Claude Code Projects (not that they offer anything you cant do today without them).
The approach doesn't require alchemy or some stupid metaphor philosophy approach. Just a proper subagents implementation, and integrated-monitoring. Some popular harnesses fuck these things up.
Here's 2 extensions to get it right in Pi:
1. subagents
github.com/nicobailon/pi-sub…
2. Monitors/polling/loops
github.com/joelhooks/pi-unti…
Michael Ramos retweeted
Replying to @backnotprop
in pi i use this constantly for loops github.com/joelhooks/pi-unti…
Codex needs actual monitors. Scheduled reminders as implemented eats a whole turn of tokens every time. It is not viable for any type of polling situation.
Instead I need to set a programmatic monitor that only alerts the agent when needed.
And we all need MCP events/triggers/tasks/whatever asap
Michael Ramos retweeted
The latest Plannotator is out
@MermaidChart's 12.0 release is awesome so we've shipped it plus theming sync and new interaction capabilities.
If you are ab testing engineer workflows using your product - you do not care about engineers.
Replying to @Steve_Yegge
Yep you're part of the experiment. Yesterday I found out Anthropic was quietly running A/B tests with my system prompts in Claude Code without my knowledge 😢 lnkd.in/p/gahaK-Uh
I had edited down a meaner stance on this awhile back
backnotprop.com/blog/do-not-…
Jev is a 🐠 / not good at poker.
Up to you to decide what intelligent and consequential decisions you're willing to have this thing make for you.
writeup: nitter.cf/backnotprop/status/210…
Here is my writeup on jev & making decisions in a game of poker
backnotprop.com/blog/jev-pok…
Here is my writeup on jev & making decisions in a game of poker
backnotprop.com/blog/jev-pok…
This is the future in that we really only need to interface with one agent. Age of orchestration turns into us talking to a single assistant. Easier on us cognitively.
But you could already do this very well in CC/subagents and I hope the core does not degrade. I just need a simple agent interface and high class sub agents implementation.
Poker eval was iterated on quite a bit. At one point I Showed the opponent's hand, yet Jev still made a poor decision.
Jev would respond correctly when I gave it very clear the state of the world, like when it was in a non-favorable position like "Opponent has nut flush. Hero has 0 outs"
And you could say well you need to tell the model versus make it derive from seeing the opponent's hand, but in my mind this is all a wash - because half the showcases on our X are about, you know, can this model make intelligent decisions?
So I don't think I'd use it to classify insecure code or make decisions about how to navigate across a browser for any consequential task.
I do have plenty of fun use cases in mind though.
Michael Ramos retweeted
Today we introduce TIN: a powerful and reliable full-text search extension for Postgres.
TIN works with complicated WHERE clauses, replication, backups, and maintains correct transaction visibility.
It's also mind-blowingly fast.
Michael Ramos retweeted
Replying to @backnotprop
Thanks for mentioning GLiCLass. We initially designed it for information extraction workflows. But actually, the long-term strategy was to develop systems like Jev, we also introduced RL-based training for such models in our paper: arxiv.org/pdf/2508.07662
Also, please check our newer, more generalist model: huggingface.co/knowledgator/…
I deleted a tweet that was starting to go viral. I misspoke on the benchmark comparison and it would have been wrong to spread that.
Still worth knowing about GLiClass...
github.com/knowledgator/glic… huggingface.co/models?search…
There are open models for zero-shot performant classification. Maybe useful as a foundation for Jev-like systems.
What Jev has done is impressive in terms of the RL and global calibration.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution