@backnotprop

Cofounder, AI @EQTYLab / prev dc - complex systems / veteran / For fun: @plannotator

CA
Joined March 2024
Anthropic code review this, clanker review that ... why don't you shut up and review+annotate your own code.... (yes im a loser who still manually reviews code) Originally inspired by a bunch of feature requests and then seeing @dillon_mulroy tweet a similar cool ux. @plannotator for reviewing plans (primary focus) and code, fully oss. OpenCode, @badlogicgames 's pi.dev, and Claude Code and other clankers
18
17
10
317
73,905
I don’t think we are talking/arguing about this enough yet. I think we will and I don’t think this is the same type of industrial/economic advantage history offered to companies that spend more. It feels much worse. Could be wrong. nitter.cf/backnotprop/status/210…
for arguments sake let's accept that spending a lot of tokens is truly more productive we're in a weird stage where early stage startups are disadvantaged. they cannot afford to do this it's reverting back to the days when you had to raise money to buy servers to even launch
2
234
Michael Ramos retweeted
show me your custom cloudflare dashboards for birthday week ❤️‍🔥
7
11
3
96
7,806
Sharing some prompt injection "jailbreak" benchmarks I've been running with Jev, competing against a bunch of standard guardrail classifiers including Meta's. Part of it is a standard open suite that's getting dated. Jev wins that one outright, and wins the newest attack set too. It loses on the older sets and can't run at a tight false-alarm budget. It failed some of our internal benchmarks when something more nefarious is going on & is obscured, but there's really not any good off the shelf classifier that could be used for those type of scenarios. You can see a glimpse of this with how every model performed poorly at allenai's WildJailbreak. Since the poker eval I've been trying to pin down where Jev fits. My hunch is perfect-information calls with little strategy in them. Still testing it.
9
1
26
1,308
We are in the age of orchestration. It is very effective to work through 1 agent per project, while it delegates tasks to durable subagents. Hence Cursor/Claude Code Projects (not that they offer anything you cant do today without them). The approach doesn't require alchemy or some stupid metaphor philosophy approach. Just a proper subagents implementation, and integrated-monitoring. Some popular harnesses fuck these things up. Here's 2 extensions to get it right in Pi: 1. subagents github.com/nicobailon/pi-sub… 2. Monitors/polling/loops github.com/joelhooks/pi-unti…
22
6
169
13,907
Codex needs actual monitors. Scheduled reminders as implemented eats a whole turn of tokens every time. It is not viable for any type of polling situation. Instead I need to set a programmatic monitor that only alerts the agent when needed. And we all need MCP events/triggers/tasks/whatever asap
10
39
2,466
Michael Ramos retweeted
The latest Plannotator is out @MermaidChart's 12.0 release is awesome so we've shipped it plus theming sync and new interaction capabilities.
4
7
1
129
5,929
Just spotted this frontier lab billboard in SF
3
2
14
880
This is the future in that we really only need to interface with one agent. Age of orchestration turns into us talking to a single assistant. Easier on us cognitively. But you could already do this very well in CC/subagents and I hope the core does not degrade. I just need a simple agent interface and high class sub agents implementation.
Projects now run from one conversation, starting in Claude Code. You describe what needs doing, and Claude directs parallel threads that keep working after you close your laptop. In beta today for select Pro and Max users in cloud sessions; coming to all Claude users soon.
2
22
2,126
Haiku 4.5 with reasoning turned off
1
1
383
Poker eval was iterated on quite a bit. At one point I Showed the opponent's hand, yet Jev still made a poor decision. Jev would respond correctly when I gave it very clear the state of the world, like when it was in a non-favorable position like "Opponent has nut flush. Hero has 0 outs" And you could say well you need to tell the model versus make it derive from seeing the opponent's hand, but in my mind this is all a wash - because half the showcases on our X are about, you know, can this model make intelligent decisions? So I don't think I'd use it to classify insecure code or make decisions about how to navigate across a browser for any consequential task. I do have plenty of fun use cases in mind though.
1
5
301
Michael Ramos retweeted
Today we introduce TIN: a powerful and reliable full-text search extension for Postgres. TIN works with complicated WHERE clauses, replication, backups, and maintains correct transaction visibility. It's also mind-blowingly fast.
58
114
31
1,978
501,475
Michael Ramos retweeted
Replying to @backnotprop
Thanks for mentioning GLiCLass. We initially designed it for information extraction workflows. But actually, the long-term strategy was to develop systems like Jev, we also introduced RL-based training for such models in our paper: arxiv.org/pdf/2508.07662 Also, please check our newer, more generalist model: huggingface.co/knowledgator/…
1
5
111
I deleted a tweet that was starting to go viral. I misspoke on the benchmark comparison and it would have been wrong to spread that. Still worth knowing about GLiClass... github.com/knowledgator/glic… huggingface.co/models?search… There are open models for zero-shot performant classification. Maybe useful as a foundation for Jev-like systems. What Jev has done is impressive in terms of the RL and global calibration.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
3
1
30
2,798