@ceefryingpan

building Luke (YC P26)

Joined February 2026
Charles Pan retweeted
SotA forward deployed engineering in 2026: "try this prompt with your guy: ..."
2
1
8
257
I am super against this and I don't understand the backlash against plan mode People still don't put enough effort into planning with their agent and then are surprised when it makes assumptions or does things that they're not expecting
we’re thinking of killing plan mode and using the shift+tab hotkey to adjust effort levels I don’t think the models need plan mode anymore, but if you’re a plan mode diehard would love to get your feedback on why
2
3
280
Just like how you'd sit down with a teammate and sketch out a plan/discuss trade offs when working on a large feature, we should be doing the same with agents
1
26
this is one of the best use cases for jev I've seen
a jev lint rule that detects if a new component is a potential duplicate of an existing file. this check takes less than a second + costs less than $0.01 (for our codebase at least) jev makes it feel actually reasonable to run these checks on every code change.
2
166
what are evals and benchmarks going to look like in the jev world?
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
1
170
I've noticed that forcing agents to keep its answers around 150-200 words dramatically improves the quality of its responses
2
2
5
186
this is a great use case
I tested Jev on the most annoying problem in personal finance tools: getting a good payee description from raw bank data It probably 95% of the way there on the first try, and there's lots of different ways I can improve how to ask it
1
2
131
Luke just became your new favorite way to ship PRs. Today we're releasing Luke v0.6, powered by @OpenAIDevs GPT-Live-1. It feels like you're talking to a real EM that can manage your dev work for you.
3
1
1
12
684
the fact that Rippling makes you pay to use their MCP is incredibly stupid
1
2
127
need a good system for splitting up big feature work into smaller PRs, each handled by its own agent has anyone built this yet?
1
1
5
98
we got agents congratulating each other before GTA 6
we got agents thanking each other before GTA 6
1
1
4
307
if AI solves P vs NP everything is over
SITUATION DETECTED: OpenAI is using the internal model that produced its Navier–Stokes proof to attempt the Riemann Hypothesis and P vs NP.
1
3
119
Luke is about to sound a whole lot better
GPT-Live-1 is now available in the API. Bring ChatGPT’s natural back-and-forth to your app, with voice agents that listen while they speak and work with the models and harness you choose.
1
3
10
625
personal agents are the new model routers
Introducing Muse, the personal agent that understands your goals and works 24/7 to get things done for you.
1
2
443
whoever figures out how to represent a codebase visually will make a lot of money mermaid diagrams are a good start but just aren't it
1
3
67
I remember the days when Claude Code was the undisputed best coding agent the Codex comeback story needs to be studied
1
1
7
343
Charles Pan retweeted
I don’t think there’s any benchmarks that accurately reflect the current capabilities of LLMs
264
116
35
3,367
133,286