@RoutekitShell

AI that follows your rules. Governed agent workflows • Stories • RAG • MCP • Audit trails • AGPL

Joined July 2026
AI models are remarkably capable. Capability isn’t the problem. Reliability is. Routekit is the system around the model that turns probabilistic AI capability into predictable software delivery. Context and provenance. Durable state. Bounded tools and permissions. Explicit workflows. Specialized agents where reasoning is useful. Deterministic enforcement where it isn’t. Validation, recovery and observable execution. Your models. Your keys. Your files. Your infrastructure. AGPL. Open source. No platform lock-in. The models will change. The system that makes them reliable is yours. Routekit — reliable AI work, by design. routekit.dev
174
A well-designed governed system shouldn’t constantly announce that it is governing you. In normal operation, the safe path should also be the easiest path. Authority gets established as part of the workflow, provenance gets captured because retrieval naturally produces it, validation happens because advancing state invokes it, and permissions are scoped because the work itself establishes the scope.
25
Product management is full of work we treat as inherently conversational that is actually process-defined. Understand the customer. Identify the problem. Establish the outcome that matters. Research the existing system. Evaluate constraints and opportunities. Define scope. Surface assumptions and dependencies. Establish acceptance criteria. Decide whether the work is worth doing and whether it’s ready. There’s judgment throughout that process. But there’s also a process. I’m much less interested in having AI write user stories faster than I am in making sure we understand why something should be built, what outcome it should create, and what must be true before we build it. The opportunity isn’t automating product management. It’s making good product practice executable.
11
Experience still matters. AI can write better code than I can. I’m a designer who somehow ended up building a software engineering application, so that isn’t particularly surprising. What surprised me is how quickly writing the code stopped being the hard part. What should we build? How should the work be decomposed? What assumptions are we making? What does “good” look like? How will we know when it’s wrong? Those are judgment problems. AI can produce remarkable work. But give it poor context, weak structure and no meaningful definition of quality and it can produce the wrong thing remarkably well. Experience is knowing what should be built, where it’s likely to fail, and how to recognize good from merely plausible. Better models don’t make that experience obsolete. They make encoding it into the systems around them more important. Says the designer who vibe coded a governed agent system.
1
42
Guidance isn’t enforcement. I got a very practical reminder of this while debugging Routekit this week. The Dispatcher (Claude Code in this case) had suddenly become incredibly verbose. Walls of commentary instead of getting into the work. I found a surprisingly effective workaround: Every prompt ended with: "BE CONCISE!!!" And it mostly worked. But I had to remember to do it. Every. Single. Time. Worse, I needed the Dispatcher to help me debug the system that was supposed to put Routekit into the right operating context in the first place. So, "BE CONCISE!!!" became the temporary patch that made Routekit usable _enough_ to help fix its own harness. Eventually the bugs were squashed and now the agent gets routed into the appropriate Skill almost immediately, with the context it needs to do the job. No more yelling "BE CONCISE!!!" at every prompt. The instruction wasn’t useless. It was actually useful diagnostic evidence: Routekit could behave the way I wanted under the right conditions. But the fix wasn’t a better prompt. The fix was engineering the system that establishes those conditions. Prompts are great for guidance. If a behavior actually matters, enforce it in the system.
84
Routekit Shell retweeted
What a fascinating reversal. We are now using code to constrain reasoning agents. We used to be reasoning agents that constrained code.
32
29
5
536
29,654
DHH spent 20 years dismissing product management. Then 1h21 into his Pragmatic Engineer interview, he caught himself and admitted he was wrong. He was listing what matters now that AI writes the code: figuring out what to build, how to build it, which customers to talk to, where to focus. Then his exact next words: "It's product management. It's so funny for me too because historically I've not necessarily had the highest esteem for product management as a function. I thought there was a lot of BS." He explained why. Implementation was always the constraint. Engineers needed four weeks to ship anything, so PMs spent those weeks talking, planning, strategizing. Nothing looked like output until the code landed. "They were underutilized. They were not the constraint." They were rate-limited. The loudest PM-skeptic changing his mind (and DHH has strong opinions!) -- that should count for something. let's toast to that 🥂
26
67
15
880
97,695
The model isn’t the system. We spend an enormous amount of time talking about which model is smartest. Claude vs. GPT vs. Gemini. Benchmarks. Context windows. Reasoning scores. How long an agent can run. The model matters. But once you’re trying to make AI do real work predictably, it’s only one component of the system. - What context does it receive? - What state persists? - What tools can it use? - What is it allowed to change? - What happens next, and who decides? - How is its work validated? - What happens when something fails? - When does a human need to get involved? Those aren’t prompting problems. They’re system design problems. A better model can improve the reasoning inside the system. It doesn’t automatically give the system context, state, permissions, provenance, workflow, validation or recovery. And as models improve, the surrounding architecture becomes more important, not less. Use the cheapest model capable of the reasoning you need. Spend expensive intelligence where it actually changes the outcome. The model provides capability. The harness turns that capability into a system you can depend on.
97
We’re recreating the Second Brain productivity trap with AI memory. Capture everything. Summarize everything. Put it in Obsidian. Add embeddings. Retrieve two relevant notes and declare that your AI now “knows you.” We did this before. We reorganized our notes, redesigned our PARA folders, added backlinks and spent Sunday afternoon watching videos about how to improve the system that was supposed to save us time. The outline looked fantastic, so the experiment was a success. Retrieval isn’t the interesting test of AI memory. The interesting test is whether retained knowledge improves future work when you don’t already know what should be retrieved. Did it preserve the right decision? Did it know what was superseded? Can it distinguish evidence from speculation? Did it surface the constraint you forgot existed? Did the next worker produce a better outcome because of it? Otherwise we’ve taken the Second Brain, given it a vector database, and started calling it intelligence.
1
1
146
It’s pretty clear that we collectively have a box of apples and oranges when it comes to our respective definitions of “agent harness.” But it's more useful as a description of the larger system responsible for turning model capability into reliable work: context, durable state, knowledge, tools, permissions, workflow, orchestration, validation, recovery and observability. In that system, Claude Code, Codex, etc are execution runtimes. The model is another component. Neither is the whole harness. This distinction matters because commentary like “the harness is getting in the way” is becoming more common, usually to describe bloated context or an overly opinionated coding agent. That’s not an argument against harnesses. That’s a harness design problem. A good harness shouldn’t maximize the machinery around the model. It should give each piece of work the minimum structure required to produce a reliable outcome. Maybe we need better names for these layers before “harness” becomes the next word that means everything and therefore nothing.
2
44
Capability isn’t reliability.

 AI can write production-quality code.

It can produce excellent research, analyze a contract, diagnose a bug, create a marketing plan, or make a genuinely insightful product recommendation.
That proves capability.

 Now do it again.

 Give it slightly different context. A larger codebase. Conflicting information. A missing requirement. An unusual edge case. Let it run longer. Give it tools that can actually change things.

 Does it still produce the same quality of outcome? That’s reliability.

 We spend a lot of time benchmarking what models *can* do. I think the harder engineering problem is turning that capability into something predictable.
Good context. Durable state. Bounded permissions. Explicit workflows. Validation. Recovery. Knowing when to involve a human.

 The goal isn’t to eliminate uncertainty. These are probabilistic systems.

 The goal is to build systems where uncertainty doesn’t quietly become failure.

 A great result is impressive.

 A great result you can depend on is useful software.
26
AI isn’t your teammate. It’s software.
 And anthropomorphizing AI is pushing us toward some bad product and engineering decisions.
 We talk about agents as coworkers. We give them names and personalities. We build interfaces around conversations. We celebrate how long they can work autonomously.
 But software has always had a simpler job: help us accomplish something while demanding as little of our attention as possible.
 An AI agent that needs constant prompting, clarification, supervision and correction isn’t a particularly good teammate.
 More importantly, it isn’t particularly good software.
 The opportunity with AI isn’t to give every person a team of synthetic coworkers they now have to manage.
It’s to build software that understands enough context, maintains enough state, has the right tools, operates within clear constraints, and can validate enough of its own work that we can safely give it something increasingly valuable:
 Our attention.
 The best AI systems may ultimately be the ones we interact with the least.
1
50
X may be the town square, but we all know there’s plenty of snake oil being sold in the bazaar. I’m wildly optimistic about AI. I’m also increasingly skeptical of some of the ways we talk about building with it. So starting next week, I’m going to spend some time exploring a few ideas I keep coming back to: • AI isn’t your teammate. It’s software. • Capability isn’t reliability. Doing something well once is very different from doing it well predictably. • The model isn’t the system. Context, state, tools, permissions, workflow, validation and recovery matter. • Guidance isn’t enforcement. If a constraint matters, build it into the system instead of asking the model nicely. • Autonomy isn’t the goal. Reliable delivery of validated outcomes is. • Experience still matters. AI can generate an extraordinary amount of work. Judgment is knowing whether it’s the right work. • Make the deterministic parts deterministic. Use models where reasoning and judgment are actually required. These aren’t theories from watching demos. They’re ideas I’ve developed while building with this stuff, getting it wrong, fixing it, and trying to make increasingly autonomous AI work predictably. More next week.
1
56
Most AI agent deployments fail because companies confuse 'model intelligence' with 'execution reliability.' If your agentic workflow relies on the LLM to self-police its own file edits, permissions, and code hygiene, you aren't automating—you're just outsourcing your technical debt creation. Real agentic leverage requires deterministic CLI guardrails around the model.
51
AI applications shouldn’t depend on one model. They shouldn’t depend on one cloud. They shouldn’t depend on one vendor. They shouldn’t depend on one browser. They shouldn’t depend on one knowledge store. Routekit.dev is the execution layer that keeps all of those replaceable. #GovernedAI #AIEngineering
37
Everyone talks about context windows. Almost nobody talks about context discipline. Those are different problems. #GovernedAI #AIEngineering
21
AI doesn’t need more freedom. It needs better guardrails. RouteKit Shell (rks) turns: “Build me…” into - scoped work - cited reasoning - reviewable plans - governed execution - automatic Git workflows - visible token costs AGPL. Built in public. routekit.dev
55