@humanlayer_dev

Multiplayer coding agent IDE for teams solving hard problems in complex codebases ai that works pod @ https://nitter.cf/t.co/1SEN2TwqzK

San Francisco, CA
Joined August 2024
👀
superthread of every resource I have published about advanced claude code usage, research-plan-implement, and advanced context engineering for coding agents 👇 (mostly chronological)
3
27
30,697
humanlayer retweeted
Getting hands-on with HumanLayer. @humanlayer_dev From what I’ve seen so far, I’d describe it as a software factory, it can plan, research, and drive implementation. Excited to build an agent harness using HumanLayer and explore what’s possible. Would love to hear your thoughts on this!
1
3
179
humanlayer retweeted
The almighty @saasmakermac shared his dev process with me yesterday. Here's some of mine. Messy, constantly changing, but fun! Featuring @AmpCode @humanlayer_dev @coderabbitai and AL 🫡 Also probably outdated by the time this video is encoded.
5
4
1
47
20,617
humanlayer retweeted
cool product out of riffing w/ @mattpocockuk this week - decided to open source the pull request skill that bundles with @humanlayer_dev - /show-me bundled with some steering to cut out a lot of the slop and noise that comes with most agent prs It's a small piece of a much larger puzzle, but rather than just share the SKILL.md contents with Matt, I figured we'd just give it to all of you 🙂 enjoy npx skills add humanlayer/skills --skill visual-pr
46
85
8
1,262
87,582
humanlayer retweeted
coming soon to a humanlayer new you: live multiplayer prompting - co-author a single prompt with your team. A few weeks ago we shipped "send a prompt to your colleague's session" - now we made each individual prompt collaborative. No cloud agents required, works wherever you run your agent - across laptops, mac minis, cloud compute, we even have this working in coding agent sessions that run in disposable compute like github actions. built on durable streams and @rivet_dev actors on ECS - incredible work from @0xblacklight That's in addition to all the other things we made collaborative in the last 2 months: - live google-docs style commenting on plans, right in the IDE, or plan/collab from your phone - diffs stream live as the coding agent is working, so you can comment on changes wayyy before PR shift collaboration left, shift alignment left, go check it out @humanlayer_dev - free for small teams go check it out today
27
9
1
127
17,227
Replying to @dexhorthy
@dexhorthy has fresh evidence that unattended coding agents still turn healthy codebases into radioactive spaghetti. At @humanlayer_dev, a lightly supervised software factory metastasized into 30,000 to 40,000 lines with an "insanely complicated state machine" running four Unix processes. They archived the whole thing. His new SlopCodeBench data shows every model accumulated defects as challenges piled up. That is the pattern. He is presenting "There Is No Software Factory Without Better Verifiers" at AGNTCon + MCPCon NA in San Jose on Oct 22. bit.ly/46ifPK1
3
2
31
14,432
humanlayer retweeted
One of our most useful skills @autumnpricing today is called /explain (based on /show-me from @humanlayer_dev) It started off just for code explanations, but I've found myself using it in every interaction today. Wanted to share some thoughts on a couple of the form factors which were the most intuitive! 🧵
2
1
14
1,174
I think @humanlayer_dev has been telling everyone this for close to 9 months now 1. Many folks CORRECTLY understood that LLMs can operate much more effectively in complex (and even bad) codebases than humans can 2. They therefore MISTAKENLY inferred that code quality & complexity and program design no longer matter 3. What they missed was that LLMs can also create much more bad code and much more complex systems faster than humans can. 4. The derivative of (3) is MUCH LARGER than the derivative of (1). Models do get better at working in bad codebases with each new generation. But they get a lot better at running unsupervised for a long time and producing exceptionally complicated programs. 5. The debt associated with (4) compounds superlinearly on large teams
All models, no matter how smart, will eventually build systems that they can no longer understand or maintain, if you let them. Fable 5 finally outbuilt itself, and flailed on me for a week. Fable 5.1 looks like it will fix it. For now. But you have to keep an iron grip on system size, or it'll run away from you.
15
33
2
358
40,628
humanlayer retweeted
at @humanlayer_dev we call this the alley-oop pull request
7
2
48
8,528
humanlayer retweeted
one of the biggest Agentic Engineering videos on youtube instead of watching Netflix tonight, watch this 58 minute masterclass you will be ahead 99% of people
damn already 200k views in 3 weeks sharing how we build with agents at @humanlayer_dev
6
19
369
31,717
.@grok in @humanlayer_dev has been achieved internally
2
1
7
1,574
humanlayer retweeted
wake up babe a dex video just dropped
1
3
876
humanlayer retweeted
Take 2 at comparing a plan produced by @humanlayer_dev and one without, on a nontrivial (3 repos) change this time. Humanlayer wins, big time. Not only did it plan for a nicer serialization, I also understood the plan and the high-level ideas/decision points much faster.
1M+ tokens, 4 huge planning md files and few hours later, comparing code generated from a @humanlayer_dev QRSPI session with a one-shot impl from a terse linear ticket... "the plan's hazard audit is wrong at 2/3 sites..." Not a great start.. what am I doing wrong @dexhorthy ?
1
1
1
657
lots of coding agent memory systems are unhelpful - at best i have said this a lot and some folks think I'm against agent memory in entirety which is not true - I just think that users should be on the loop that generates memories and should be able to enable / disable the system or individual memories at will and should be able to dismiss disruptive or unhelpful memories putting a user on the loop also lets you build a better data flywheel over time to improve the quality of recommendations live in @humanlayer_dev btw lmk if you want beta access
5
2
1
21
4,190
humanlayer retweeted
saturday internal tool tinkering is meditative
2
1
13
1,019
humanlayer retweeted
Really enjoyed this conversation with @dexhorthy on agentic coding. He and the team at @humanlayer_dev are pushing the frontier here and understand what actually works, what breaks, and where the hard problems still are.
On episode 12 of High Leverage, Joe Ruscio (@josephruscio) and Dexter Horthy (@dexhorthy) of @humanlayer_dev explore the promise and limitations of autonomous coding. Their conversation covers dark software factories, code review bottlenecks, program design, technical debt, and the continuing importance of experienced engineers. Tune in! hubs.ly/Q04rzHBv0
2
1
5
1,052
humanlayer retweeted
Replying to @dexhorthy
@dexhorthy @humanlayer_dev so stoked to be able to learn from your work substack.com/@dvinubius/note…
2
3
487
humanlayer retweeted
On episode 12 of High Leverage, Joe Ruscio (@josephruscio) and Dexter Horthy (@dexhorthy) of @humanlayer_dev explore the promise and limitations of autonomous coding. Their conversation covers dark software factories, code review bottlenecks, program design, technical debt, and the continuing importance of experienced engineers. Tune in! hubs.ly/Q04rzHBv0
2
2
5
6,753
humanlayer retweeted
The last AI that works unconference brought together 70+ founders, systems engineers, and AI tinkerers to share the latest learnings and push the frontier together at @ycombinator's dogpatch HQ. This weekend me and @vaibcode are running it back (signup below!) with a whole bunch of great engineers and plenty of room for audience-driven content - 8 hours of talks, breakout sessions, and some legendary Keynote Speakers - @kwindla - @trydaily - @swyx - @aiDotEngineer - @lenadroid - @Akamai - @bdougieYO - @papercompute - @JohnKutay - @Rippling - Leif from @SierraPlatform - @codeshaunted - @boundaryML - @0xblacklight - @humanlayer_dev
10
8
5
54
15,231
. @dexhorthy opens "Harness Engineering is not Enough: Why Software Factories Fail" with a bold claim about agentic coding: no amount of harness engineering or loops can fix what is fundamentally a model training issue. It's on @aiDotEngineer's YouTube. Dex is Co-Founder at HumanLayer. The talk gives you a mechanical explanation for why coding agents erode your codebase over time, traced back to how the models are trained, plus the workflow HumanLayer uses instead. - They tried lights off and it broke. In July 2025 HumanLayer stopped reading the code. Months later they hit a bug the agent couldn't solve, and had to dig into a codebase nobody had read since. - Brownfield starts sooner than you think. Dex argues agents begin to struggle on a codebase after roughly three to six months, not after ten years. - The training loop can't see maintainability. SWE-bench-style rewards are binary: did the test pass, did you break anything else. Nothing in that scoring penalizes a needless try/catch or a cast that exists only to make a test go green. - The cost function of bad architecture is measured in months and years. That gap is why the reward signal can't propagate back to the coding episode that caused it. - Cloud Code's edge was training against its own harness. Same tools as ADER and CodeBuff. The difference was a model lab RL'ing the model inside the harness it shipped in. - A judge model has a ceiling. If the model knew what good code looked like, it would have written it the first time. - Newer benchmarks are reaching for this. SWE Marathon from Abundant AI, Deep SWE from Data Curve, and frontier code from Cognition, which penalizes tests that don't fail on the pre-patch code. - Plan up front so review stays cheap. Product review, system architecture, then program design (types, method signatures, call stacks), then vertical slices for order of implementation. - Program design is the underemphasized step. People assume the model can cook once architecture is right. Dex says the layout and call stacks are where the work is. - You don't have too many PRs, you have too many bad ones. A good PR is a joy to review. Thirty minutes of alignment up front saves hours of review, which makes reading every line still feasible. I'm working through the published talks from AI Engineer World's Fair sharing summaries and takeaways. Follow for more!
3
1
4
489
humanlayer retweeted
Giving @humanlayer_dev a first try and holy shit, the product workflow hits different
1
1
2
240