For 4–5 months, HumanLayer ran a “YOLO pull request” experiment, reviewing the plan instead of the code. It ended with an unusable codebase and a six-service architecture that the model had over-designed without guidance. In the rebuild, co-founder Kyle hand-typed the data architecture for the first two weeks. Dex now treats planning as expected-pain management and still reads the code, because maintainability has no fast oracle and bad code degrades every future run of a software factory. He backs compounding factory improvements over big-bang rewrites, and, as a founder put it to him, there will probably always be alpha in reviewing something.
Dex Horthy is CEO and co-founder of HumanLayer, and the host introduces him as the author of “12 factor agents”. His journey started around the launch of Opus 4 and the Claude Agent SDK. That was the first time you could use Claude’s small agent loops inside a broader, more deterministic piece of software. He jokes that today the Claude CLI is basically the JVM of agent coding, with around 200 flags. HumanLayer was building terminal multiplexers for Claude Code, meaning tools to run many sessions in parallel with an inbox of approvals waiting on you. Much of that work was about getting agents to work for longer.
Dex says the biggest change between summer 2025 and now is “the planning metagame.” In July 2025, writing a plan file was the highest-leverage move because it reliably kept the model working longer. The conversation also raised Spec Kit as an example of tooling that tried to micromanage the agent step by step. A plan was a “compaction of intent” that shrank everything you wanted to build into a couple of hundred lines of markdown. You could keep it updated as work was done and resume later from a clean context window.
Dex coined “the dumb zone” for context-window saturation. At the time it began at around 100,000 tokens, after which results got much worse. The host adds that even with larger limits there is probably still a dumb zone: too much focus is no focus. Being deliberate about what goes into context still pays off.
HumanLayer had just come away from a talk by Sean Grove, then a researcher at OpenAI, arguing that the spec is the new code. His analogy: you compile Java into a jar and commit the source, not the jar. Having a long chat with an agent, throwing away the chat and shipping only the code is like checking in the compiled jar and throwing away the source. The prompts are the high-leverage artifact.
The labs seemed to have more alpha than anyone, and OpenAI said this approach was working for them. So HumanLayer tried “YOLO pull requests”: if you read the plan, the code matters less. There were problems from the start. The plans contained every line of code as inline diffs. Larger teams, some with around 50 developers, reviewed the plan and then also reviewed the code because the implementation drifted. They were “reviewing basically the same, almost the same shape of code twice.” HumanLayer ran the experiment for 4–5 months before realizing the codebase was unusable. It got slower and slower, and every change caused a regression somewhere else.
HumanLayer also let the model decide the system architecture, with a human in the loop reviewing its proposals. Left to decide, it produced a six-service system: a desktop app, a Go daemon, Claude Code sessions, MCP servers for approvals, and SQLite, all wired over Unix sockets. The desktop app’s socket link to the daemon went through Rust. Every approval travelled from the UI to the MCP server, to Claude, to the Go daemon, and back to the UI.
The host pushed back: is the lesson that models can’t architect, that they can’t architect yet, or that they can’t without enough guidance? Dex answered with what happened next. The team rebuilt the whole product from scratch. For the first two weeks, co-founder Kyle typed every character of the data architecture by hand in VS Code. The new version is built on ElectricSQL and durable streams. Data flows in one direction, and a sync engine handles all signaling. Dex says those early decisions “constantly cascade into the future” and have held up well.
Dex paraphrases a debugging maxim he attributes loosely to “Donald Knuth or one of these like OG like C language guys”: if you write the cleverest code you can, you are by definition not smart enough to debug it. He cites a recent case of a prominent advocate of not reading the code, named Steve. Steve reported that Fable built a system so complicated that the model itself could not debug it.
The host adds a counterargument. Tessl has about 20 engineers and compliance requires a named reviewer, but reviewing code by hand doesn’t build a mechanism that produces code to your preferences. He describes specs as a slider between determinism and adaptability.
Dex frames planning as a stage in a factory. Before AI, humans did sprint planning and architecture alignment to cut rework and review time, because building and reviewing each took hours or days. When teams say they are drowning in PRs, his answer is that they have too many bad PRs, because people stop thinking while they code. There is a spectrum of effort:
The deciding question is expected pain: how likely a change is, and how painful it would be. A button color takes one prompt to fix. A wrong database schema means throwing the work out. HumanLayer has several planning workflows for different feature sizes: a plain design doc, a design plus the order of work, and a PR flow with 3–4 stages. Each stage should be “slicing off 50% of the potential bad outcomes.”
On capturing intent rather than just decisions, Dex points to Matt Pocock’s “grill me” skill. He says it is among the top 25 GitHub repos by stars of all time, and in the top three for repos that mostly just contain skills. He takes that as a sign that people are bad at specifying intent. His own approach is to turn on voice mode, ramble for 60 seconds, and let the AI interview him about alternatives and open questions.
Asked whether good intent capture plus good verifiers would remove the need for review, Dex cites a Haskell-community post arguing that a sufficiently detailed spec is indistinguishable from code. He notes that the post is about implementation specs, where you are dictating an algorithm. He agrees that a spec with verification is powerful. Give the model “back pressure”, meaning a way to get feedback without a human, and “it will move mountains for you”. You can also run more work in parallel.
He is using this for a code-search harness experiment built around a model with only a 32K–64K input context window. He uses an evolutionary approach, hill-climbing the Pareto frontier of cost, speed and accuracy across 45 code-search evals, and he is not reading any of that code. If it works, he’ll have the model explain the algorithm, turn that into a spec, and re-implement it with clean architecture. Skipping code review is fine for prototyping and proving what’s possible.
In the earlier days, people said the code might be slop, but fixing it would be “GPT seven’s problem.” Dex replies: “folks, we’re at six.” Astra (as captioned) got much better at computer use, video game slop and Blender, but not much better at maintainable code. His explanation: training a skill with RL needs an oracle, and you don’t learn that code is unmaintainable until two months later. He discussed this with Addy Osmani, whose take was that maintainability has no fast oracle.
Slop Code Bench comes from a lab at the University of Wisconsin; HumanLayer helped with runs and inference. It gives a model a codebase and a series of features, and the model has to build each feature on top of its own previous code. The benchmark is far from saturated: GPT 5.5 scored 14.8% in May. The lab was still preparing to publish Astra’s results, and in the first run Astra scored 16.3%. Models are improving here, “but not as fast as everyone thinks.”
The host describes Tessl’s path from spec-driven development to context and skills. In his view, skills hold specs (how a product works), workflows (how to perform actions) and opinions (policies and choices). Dex gives an example from Sprout Social, where his boss Allen had a rule called “three copies pressed hard”. Don’t wrap the log-then-metric-then-action lines at the top of an endpoint in an abstraction. Duplicate the three lines instead, because abstractions leak and pick up responsibilities. That kind of taste is hard to put into a model.
Dex says CLAUDE.md or AGENTS.md files end up as long lists of bullet points that mostly don’t get followed, because there is too much context. HumanLayer has deterministic detectors for some patterns, such as react doctor and oxlint warnings, and simple non-deterministic ones for others. Every night, 4–5 “agents” run on crons and GitHub Actions. Each is the same harness with a different user message, and the team wakes up to five PRs that each improve the codebase along one dimension. The more opinions you have, the more clever your context engineering has to be.
Dex is clear that he is “very bullish on this compounding thing”. Later he adds, “I 100% agree on this idea of compounding engineering.” He objects to two things. The first is the org that throws everything out to build a software factory from scratch; big design up front always loses to getting 1% better every day or week. The second is letting slop in. A bad car leaves the line, but bad code “degrades every future piece of work that goes through that factory, because there’s more bad patterns for the model to see.” So the first priority is to stop the bleeding.
Human review has useful exhaust. Every PR comment that says “we don’t do it this way”, and every session trace where someone tells the model it got something wrong, should feed into better base skills and code-review skills. Done well, you shift “from like caring about your position to caring about your velocity.”
The host adds that the existing codebase is one of the main pieces of context guiding the factory, which is a counterpoint to treating tech debt as deflationary. Dex won’t predict when code stops being the source of truth. His Fortune 500 and large-company customers like HumanLayer’s approach of not mortgaging the codebase for the future.
At a San Francisco dinner about six weeks earlier, everyone raised a hand when asked if AI writes 100% of their code. About 50% raised a hand when asked if they still read the code. Others had, for example, hand-written 70 custom linter rules. Dex guessed one to two years until people stop reviewing code. Then a veteran founder at the dinner said there will probably always be alpha in reviewing something. The reasoning, as retold in the conversation: if you review nothing, you get “the same product, the same email, the same company that everybody else gets.” Dex: “For now, I think it’s still the code, but eventually it might be something else.”
On taste, Dex is skeptical of the idea that taste is the new moat: “I think that’s really blurry.” He thinks understanding the problem better than anyone else is what counts, and that “taste is just a preference that you haven’t bothered writing down.” Written-down taste is still a differentiator, but it’s “not this elusive human versus AI difference.”
The host’s view is that software engineering is a creative profession, and that skills capturing taste and preferences will become the units of development. He compares this to Tessl’s idea of “a spec and a shadow spec” and to delegating garbage collection in Java. Dex names two bottlenecks for builders. First, how fast the model can “brain drain” the human on their tastes and preferences. Second, the fastest and most visual way for the model to show what it has done, which he calls letting the model “context engineer the human”. The goal is something 50–80% faster than reading every line that gives almost as much confidence.
Dex expects context engineering to stay true as long as we use transformer-based LLMs, because you get better quality with less context. His prediction is that the software factory stack will decompose into good interfaces and open components, as happened in the Kubernetes world. He contrasts that with building everything vertically on Claude: Claude Code as harness, Claude automations as orchestrator, and managed agents and sandboxes as runtime. He expects dev tools to stay open or become more open. His advice is to prepare for that, and to be careful with the code, especially on critical systems. The host recalls Matt, the founder of Netlify, arguing on an earlier episode that the open web won, and thinks open and composable will win here too. The conversation closes on open “catching up faster than anyone is ready to notice yet.”
These quotes are from Dex Horthy, with two exceptions. The “alpha in reviewing something” line is from a founder Dex met at a San Francisco dinner, as Dex retells it. The “slider between determinism and adaptability” line is the host’s.
The biggest change between summer of 2025 and now is actually like the planning metagame.
We ran this for like 4 or 5 months before we realized, like, oh, this code base is actually unusable now.
It was just like this incredibly complex system that the models developed.
Please don’t yolo a two sentence prompt and not read the code and send it to someone else to review.
I think about this in terms of expected pain, like, what is the chance you’ll have to change something later? And how painful is that?
I never argue with a model about like how what color a button is going to be or something like that, because I know that if it’s not the color I like, I’ll just send one prompt and it will fix it. But if like the database schema is wrong, then we’re probably going to have to throw it out and start over.
There’s no fast oracle for software maintainability.
If you ship bad code in your software factory, it degrades every future piece of work that goes through that factory, because there’s more bad patterns for the model to see.
You’re basically shifting from like caring about your position to caring about your velocity.
There will probably always be alpha in reviewing something. We don’t know what it is.
For now, I think it’s still the code, but eventually it might be something else.
Part of me feels that taste is just a preference that you haven’t bothered writing down.
You will get better quality if you use less context.
I always think of specs as this kind of this slider between determinism and adaptability.
| Time | Topic |
|---|---|
| 00:00 | Introduction |
| 02:15 | Meet Dexter Horthy, CEO of HumanLayer |
| 06:03 | The “dumb zone”: why more context makes models dumber |
| 06:40 | Sean Grove’s “the spec is the new code” |
| 09:22 | The YOLO pull-request experiment that broke their codebase |
| 11:09 | Letting the model own the architecture |
| 17:03 | Planning as expected-pain management |
| 28:15 | Slop Code Bench and the maintainability oracle problem |
| 37:59 | Why software factories aren’t like car factories |
| 42:49 | “There will always be alpha in reviewing something” |