State of Agentic

Spec-driven development

The Gate Moves

Write the spec first, then let the agent loose. That's the recipe everyone is selling. Anthropic's own people are doing the opposite right now — and that's the interesting bit.

Status: 2026-09-13

Three words up front, in case they're new.

Agent — a program that uses a language model to actually get things done: read files, write code, run the tests, fix what broke. Not autocomplete. More like someone working through a ticket.

Spec — the description of what's supposed to come out the other end. Can be one page. Has to be checkable.

PR (pull request) — a bundle of changes somebody looks at and signs off before it goes into the real codebase. The place where human review traditionally happens.

Two trends, pointing opposite ways

One camp says: if agents write the code, the most valuable thing you as a human produce is the spec. You write it before anything gets built. That's spec-driven development, SDD for short.

The other thing happening: the team behind Claude Code — the most widely used coding agent out there — is walking away from its own Plan Mode. Plan Mode means the agent writes a plan first, you read it, you approve, then it codes. What replaced it is Auto Mode: the agent just goes, and you look at the result.

Both of these are documented. The obvious move is to pick a side. I think that's wrong. They're not talking about the same thing.

Three recipes that look suspiciously alike

SDD isn't one method, it's a family. The resemblance is hard to miss:

  • GitHub Spec Kit runs Constitution → Specify → Plan → Tasks → Implement → Converge, and supports 30-plus different agents. (No verified publication date for it, so treat the timing as unknown.)
  • HumanLayer started with RPI — Research, Plan, Implement (2025-08-29) — and grew it into QRSPI: Questions, Research, Design, Structure, Plan, Implement. Their line for it: “Do not outsource the thinking.”
  • Dex Horthy in “Why Software Factories Fail” (2026-07-22): Product Review → System Architecture → Program Design → Vertical Slices. A vertical slice means one thin strip through every layer at once — one working feature end to end — instead of building all the database stuff, then all the API stuff, then all the screens.

Why put the human effort at the front? It's arithmetic, not taste. Horthy's version: one bad line in a plan turns into hundreds of bad lines of code. You can review a 200-line plan. Nobody reviews a 2,000-line PR — they skim it and click approve. Thirty minutes of planning buys back hours of review.

Boris Cherny says the same thing from inside Anthropic: iterate on the plan until it's good, and the implementation “will one-shot almost every time” (2026-03-04). One-shot meaning it works on the first try, no back-and-forth.

Horthy's paper is also the best argument anyone has published against full autonomy. HumanLayer ran a lights-off software factory — agents building and shipping with nobody watching — from July 2025, and called it failed that November. The Faros AI numbers they cite: 31.3% more PRs merged with no review at all, and 242.7% more incidents per PR. Incidents being the things that page you at night.

And yet the counter-trend is real

February 2026, Cherny says he starts about 80% of his sessions in Plan Mode. In the same conversation he says “plan mode has a limited lifespan,” because technically it's just a sentence injected into the prompt, and Claude had started doing it on its own anyway.

By June he says Auto Mode has replaced it for him. His reason is better than “it's faster.” He says Auto Mode is safer than asking permission, because when you ask a human 99% of the time they just click yes. A confirmation dialog nobody reads isn't a safety mechanism. It's theatre.

A month later Cat Wu's summary of how people at Anthropic actually work: “almost every single person uses auto mode” (via Willison, 2026-07-21).

And there are numbers behind it, not just vibes. Anthropic's own write-up (2026-03-25) describes a classifier — a small model that judges every single action the agent wants to take — scoring 0.4% false positives and 17% false negatives across 10,000 real tool calls. False negative meaning it waved through something it shouldn't have. They built it because they'd measured that users were confirming about 93% of prompts anyway. Underneath sits a hard containment boundary that doesn't depend on any model's judgement (2026-05-25).

So: same year, same industry. One side adding structure at the front, the other tearing it out.

They're about different things

Plan Mode approves actions. A spec describes intent. Automating the first says nothing about giving up the second.

Read the Anthropic material closely and it says the gate is moving, not going away. Planning slides into the agent. The human checkpoint slides to the result. You can see it on the timesheet: roughly 15% of the effort on implementation, 85% on fixing, testing and validating (Orosz, 2026-07-28). Human code review gets peeled back layer by layer, but the core still needs a code owner's approval.

And it's not a return to heavyweight documents either. The Claude Code team calls the old norm — six months of product requirements before anyone writes a line — dead. Idea to shipped is about a week now (2026-07-01). Cherny writes no PRDs at all and builds dozens of throwaway prototypes instead (2026-03-04).

So the spec that survives 2026 isn't the 40-page document. It's the 200-line plan you can actually read in one sitting.

Which makes the synthesis pretty boring, honestly. Automate approval for reversible actions inside a boundary that holds. Spend the attention you just freed up on the two questions no classifier will ever answer for you: what does “done” mean here, and is the finished thing actually right.

The bill always arrives at verification

Here's the thing both camps run into. Cherny: “The verification is probably the single most important thing.”

Anthropic's guidance for long-running agents splits the roles — planner, generator, evaluator — for a simple reason. An agent grading its own homework gives itself an A. Confidently.

Thorsten Ball names the ceiling on this: ask an agent for irrefutable proof that it worked, and a capable enough agent will find a way to hand you that proof. The proof is gameable. Greptile measured a version of the same asymmetry: Opus caught 53.7% of the high-severity bugs in its own code, but 62.0% in code written by Codex. Models are worse at reviewing themselves.

Which is exactly why a spec with checkable acceptance criteria isn't paperwork. It's the oracle — the independent yardstick the evaluator measures against. Without one you're asking the agent whether it's happy with itself.

Two of Willison's rules survive this whole argument intact. Green tests don't prove the behaviour you wanted (2026-03-06). And a test you never saw fail proves nothing at all (2026-02-23) — if it passed against the old code too, it isn't testing your change.

What that looks like in a small project

My own side project is built almost entirely by agents, so I had to make this concrete. It turned into gates — checkpoints a change has to pass.

The starting rules were simple. The spec is the leverage. Whoever builds doesn't check — the reviewer runs in a fresh session and sees only the diff and the criteria, never the builder's reasoning, because otherwise it just agrees. And a loop with no resistance in it is an agent nodding along to itself.

Then two gates showed up because the obvious ones weren't enough.

G1 is the plain green gate: typecheck, tests, lint. It proves nothing about whether a new test would have caught anything. A test that can't fail sails straight through. So G1r copies the new and changed test files onto a throwaway copy of the old code and demands at least one genuine failure per file. Red first, then green.

GN is a usage gate, sitting between review and deploy. It exists because on 2026-08-21 every single defect I found myself was of a kind no code gate can catch: a feature that shipped and was invisible, information parked in a tooltip on a touch device where there's no hover, a threshold that matched 74% of all objects, a default value taken from a datasheet instead of from practice, a frozen display, a component that was never installed. All of it green. All of it broken.

What nobody has settled

  • There is no trustworthy productivity number. METR couldn't reproduce its own 19% slowdown finding and threw out the experiment design (2026-02-24). People asked to estimate their own speedup overshoot by roughly 40 points. And a big speedup on tasks is perfectly compatible with almost no gain in value (2026-05-08).
  • The quality evidence contradicts itself. GitClear reports duplication up 81% and refactoring collapsing from 21% to 3.8%. Greptile reports AI pull requests getting reverted less often than human ones — 1.19 for Codex, 1.80 for Claude, 2.72 for humans, per thousand. A revert means someone had to undo it. Probably different populations being measured. Nobody has resolved it.
  • Some of this is shakier than the rest, and I'd rather say so. Spec Kit's publication date is unverified. Cherny's “auto mode replaced plan mode” comes from a video whose captions carry no speaker labels, so the attribution is inferred, not certain. The Brockman quote below sits behind a paywall and was pulled by fetch rather than read in full.
  • The most interesting question is a guess, not a finding. Where correctness can be measured by machine — kernels, compilers, test suites, benchmarks — humans have already stopped reading. OpenAI ran agent-written kernel optimisations in production for a month without anyone reading the diff, then read it afterwards (2026-09-04). Where there's no oracle — does this product make sense, is this architecture sane, is the security posture sound, is this even the right feature — nothing has replaced the human. So the question isn't “will models get good enough.” It's whether anyone builds oracles for the messy half. That framing is mine, and I'm carrying it without a source on purpose.

Shelf life

This is where the research stood on 2026-09-13, looking at roughly February through August 2026. It will go stale quickly, and the subject matter says so itself: scaffolding is a wear part, no part of Claude Code is older than six months, and Cherny's advice to anyone learning this stuff is “Forget all of the things that you learned about past models.”

Spec-Driven Development

Das Gate wandert

Erst die Spec schreiben, dann den Agent ranlassen. So lautet das Rezept, das gerade alle verkaufen. Anthropic selbst macht gerade das Gegenteil — und genau das ist der interessante Teil.

Stand: 2026-09-13

Drei Begriffe vorweg, falls sie neu sind.

Agent — ein Programm, das ein Sprachmodell benutzt, um wirklich zu arbeiten: Dateien lesen, Code schreiben, Tests laufen lassen, kaputte Sachen reparieren. Keine Autovervollständigung. Eher jemand, der ein Ticket abarbeitet.

Spec — die Beschreibung dessen, was am Ende rauskommen soll. Darf eine Seite sein. Muss überprüfbar sein.

PR (Pull Request) — ein Paket Änderungen, das jemand anschaut und freigibt, bevor es in den echten Code wandert. Die Stelle, an der klassischerweise ein Mensch draufschaut.

Zwei Trends, zwei Richtungen

Das eine Lager sagt: Wenn Agents den Code schreiben, ist das Wertvollste, was du als Mensch beisteuerst, die Spec. Und die schreibst du, bevor irgendwas gebaut wird. Das ist Spec-Driven Development, kurz SDD.

Und gleichzeitig: Ausgerechnet das Team hinter Claude Code — dem meistgenutzten Coding-Agent überhaupt — verabschiedet sich vom eigenen Plan Mode. Plan Mode heißt, der Agent schreibt erst einen Plan, du liest ihn, gibst frei, dann wird gecodet. Abgelöst hat ihn der Auto Mode: Der Agent macht einfach, und du schaust hinterher drauf.

Beides ist belegt. Der Reflex ist, sich für eine Seite zu entscheiden. Ich halte das für falsch. Die zwei reden schlicht über verschiedene Dinge.

Drei Rezepte, die verdächtig ähnlich aussehen

SDD ist keine Methode, sondern eine Familie. Die Ähnlichkeit ist schwer zu übersehen:

  • GitHub Spec Kit fährt Constitution → Specify → Plan → Tasks → Implement → Converge und unterstützt über 30 verschiedene Agents. (Ein verifiziertes Veröffentlichungsdatum gibt es nicht, also nimm die zeitliche Einordnung mit Vorsicht.)
  • HumanLayer fing mit RPI an — Research, Plan, Implement (2025-08-29) — und hat daraus QRSPI gemacht: Questions, Research, Design, Structure, Plan, Implement. Ihr Leitsatz dazu: „Do not outsource the thinking.“
  • Dex Horthy in „Why Software Factories Fail“ (2026-07-22): Product Review → System Architecture → Program Design → Vertical Slices. Vertical Slice heißt: ein dünner Streifen quer durch alle Schichten auf einmal, also ein Feature komplett lauffähig — statt erst alles an der Datenbank, dann alles an der Schnittstelle, dann alle Bildschirme.

Warum den Menschen nach vorne setzen? Das ist Rechnen, nicht Geschmack. Horthys Version: Eine schlechte Zeile im Plan wird zu hunderten schlechter Codezeilen. Einen 200-Zeilen-Plan kannst du lesen. Einen 2.000-Zeilen-PR liest niemand — der wird überflogen und durchgewunken. Dreißig Minuten Planung holen Stunden Review zurück.

Boris Cherny sagt dasselbe von innen: den Plan iterieren, bis er gut ist, dann läuft die Implementierung „fast jedes Mal“ im ersten Anlauf durch (2026-03-04). One-Shot nennt er das — funktioniert sofort, kein Hin und Her.

Horthys Papier ist gleichzeitig das beste veröffentlichte Argument gegen volle Autonomie. HumanLayer hat ab Juli 2025 selbst eine Lights-off-Factory betrieben — Agents bauen und liefern aus, niemand schaut zu — und sie im November für gescheitert erklärt. Die Faros-AI-Zahlen, die sie zitieren: 31,3 % mehr PRs, die komplett ohne Review durchgingen, und 242,7 % mehr Incidents pro PR. Incidents sind die Dinger, wegen denen nachts das Telefon klingelt.

Und trotzdem ist der Gegentrend echt

Februar 2026: Cherny sagt, er startet rund 80 % seiner Sessions im Plan Mode. Im selben Gespräch sagt er, „plan mode has a limited lifespan“ — technisch sei das nur ein Satz, der in den Prompt geschoben wird, und Claude mache das inzwischen sowieso von allein.

Im Juni sagt er, für ihn habe Auto Mode den Plan Mode abgelöst. Seine Begründung ist besser als „geht schneller“: Auto Mode sei sicherer als Nachfragen. Denn wenn du einen Menschen fragst, klickt der in 99 % der Fälle einfach auf Ja. Ein Bestätigungsdialog, den keiner liest, ist keine Sicherheitsmaßnahme. Das ist Theater.

Einen Monat später fasst Cat Wu die interne Praxis so zusammen: „almost every single person uses auto mode“ (via Willison, 2026-07-21).

Und dahinter stehen Zahlen, nicht nur Bauchgefühl. Anthropic beschreibt (2026-03-25) einen Klassifikator — ein kleines Modell, das jede einzelne Aktion bewertet, die der Agent ausführen will. Ergebnis über 10.000 echte Tool-Calls: 0,4 % False Positives, 17 % False Negatives. False Negative heißt, es hat etwas durchgelassen, das es nicht hätte durchlassen dürfen. Gebaut haben sie das, weil sie gemessen hatten, dass Nutzer ohnehin rund 93 % aller Rückfragen bestätigen. Darunter liegt eine harte Containment-Grenze, die von der Einschätzung keines Modells abhängt (2026-05-25).

Also: gleiches Jahr, gleiche Branche. Die eine Seite baut vorne Struktur auf, die andere reißt sie raus.

Es geht um verschiedene Dinge

Plan Mode gibt Aktionen frei. Eine Spec beschreibt Absicht. Das eine zu automatisieren sagt nichts darüber, das andere aufzugeben.

Wenn man das Anthropic-Material genau liest, steht da: Das Gate wandert, es verschwindet nicht. Die Planung rutscht in den Agent. Der menschliche Kontrollpunkt rutscht ans Ergebnis. Man sieht das am Stundenzettel: grob 15 % des Aufwands auf Implementierung, 85 % auf Fixen, Testen, Validieren (Orosz, 2026-07-28). Menschliches Code-Review wird Schicht für Schicht abgebaut, aber den Kern muss weiter ein Code Owner freigeben.

Und es ist auch keine Rückkehr zu dicken Vorab-Dokumenten. Das Claude-Code-Team erklärt die alte Norm — sechs Monate Anforderungsdokument, bevor jemand eine Zeile schreibt — für tot. Von der Idee zum Ausgelieferten ist es heute etwa eine Woche (2026-07-01). Cherny schreibt gar keine PRDs mehr und baut stattdessen Dutzende Wegwerf-Prototypen (2026-03-04).

Die Spec, die 2026 überlebt, ist also nicht das 40-Seiten-Dokument. Es ist der 200-Zeilen-Plan, den du in einem Rutsch lesen kannst.

Womit die Auflösung ehrlich gesagt ziemlich unspektakulär ist. Freigaben für umkehrbare Aktionen automatisieren, innerhalb einer Grenze, die hält. Und die Aufmerksamkeit, die dadurch frei wird, in die zwei Fragen stecken, die dir kein Klassifikator beantwortet: Was heißt hier eigentlich „fertig“, und ist das Fertige richtig.

Die Rechnung kommt immer bei der Verifikation

Und hier landen beide Lager. Cherny: „The verification is probably the single most important thing.“

Anthropics Leitlinie für lang laufende Agents trennt die Rollen — Planner, Generator, Evaluator — aus einem simplen Grund. Ein Agent, der seine eigene Arbeit benotet, gibt sich eine Eins. Selbstbewusst.

Thorsten Ball benennt die Obergrenze: Verlang von einem Agent einen unwiderlegbaren Beweis, dass es funktioniert, und ein hinreichend fähiger Agent findet einen Weg, dir genau diesen Beweis zu liefern. Der Beweis ist manipulierbar. Greptile hat eine Variante davon gemessen: Opus fand 53,7 % der schweren Bugs im eigenen Code, aber 62,0 % in Code, den Codex geschrieben hatte. Modelle sind schlechter darin, sich selbst zu reviewen.

Und genau deshalb ist eine Spec mit überprüfbaren Abnahmekriterien kein Papierkram. Sie ist das Orakel — der unabhängige Maßstab, gegen den der Evaluator misst. Ohne das fragst du den Agent, ob er mit sich zufrieden ist.

Zwei Regeln von Willison überstehen den ganzen Streit unbeschadet. Grüne Tests belegen nicht das Verhalten, das du wolltest (2026-03-06). Und ein Test, den du nie hast fallen sehen, belegt gar nichts (2026-02-23) — wenn er gegen den alten Code genauso durchlief, testet er deine Änderung nicht.

Wie das in einem kleinen Projekt aussieht

Mein eigenes Nebenprojekt wird fast komplett von Agents gebaut, also musste ich das konkret machen. Rausgekommen sind Gates — Kontrollpunkte, die eine Änderung passieren muss.

Die Ausgangsregeln waren simpel. Die Spec ist der Hebel. Wer baut, prüft nicht — der Reviewer läuft in einer frischen Session und sieht nur den Diff und die Kriterien, nie die Begründung des Bauers, weil er sonst einfach zustimmt. Und ein Loop ohne Widerstand ist ein Agent, der sich selbst zunickt.

Dann kamen zwei Gates dazu, weil die naheliegenden nicht reichten.

G1 ist das reine Grün-Gate: Typecheck, Tests, Lint. Es belegt nicht, ob ein neuer Test überhaupt irgendwas gefangen hätte. Ein Test, der gar nicht fallen kann, segelt da durch. Also kopiert G1r die neuen und geänderten Testdateien auf eine Wegwerf-Kopie des alten Codes und verlangt je Datei mindestens einen echten Fehlschlag. Erst rot, dann grün.

GN ist ein Nutzungs-Gate zwischen Review und Deploy. Das gibt es, weil am 2026-08-21 sämtliche Fehler, die ich selbst gefunden habe, von einer Sorte waren, die kein Code-Gate fängt: ein Feature, das ausgeliefert und unsichtbar war. Information, die in einem Tooltip steckte — auf einem Touchgerät, wo es kein Hover gibt. Eine Schwelle, die auf 74 % aller Objekte zutraf. Ein Vorgabewert aus dem Datenblatt statt aus der Praxis. Eine eingefrorene Anzeige. Ein Baustein, der nie installiert wurde. Alles grün. Alles kaputt.

Was niemand geklärt hat

  • Es gibt keine belastbare Produktivitätszahl. METR konnte das eigene Ergebnis — 19 % Verlangsamung — nicht reproduzieren und hat das Experimentdesign verworfen (2026-02-24). Leute, die ihren eigenen Speedup schätzen sollen, liegen rund 40 Prozentpunkte daneben. Und ein großer Speedup bei Aufgaben verträgt sich problemlos mit fast keinem Zugewinn an Wert (2026-05-08).
  • Die Qualitätsbefunde widersprechen sich. GitClear meldet 81 % mehr Duplikate und einen Refactoring-Anteil, der von 21 % auf 3,8 % einbricht. Greptile meldet, dass KI-PRs seltener zurückgenommen werden als menschliche — 1,19 bei Codex, 1,80 bei Claude, 2,72 bei Menschen, je tausend. Zurücknehmen heißt, jemand musste es rückgängig machen. Wahrscheinlich werden da verschiedene Populationen gemessen. Aufgelöst hat es niemand.
  • Ein Teil davon ist wackliger als der Rest, und das sage ich lieber dazu. Das Veröffentlichungsdatum von Spec Kit ist unverifiziert. Chernys „Auto Mode hat Plan Mode abgelöst“ stammt aus einem Video, dessen Untertitel keine Sprecher ausweisen — die Zuordnung ist erschlossen, nicht sicher. Das Brockman-Zitat unten liegt hinter einer Paywall und wurde per Fetch geholt, nicht im Volltext gelesen.
  • Die interessanteste Frage ist eine Vermutung, kein Befund. Wo Korrektheit maschinell messbar ist — Kernel, Compiler, Testsuiten, Benchmarks — hat der Mensch schon aufgehört zu lesen. OpenAI hat agent-geschriebene Kernel-Optimierungen einen Monat in Produktion laufen lassen, ohne dass jemand den Diff gelesen hat, und ihn erst danach angeschaut (2026-09-04). Wo es kein Orakel gibt — ergibt dieses Produkt Sinn, ist die Architektur gesund, stimmt die Security-Lage, ist das überhaupt das richtige Feature — hat nichts den Menschen ersetzt. Die Frage lautet also nicht „werden die Modelle gut genug“. Sie lautet, ob jemand Orakel für die unscharfe Hälfte baut. Diese Rahmung ist meine, und ich führe sie bewusst ohne Quelle.

Haltbarkeit

Das ist der Stand vom 2026-09-13, mit Blick auf etwa Februar bis August 2026. Das wird schnell alt, und der Gegenstand sagt das selbst: Scaffolding ist ein Verschleißteil, kein Teil von Claude Code ist älter als sechs Monate, und Chernys Rat an alle, die das gerade lernen, lautet „Forget all of the things that you learned about past models.“

SourcesQuellen

  1. We Cut 80% of Claude Code’s Prompt (YC Startup School) Boris Cherny · 2026-07-27 · Transkript: ycrootaccess.com/p/boris-cherny-building-claude-code
  2. Building Claude Code with Boris Cherny The Pragmatic Engineer · 2026-03-04
  3. Inside Claude Code With Its Creator YC Lightcone · 2026-02-17 · Video nur Kapitelbeschreibung ausgewertet
  4. Reflecting on a year of Claude Code Boris Cherny + Cat Wu · 2026-06-08 · erschlossen Untertitel ohne Sprecherlabels
  5. Claude Fable, Claude Tag, and Anthropic’s Culture Cat Wu, Thariq Shihipar · 2026-07-15 · Bericht von Simon Willison
  6. Claude Code auto mode Anthropic · 2026-03-25
  7. How we contain Claude Anthropic · 2026-05-25
  8. Harness design for long-running application development Prithvi Rajasekaran · 2026-03-24
  9. Why Software Factories Fail Dex Horthy (HumanLayer) · 2026-07-22
  10. Advanced Context Engineering (ACE-FCA) Dex Horthy · 2025-08-29 · Ursprung von RPI
  11. GitHub Spec Kit GitHub · Datum unverifiziert
  12. Agentic Engineering Patterns Simon Willison · Red/Green-TDD 2026-02-23 · Agentic manual testing 2026-03-06
  13. Models are worse at reviewing their own code Greptile · 2026-07-21
  14. Joy & Curiosity #95 Thorsten Ball · 2026-08-16
  15. The Final Bottleneck Armin Ronacher · 2026-02-13
  16. How building software is changing at Anthropic Gergely Orosz · 2026-07-28
  17. Changing our Developer Productivity Experiment Design METR · 2026-02-24 · dazu: Task Substitution and Uplift, 2026-05-08
  18. Write-Only Mode: AI Code Quality in 2026 GitClear · 2026-01
  19. Rise of the Overnight Agents Greptile · 2026-05-05
  20. An Interview with OpenAI President Greg Brockman About Astra and Alignment Ben Thompson (Stratechery) · 2026-09-04 · Paywall per Fetch geholt, nicht im Volltext gelesen
  21. Internal working document on the agentic process, as of August 2026Internes Arbeitsdokument zum agentischen Prozess, Stand August 2026 Gates G0–GN, Rot-vor-Grün-Gate, Nutzungs-Gate; GN-Befunde vom 2026-08-21 · keine öffentliche URL