Spec-Driven Development: What It Fixes, What It Costs, and When to Use It
By Sergey Nosov
29 September 2026
AI coding assistants made it cheap to produce code. They did not make it cheap to produce the right code. On 2 February 2025, Andrej Karpathy gave a popular way of working a name: vibe coding. In his words, you “fully give in to the vibes, embrace exponentials, and forget that the code even exists.” For a prototype or an experiment, that is a reasonable way to work. For a production system, it breaks down in predictable ways. The agent forgets context between sessions. Two developers prompt for the same feature and get two different designs. Months later, nobody can say what the code was supposed to do.
Spec-driven development is one structured answer. Its premise is simple. Write down what you want and why, in a form an agent can follow. Then make that document, not a chat history, the thing everyone works from. GitHub released its Spec Kit toolkit on 2 September 2025. In the announcement, Den Delimarsky summed up the problem it addresses: “We treat coding agents like search engines when we should be treating them more like literal-minded pair programmers.”
This article explains what spec-driven development is and walks through its workflow with an example. It compares the approach with practices you already know and looks honestly at the evidence, the limitations, and the tools as of September 2026. Used in the right place, the approach earns its overhead. Used everywhere, it turns into the heavyweight process that agile teams worked hard to leave behind.
What Spec-Driven Development Is
Spec-driven development inverts the usual relationship between documentation and code. In most projects the code is the truth, and the documentation chases it and falls behind. In spec-driven development the specification is the source of truth, and the code follows from it. When requirements change, you change the specification first and then regenerate or revise the affected code.
Four principles hold the approach together:
- The specification is the source of truth. When a question comes up about what the system should do, you consult the spec, not the code.
- Intent comes before implementation. The spec says what the system should do and why. How it does so is decided later, in a technical plan.
- Specs are living artifacts. They sit in version control next to the code, change with the project, and actively drive the agent’s work.
- People validate at checkpoints. Each phase ends with a person reviewing its output before the next phase begins. This is structured collaboration, not autopilot.
The idea is older than coding agents. Contract-first design has long treated an agreed interface specification as the authoritative artifact from which implementations follow. What is new is the reader: the spec is now written for an agent as much as for people.
The term also covers practices of very different ambition. Birgitta Bockeler, a Distinguished Engineer at Thoughtworks, distinguished three levels on 15 October 2025:
- Spec-first: “A well thought-out spec is written first, and then used in the AI-assisted development workflow for the task at hand.”
- Spec-anchored: “The spec is kept even after the task is complete, to continue using it for evolution and maintenance of the respective feature.”
- Spec-as-source: “The spec is the main source file over time, and only the spec is edited by the human, the human never touches the code.”
Spec-as-source is the most radical of the three. It assumes that regenerating code from a spec is safe, and today’s agents do not produce the same code twice from the same spec.
The Workflow, Phase by Phase
Tools name their steps differently, but most spec-driven workflows follow the same arc. I find it clearest in six phases: constitution, specify, clarify, plan, tasks, and implement. To make them concrete, suppose a product manager asks for “a notification system so users know when important events happen.”
1. Constitution: The Rules That Do Not Change
The constitution records a project’s non-negotiable standards once. You stop repeating them in every prompt, and the agent applies them to every feature. A short example:
# constitution.md
## Coding standards
- TypeScript in strict mode
- Public functions documented
- Functions of 50 lines or fewer
## Architecture
- REST APIs described in OpenAPI
- Data access via repositories
- No database calls in controllers
## Quality
- Every feature ships with tests
- No unmaintained dependencies
- Runtime versions pinned here
Keep it short. One or two pages is plenty; a constitution nobody can remember is a constitution nobody enforces.
2. Specify: What and Why, Not How
The specification captures intent: the purpose, user stories, acceptance criteria, edge cases, and how you will know the feature succeeded. It deliberately leaves out schemas, endpoints, and code.
# Spec: User Notifications
## Purpose
Users need timely awareness of
account activity and system
events without checking
manually.
## User stories
- As a user, I want to be
notified when someone
comments on my content.
- As a user, I want to choose
which notifications I get:
email, in-app, or both.
3. Clarify: Find What the Spec Hides
This is the phase I would never skip. The agent reads the spec and asks the questions it leaves open:
- How long should notifications be kept?
- What happens to a user’s notifications when the account is deleted?
- Should notifications be batched during high-volume periods?
- How do users reach their notification history?
- Which notification types must be delivered immediately?
Without this step, each person on the team answers those questions privately and differently. The backend developer assumes notifications live forever, the product manager assumes thirty days, and whoever owns data retention assumes something else again. With it, the answers go back into the spec:
## Clarifications
### Retention
- Keep notifications 90 days,
then archive them
### Account deletion
- Delete all notifications
when an account closes
### Delivery
- Immediate: security alerts,
mentions
- Batched hourly: comments,
likes
Clarification is requirements auditing: a deliberate hunt for ambiguity, contradiction, and gaps before they turn into code. It is the subject of my book Finding What Requirements Hide, and in my view it is where spec-driven development earns its keep. Note that GitHub’s Spec Kit now treats clarification as optional, one of the extra steps you add “when you need extra quality gates.” For any feature that matters, I would switch it on.
4. Plan: Now Decide How
Only now does the work turn technical. For the notification system, the plan settles the architecture, the data model, and the interfaces:
- a notifications table with an identifier, the user, the type, a payload, a read flag, and the creation time;
- endpoints to list notifications and to mark one as read;
- a background job that delivers notifications;
- a real-time channel for in-app delivery.
Notice where the choice of transport lives. A spec that says “use WebSockets” has let implementation leak into intent. The spec should say that users see important notifications immediately, and the plan decides whether that means WebSockets, server-sent events, or polling.
5. Tasks: Small, Ordered, and Reviewable
The plan breaks down into tasks with an explicit order:
- Create the database migration for the notifications table.
- Implement the notification repository.
- Build the REST endpoints and their OpenAPI description.
- Create the delivery service and its background job.
- Implement the real-time channel.
- Build the notification component in the user interface.
- Write the end-to-end integration tests.
Unit tests travel with each task; the constitution already requires them.
6. Implement: One Focused Change at a Time
The agent works through the tasks, and you review each change as it lands: the migration in one pull request, the repository in the next, and the endpoints after that. Seven focused reviews are easier to do well than one review of a pull request that touches thirty files. Reviewing code you did not write is a discipline of its own, and it is the subject of my book Code You Did Not Write.
How It Relates to TDD, BDD, Waterfall, and Big Design Up Front
Spec-driven development overlaps with practices most teams already know, and the overlaps explain both its value and its risks.
Test-driven development works at the level of units. A short cycle of failing test, passing code, and refactoring uses tests to shape the design of small pieces of code. Behavior-driven development works at the level of user-visible behavior, with scenarios written in plain language that stakeholders can read and a test runner can execute. Spec-driven development works at the level of features and systems. Its main artifact is the specification, which guides an agent in producing both code and tests.
These are layers, not rivals. A team can use a spec to agree on a feature, behavior scenarios to pin down end-to-end behavior, and test-driven development for the units underneath.
One difference matters more than the rest. A behavior-driven scenario is executable: when the system stops behaving as described, the scenario fails. A spec in spec-driven development is prose that an agent is asked to follow, and by default nothing checks that the code still matches it. Tools have started to close that gap. When Kiro became generally available on 17 November 2025, it added property-based testing that, in Kiro’s words, “measures whether your code actually matches what you specified.” Until such checks are routine, treat the spec as a promise, not a proof, and keep real tests in the loop.
The most common objection is that spec-driven development is waterfall with better branding. It deserves a serious answer. Waterfall moves through sequential phases with gates, tries to specify everything up front, and shows working software only after months. Spec-driven development, done well, loops. You capture intent at the level of detail the task needs, see code within hours or days, and change the spec whenever you learn something. Big design up front assumes you can predict every requirement. Spec-driven development assumes you cannot, and gives you a cheap place to record what you learn.
Done badly, though, it feels exactly like waterfall. When Bockeler asked Kiro to fix a small bug, its requirements document turned the fix into four user stories with sixteen acceptance criteria in total. She compared it to “using a sledgehammer to crack a nut.” Kiro’s documentation now describes a separate bugfix spec alongside feature specs. Thoughtworks placed spec-driven development in the Assess ring of its Technology Radar in November 2025. It added a sharper warning: “We may be relearning a bitter lesson—that handcrafting detailed rules for AI ultimately doesn’t scale.”
What the Evidence Says
The strongest case for spec-driven development is not speed. A widely cited controlled study of AI coding tools, published by METR on 10 July 2025, did not test spec-driven development, but its result is a useful caution. Sixteen experienced open-source developers completed 246 real issues in large, mature repositories, which averaged more than 22,000 stars and more than a million lines of code. When the developers could use AI tools, mainly Cursor Pro with Claude 3.5 and 3.7 Sonnet, they took 19 percent longer. Before the study they had expected AI to speed them up by 24 percent, and afterward they still believed it had sped them up by 20 percent.
METR itself cautions against generalizing: its developers and repositories do not represent most software work. In a 24 February 2026 update, METR said it believes developers are likely “more sped up from AI tools now” than in early 2025. It also said its newer data gives “an unreliable signal,” in part because more developers now decline to take part in a study that requires some work without AI.
The lesson for spec-driven development is to adopt it for alignment, consistency, and a durable record of intent, not in the expectation of writing code faster. If speed is the only goal, the evidence does not yet promise it.
The Honest Limitations
- Agents do not always follow the spec. Bockeler reported that despite all the files, templates, and checklists, “I frequently saw the agent ultimately not follow all the instructions.” The checkpoints exist to catch exactly this.
- The same spec does not produce the same code. Generation is nondeterministic, which is why regenerating code from a spec is not like recompiling source code.
- Agents over-apply rules. A constitution rule such as “prefer functional patterns” can turn a simple function into an elaborate composition that nobody wants to maintain.
- Specs can drown their reviewers. “Spec-kit created a LOT of markdown files for me to review,” Bockeler wrote, and they were repetitive. A spec nobody reads protects nobody.
- Writing specs is a different skill. Developers who are fast in code are not automatically precise in prose, and the approach feels slower at first. The objections I hear most, “this breaks my flow” and “writing specs takes longer than just coding,” have real truth in them early on.
- Some work gains nothing. Visual design, novel algorithms that depend on deep expertise, and quick exploration rarely benefit from a written spec.
When to Use It, and When Not To
Ask one question of any task: will writing the specification first reduce the total effort? It usually will for:
- complex features in existing systems, where architectural constraints matter;
- work that needs several groups, such as product, engineering, design, and legal, to agree;
- efforts that span days or weeks, where the agent needs context that survives between sessions;
- regulated work that needs an audit trail of intent;
- legacy modernization, where the spec captures what the old system actually does;
- distributed teams, where alignment is expensive.
It usually will not for:
- prototypes and exploration, where learning matters more than structure;
- small bug fixes, where the ceremony outweighs the change;
- requirements that change daily, where maintaining the spec becomes its own project;
- novel algorithms that depend on deep human expertise;
- a single developer working on familiar code;
- visual design work, which a text spec describes poorly;
- an urgent deadline that leaves no room to learn a new method.
Karpathy drew a similar line on 30 April 2026. Vibe coding, he wrote, is about “raising the floor for everyone in terms of what they can do in software.” Agentic engineering “is about preserving the quality bar of professional software.” Spec-driven development is one way to do the second. Vibe coding, spec-driven development, and plain test-driven coding are complementary tools, not competing beliefs; match the method to the task.
The Tools, as of September 2026
The tooling changes monthly, so treat this as a snapshot.
- GitHub Spec Kit is open source and works with several coding agents, including GitHub Copilot, Claude Code, and Gemini CLI. GitHub announced it with four phases: specify, plan, tasks, and implement. Its current core flow starts with a constitution and runs through implementation, and it adds clarification, checklists, and consistency analysis on request.
- Kiro, from Amazon, is an agentic IDE introduced on 14 July 2025 and generally available since 17 November 2025, together with a command-line version. It writes a feature spec as three files,
requirements.md,design.md, andtasks.md, and it offers a separate bugfix spec. - Tessl pursues spec-as-source. As Bockeler describes it, the generated code even carries a comment at the top:
// GENERATED FROM SPEC - DO NOT EDIT. Tessl launched its Spec Registry in open beta and its framework in closed beta on 23 September 2025. - OpenSpec, which Thoughtworks added to the Assess ring of its Technology Radar in April 2026, takes a lighter route. Its three-step workflow is propose, apply, and archive, and Thoughtworks notes its focus “on spec deltas rather than defining a complete specification upfront,” which makes it “well-suited for existing systems.”
How to Start
- Run a pilot. Pick a few volunteers and a handful of medium-sized features. Keep high-pressure deliverables out of the pilot, and keep ordinary development available as a fallback.
- Use one tool. Comparing several tools while learning the method teaches you neither.
- Keep specs short. A one- or two-page constitution and a two- to five-page feature spec are enough for most work.
- Put tests in the spec. Quality belongs in the specification, not after it.
- Version and review specs like code. Commit them, review them in pull requests, and change them before you change behavior.
- Refine specs from what the agent does. Your first spec will not be your best. Adjust it wherever the agent misreads it.
- Measure honestly. Track time, quality, and how the team feels, and give the method more than a sprint or two before you judge it.
Takeaways
- Spec-driven development makes a written specification, not the chat history, the source of truth for AI-assisted work.
- The workflow runs from a constitution through specify, clarify, plan, and tasks to implementation, with a human checkpoint after each phase.
- Clarification is the most valuable phase: it finds ambiguities, contradictions, and gaps before they become code.
- It complements test-driven and behavior-driven development instead of replacing them, and its specs are not executable checks.
- It is not waterfall when it loops in small steps, and it becomes waterfall when it does not.
- Adopt it for alignment and a durable record of intent. The evidence does not yet promise faster coding.
- Match the method to the task.
Further Reading
- Spec-Driven Development with AI: Get Started with a New Open Source Toolkit (GitHub, 2 September 2025)
- GitHub Spec Kit (GitHub)
- Understanding Spec-Driven-Development: Kiro, spec-kit, and Tessl (Birgitta Bockeler, martinfowler.com, 15 October 2025)
- Specs (Kiro documentation)
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (METR, 10 July 2025)
- We Are Changing Our Developer Productivity Experiment Design (METR, 24 February 2026)
- Spec-Driven Development (Thoughtworks Technology Radar, November 2025)