Software development has always evolved alongside the tools we use to build software.
We moved from manual coding to IDEs, from waterfall to Agile, from manual testing to automated testing, from heavyweight releases to continuous delivery. Practices such as Test-Driven Development (TDD), Continuous Integration, code review and DevOps changed where and how we invested engineering effort.
Now we are entering another significant transition.
AI is no longer just helping engineers write code. Increasingly, AI agents can take a requirement, reason about a codebase, create a plan, implement changes, run tests, fix failures and produce a pull request with limited human intervention.
That changes something fundamental.
When humans were the primary producers of code, the amount of code an engineering team could create was constrained by human typing and reasoning capacity.
With agents, that constraint is disappearing.
The new constraint is increasingly human attention, context and verification capacity.
This raises an important question for engineering leaders:
If AI can produce code faster than humans can reasonably review it, should code remain the primary artifact that humans review?
Our recent discussions around this question led my team to explore Spec-Driven Development (SDD) and what it could mean for the broader Software Development Life Cycle.
This post shares some of the thinking that emerged from those discussions. It is not a prescription for every engineering team. It is an exploration of what currently makes sense for us and where I believe the industry may be heading.
The journey from TDD to SDD
TDD changed the way we thought about software development.
Instead of:
Requirement → Code → Test
we started thinking more like:
Requirement → Test → Code
The test became an executable expression of expected behaviour.
TDD did not eliminate the need for developers to understand the system. Quite the opposite. It encouraged developers to think about behaviour and boundaries before implementation.
Now consider the AI-agent era.
The problem is no longer simply that humans need to write code correctly.
The problem is that an AI agent can generate far more implementation than a human can comfortably inspect.
That suggests another shift:
Requirement → Specification → Tests + Code
The specification becomes the higher-level contract.
This is where Spec-Driven Development becomes interesting.
SDD is not an entirely new idea. Software engineering has long used requirements documents, RFCs, API contracts, architecture decision records and formal specifications. What is changing is their role.
With AI agents, specifications can become much more directly connected to implementation.
Microsoft describes SDD as making structured specifications the shared source of truth for humans and AI, aligning intent before AI accelerates implementation. IBM similarly describes SDD as a methodology where detailed specifications guide what is built and how it is built. (Microsoft Developer)
The important change is therefore not simply:
“Let’s write more documentation.”
It is:
“Let’s move more of our engineering intent into an artifact that both humans and agents can reason about.”
Why AI changes the code-review equation
AI-assisted development creates an interesting asymmetry.
The cost of producing code is falling rapidly.
The cost of understanding and verifying that code has not fallen at the same rate.
An engineer can ask an agent to modify several parts of a system, generate tests, refactor an existing component and update integrations in parallel.
The resulting pull request may technically be correct.
But correctness is not the only question.
A reviewer still needs to ask:
- Did the implementation solve the actual business problem?
- Did the agent make assumptions that nobody explicitly approved?
- Did it change an existing system boundary?
- Did it introduce a new dependency?
- Did it alter security behaviour?
- Did it expose or retain data that should have been private?
- Did it change an API contract?
- Did it introduce an architectural pattern that conflicts with the rest of the system?
- Did it solve the requested problem while creating another one elsewhere?
These questions are difficult to answer by staring at thousands of lines of generated code.
And this is becoming a recognised industry tension.
DORA’s research into AI-assisted software development found that AI acts as an amplifier of an organisation’s existing strengths and weaknesses. Its 2026 analysis also highlights a particularly interesting tension: AI can accelerate initial development while shifting some of the saved time into auditing and verification. (Dora)
In other words:
AI may remove coding bottlenecks without removing delivery bottlenecks.
It can simply move the bottleneck downstream.
Move the review left: review intent before implementation
This is where SDD starts becoming compelling.
Instead of asking:
“Does this code do what we want?”
we can first ask:
“Have we clearly described what we want?”
A good specification should establish things such as:
Functional behaviour
What should the system do?
Boundaries
What should it not do?
System interactions
Which services, APIs, databases or external systems can it interact with?
Constraints
What architectural or technical constraints must be respected?
Security
What authentication, authorisation, secrets management or security requirements apply?
Privacy
What data can be collected, stored, processed or exposed?
Compliance
Are there GDPR, regulatory, contractual or organisational requirements?
Testing
What behaviours must be tested?
Acceptance criteria
How do we know that the requirement has actually been satisfied?
Non-goals
What is explicitly outside the scope of this change?
That last one is particularly important for AI agents.
Humans naturally fill in gaps using context, experience and conversations.
LLMs also fill gaps.
The difference is that an AI agent can turn an assumption into hundreds of lines of code very quickly.
A specification gives the agent fewer gaps to fill.
The specification becomes a contract
I increasingly think about the specification as a contract between several participants:
Business → Product → Engineering → AI agents → Tests → Production
The specification should preserve the original intent as it moves through that chain.
Without it, requirements can gradually transform:
Business intent
↓
Ticket
↓
Developer interpretation
↓
AI interpretation
↓
Generated implementation
↓
Production behaviour
Every arrow introduces an opportunity for information to be lost or assumptions to be introduced.
SDD attempts to reduce that loss by making the specification a persistent, reviewable artifact.
The emerging academic discussion around agentic software engineering describes specifications in a similar way, as a “contract substrate” between humans and agents, while acknowledging that SDD is still an emerging discipline rather than a mature, universally validated methodology. (arXiv)
That distinction matters.
SDD is promising, but it is not yet a settled industry standard.
What does an SDD specification actually look like?
It does not necessarily need to be a giant requirements document.
In fact, I think smaller specifications are often more useful.
A specification could live directly in the repository as Markdown and describe something like:
Feature: Customer notification preferencesGoal: Allow customers to control which notifications they receive.Functional requirements: Customers can enable or disable email notifications. Customers can enable or disable SMS notifications. Existing preferences must be preserved.Security: Only the authenticated customer can modify their preferences. A customer cannot modify another customer's preferences.Privacy: Phone numbers must not be exposed in API responses unless required. Preference changes must not be written to application logs.API: PUT /customers/{id}/preferencesTesting: Test authenticated access. Test unauthorized access. Test persistence. Test invalid preference combinations.Non-goals: Do not introduce a new notification provider. Do not modify notification delivery logic
The exact format is less important than the principle.
The agent should have a clear, inspectable contract to work against.
From tickets to specifications as the unit of work
This leads to another interesting consequence.
In traditional development, we often optimise our workflow around tickets.
One engineer might have two or three tickets in progress.
But an AI-enabled engineer can potentially have several agents working simultaneously:
Engineer | +-- Agent A -> API implementation | +-- Agent B -> Tests | +-- Agent C -> Database changes | +-- Agent D -> Documentation
The limiting factor is no longer necessarily how many tickets the engineer can technically have open.
It is how many pieces of intent the engineer can actively understand, refine and govern.
That led us to an interesting experiment:
Should WIP limits apply to specifications rather than tickets?
For our team, a potential model is something like three or four active specifications per engineer.
Within each specification, multiple agents may operate.
The human remains responsible for understanding and refining the specification, resolving ambiguity and validating the resulting outcome.
This creates a different relationship between humans and agents:
Human capacity becomes the WIP constraint. Agent capacity becomes the execution capacity.
This is still an area we are experimenting with, not a universal recommendation.
What happens to code review?
This is perhaps the most controversial part of the conversation.
If the specification has been reviewed and approved, and AI agents generate the implementation, do we still need humans to review every line of code?
I don’t think the answer should simply be “yes” or “no”.
Instead, I expect code review to become increasingly risk-based.
For example:
Low-risk change
- Small scope
- Strong specification
- Automated tests
- No sensitive data
- No architectural change
Potentially:
Spec review → automated validation → AI review → merge
Medium-risk change
Additional human review may be appropriate.
High-risk change
For example:
- Security-sensitive functionality
- Authentication/authorisation
- Financial transactions
- Personal data
- Major architecture changes
- Public APIs
- Critical infrastructure
Human review remains essential.
This is a more sustainable model than either extreme.
The goal should not be:
“Humans should never review code again.”
Nor should it be:
“AI can write everything, but humans must manually inspect everything.”
The more interesting question is:
Where does human judgment create the most value?
Cloudflare’s experience with AI-assisted code review is also instructive here. It describes code review as a potential bottleneck and explores using AI to handle parts of the review process while retaining appropriate human oversight. (Cloudflare Blog)
The next evolution: specialised agents
Once we start treating specifications as contracts, another possibility emerges.
Instead of asking one general-purpose coding agent to think about everything, we can introduce specialised agents.
For example:
Specification
|
+---------+---------+
| | |
Coding Agent Security Agent Privacy Agent
| | |
+---------+---------+
|
Test Agent
|
Compliance Agent
|
Human Owner
A security agent could check authentication and authorisation requirements.
A privacy agent could check GDPR-related constraints.
A compliance agent could validate organisational policies.
A copyright or licensing agent could inspect dependency and content usage.
A testing agent could validate whether acceptance criteria are adequately covered.
This creates something resembling an AI quality pipeline around the implementation agent.
And importantly, these concerns can move earlier in the lifecycle.
Security and privacy stop being things we remember during code review.
They become properties of the specification and automated validation process.
What happens to Agile?
This is where the conversation becomes more interesting than SDD alone.
If AI agents can execute several pieces of work concurrently, our traditional assumptions about planning start to change.
Our team currently works in a Scrumban model.
We still use:
- Sprint planning, although relatively loosely
- Daily stand-ups
- Sprint reviews
- Retrospectives
As we experiment with agentic development, we are considering moving away from formal sprint planning towards a more Kanban-style continuous planning model.
The thinking is straightforward.
If implementation capacity becomes more elastic because agents can execute multiple tasks concurrently, artificially batching work into two-week planning cycles may become less useful for some types of work.
Instead:
Product + Engineering maintain a continuously prioritised backlog.
Epic owners refine business requirements with Product.
Business requirements are translated into specifications.
Specifications enter the development flow when capacity is available.
Agents execute the implementation.
Work flows continuously towards production.
That looks more like:
Prioritise → Specify → Validate → Execute → Verify → Deliver
rather than:
Plan sprint → Execute sprint → Review sprint → Repeat
But I would strongly resist presenting this as “Agile is dead” or “Kanban is the future.”
That would miss the point.
Different teams have different products, dependencies, regulatory requirements, release cadences and levels of uncertainty.
Scrum may make perfect sense for one team.
Scrumban may make sense for another.
Kanban may fit another.
XP practices may remain extremely valuable for teams where engineering feedback loops are critical.
The framework should follow the problem.
The important question is not which Agile framework wins.
It is:
Does our way of working still match the economics and constraints of how software is now being produced?
For our team, continuous planning currently appears to fit better with the way we are experimenting with AI-assisted development.
That may not be true for another team.
Not everything about Agile should disappear
Interestingly, some Agile rituals become more important, not less.
Daily stand-ups
We are retaining these.
Not because someone needs to report what they did yesterday.
They provide a daily human alignment point in a world where multiple agents may be working asynchronously.
Humans still need to understand:
- What is happening?
- What is blocked?
- What decisions need to be made?
- Are multiple agents changing the same area?
- Has an assumption changed?
Sprint reviews
We are also retaining these.
AI can accelerate implementation, but it cannot replace business alignment.
Stakeholders still need to see what has changed and determine whether it solves the problem they actually care about.
Retrospectives
Perhaps these become even more important.
When agents dramatically change how work gets done, teams need a safe environment to discuss:
- Where AI helped
- Where AI created problems
- Where humans lost context
- Where specifications were ambiguous
- Where agents generated unnecessary work
- Where the process created stress
- What should change next
The retrospective becomes a feedback loop for the human-agent system, not just the human team.
The metrics need to change too
This may be one of the biggest leadership challenges.
For decades, engineering organisations have measured things such as:
- Story points
- Tickets completed
- Velocity
- Lines of code
- Pull requests
- Deployment frequency
- Lead time
- Change failure rate
AI makes some traditional measures even less useful.
If one engineer can orchestrate five agents, counting tickets completed tells us very little.
Likewise, celebrating lines of code becomes almost meaningless.
We need to measure outcomes and system health, not AI activity.
DORA’s current AI research explicitly cautions against simplistic productivity measures and emphasises capabilities such as strong version control, small batches, quality internal platforms, team performance and software delivery outcomes. (Dora)
I would therefore consider metrics such as:
Delivery
- Lead time from approved specification to production
- Deployment frequency
- Time from business requirement to validated outcome
- Work item age
- WIP
Quality
- Change failure rate
- Escaped defects
- Production incidents
- Rework
- Specification-to-implementation defects
Specification health
- Specification review time
- Number of specification iterations
- Specification churn
- Requirements discovered after implementation started
- Percentage of work with explicit acceptance criteria
AI-specific flow
- Agent execution time
- Agent rework rate
- Human intervention rate
- Agent-generated PR volume
- Review queue time
- Percentage of changes requiring human code review
- Automated validation pass rate
But there is an important warning here.
Do not turn AI usage into a productivity leaderboard.
More tokens, more agents, more PRs or more generated code do not necessarily mean more value.
DORA has explicitly warned about treating raw AI usage or token consumption as a performance indicator. (Dora)
The objective is not to maximise AI activity.
The objective is to maximise valuable software outcomes without compromising quality, security or team health.
A possible future SDLC
Putting all of this together, I can imagine an AI-native SDLC looking something like this:
BUSINESS OUTCOME
|
v
BUSINESS REQUIREMENT
|
v
SPECIFICATION
|
+-------+-------+
| |
v v
Human Review Automated Checks
| |
+-------+-------+
|
v
AI IMPLEMENTATION
|
+-----------+-----------+
| | |
v v v
Test Agent Security Agent Privacy Agent
| | |
+-----------+-----------+
|
v
VALIDATION
|
v
Risk-based Human Review
|
v
DEPLOYMENT
|
v
OBSERVABILITY
|
v
FEEDBACK LOOP
|
+---> Specification
|
+---> Product
|
+---> Retrospective
Notice what has changed.
The centre of gravity has moved.
The engineer is no longer primarily a person who converts tickets into code.
The engineer increasingly becomes someone who:
understands the problem → defines intent → creates constraints → orchestrates agents → evaluates outcomes → manages risk → owns the system.
That is a very different job.
SDD doesn’t replace TDD
One final distinction is important.
I don’t see SDD as the replacement for TDD.
I see them operating at different levels.
SDD answers:
What should the system do?
TDD answers:
Can we demonstrate that the implementation behaves correctly?
The two can work together.
For example:
Specification
Defines the behaviour and constraints.
↓
Acceptance criteria
Defines what successful implementation means.
↓
Tests
Turn those behaviours into executable verification.
↓
AI agent
Generates implementation and tests.
↓
Automated validation
Checks conformance.
↓
Human
Reviews intent, risk and outcome.
In that sense, SDD could be viewed as moving the specification and verification conversation one level higher.
The biggest shift is not technical
The most interesting part of this transition is not Markdown files.
It is not AI coding agents.
It is not even automated code review.
It is the changing relationship between human judgment and machine execution.
For decades, software engineering has largely been constrained by how quickly humans can translate intent into implementation.
Agentic AI changes that equation.
Implementation can now scale much faster than human review capacity.
That means our engineering systems need a stronger mechanism for preserving intent.
Specifications can provide that mechanism.
They give humans a place to reason about the problem before an agent runs away with an interpretation.
They give agents explicit boundaries.
They give tests something concrete to validate.
They give reviewers a higher-level artifact to inspect.
And they give organisations a persistent record of why the system is supposed to behave the way it does.
From writing code to governing change
Perhaps the simplest way I can describe the evolution is:
Waterfall
Plan the work.
Agile
Deliver and learn continuously.
TDD
Define behaviour through tests before implementation.
AI-assisted development
Let machines accelerate implementation.
Agentic development
Let machines execute increasingly autonomous engineering tasks.
Spec-Driven Development
Define intent, constraints and boundaries clearly enough that machines can execute against them safely.
The end state is not a world where engineers stop reviewing software.
It is a world where engineers spend less of their scarce attention reviewing every implementation detail and more of it reviewing intent, constraints, risk and outcomes.
That feels like a much more scalable model for an engineering organisation where the ability to generate software is no longer the primary constraint.
The question for engineering leaders is therefore no longer simply:
“How do we get our engineers to write code faster?”
It increasingly becomes:
“How do we build an engineering system where AI can move fast without moving faster than our ability to understand and govern what it is changing?”
For our team, SDD is one experiment towards answering that question.
And I suspect we are only at the beginning.







Leave a Reply