02 — The continuum
This guide arranges them on one continuum of structure and governance rather than treating them as rivals. Autonomy and code reading are separate dials, as the two-axis map shows, and band 05 bundles two things that can come apart: a single agent can be governed, and several agents can run with no governance at all. The bands describe typical practice, not hard boundaries. As of October 2026 the left end is moving fastest: tools once used for chat-style vibe coding now edit files and run commands, so expect some cells in the table below to shift. Moving right adds structure and durability but costs more effort up front. Moving left trades durability for speed.
Chat → run → prompt again. You read little or none of the code.
Good for: Discovery, learning, demos, throwaway tools
Prompt-driven development is the more structured version: the human still breaks the work into a sequence of prompts.
The human writes; AI suggests. The human remains the primary author.
Good for: Production work in an existing codebase; coursework with learning goals
The agent plans, edits, runs tests and iterates. The human sets goals and reviews.
Good for: Multi-step internal features with a mostly clear goal
Karpathy later preferred "agentic engineering" as the framing for serious work.
The spec is the source of truth. Code, tests and docs are derived from it.
Good for: Anything others will rely on or that must be maintained
See Emerson College: domain experts who don’t read the code, but who work from a written doctrine and a data model designed up front.
Specialized agents (planner, coder, tester, reviewer) under orchestration, with human gates, audit trails, living specs and compliance checks.
Good for: Regulated, long-lived, multi-team systems
This is what earlier versions of this guide called the AI-native SDLC.
The same questions, asked of every approach. Two of them can vary independently: Emerson leaves the code unread but still verifies the work.
| Dimension | 01Vibe coding | 02AI-assisted | 03Agentic | 04Spec-driven | 05Multi-agent, governed |
|---|---|---|---|---|---|
| How the work is done | |||||
| Human role | Describes and reacts | Author; AI suggests | Sets goals, reviews diffs | Writes and owns the spec | Designs the system, owns the gates |
| Unit of work | A conversation | A commit | A task handed to an agent | A spec change | A governed pipeline run |
| Source of truth | The chat history | The code | The issue or ticket | The spec | Spec plus policy |
| Agent autonomy | Varies: often run in agentic tools that edit files and run commands; what defines it is not reading the output | Suggests only | Acts, with approval at review | Acts within the spec | Acts within guardrails and gates |
| Lifecycle coverage | Build only | Build and test | Build, test, debug; sometimes task planning | Design and requirements first; build and test follow | Plan through operate, with audit |
| How you know it’s right | |||||
| Code read by a person | Little or none | Every line | The diff, at review | Checked against the spec, not line by line | At gates, according to risk |
| Verification | Run it and look | Code review and tests | Agent-run tests plus human review | Acceptance criteria, executable specs | Automated evaluations, security scans, human gates |
| Traceability | None | Commit history | Linked issues and pull requests | Each requirement linked to its code | Full audit trail |
| Delivery automation needed | None | Version control | Automated tests and protected branches | Automated checks against the spec; deployment pipeline | Policy enforced in the pipeline; monitoring and rollback |
| Development environment | Personal sandbox or browser builder | Normal editor with an enterprise assistant | Isolated container or branch, scoped permissions, action logs | Shared institutional platform with dev/prod separation | Full pipeline with gates, staging and agent service accounts |
| Fit and risk | |||||
| Upfront investment | Minutes | Low | Moderate | High | Highest |
| Durability | Throwaway | Maintainable | Maintainable with review | Long-lived | Long-lived, regulated |
| Data and risk ceiling | Public or sandbox data only | Internal data, under existing review rules | Internal data, scoped permissions | Institutional data with IT support | Restricted data, under audit |
| People and cost | |||||
| Skills it expects | Domain knowledge; judging whether the result works | Reading and owning code line by line | Reading diffs, scoping permissions, writing tests | Writing and maintaining living specs | System design, agent orchestration, audit |
| Cost to watch | Low spend; tools orphaned when the builder moves on | Seat licenses; review time | Token spend from long agent runs | Spec upkeep; slower start | Highest token and oversight cost; platform upkeep |
| Cost of over-caution | Experiments go to personal accounts | Developers drop approved tools | Work stalls waiting for review | Specs written for throwaway work | Small tools stuck in a heavy pipeline |
Name the level of autonomy explicitly. Teams that do can decide in advance how much control to keep, and policies can say which levels are allowed for which kinds of data.
How much of the code a person reads, and how much structure surrounds the work, aren’t the same question. The bands move along both, but not in lockstep, and real projects can sit anywhere.
Positions are approximate. Band 05 reads code at gates according to risk, so it sits mid-height. The orange marker is Emerson: domain experts who don’t read the code, inside a governed environment.