Campus AI Development Multi-agent development

The continuum Band 05

05

Multi-agent development

Specialized agents under orchestration, with human gates, audit trails and compliance checks.

Also called: AI-native SDLC (this guide’s earlier name)

01Definition

Specialized agents (planner, coder, tester, reviewer) work under orchestration across the lifecycle. Humans design the system and own the gates; living specs, automated evaluations, security scans and audit trails make the work verifiable and accountable.

02Origins

Research prototypes in 2023, notably ChatDev and MetaGPT, gave separate agents roles such as product manager, architect, programmer and tester, and had them pass work between each other. Vendors and consultancies later used “AI-native SDLC” for the enterprise version.

The governance half borrows from existing frameworks: DevSecOps pipelines, NIST’s Secure Software Development Framework and the NIST AI Risk Management Framework. Of the five terms, this is the least settled.

03How it works

Person AI
01Set specs, policies, gates
02Orchestrator assigns
03Agents build, test, review
04Approve at gates
05Deploy, monitor, log
  1. 01People define specs, policies and the gates that need human approval.
  2. 02An orchestrator assigns tasks to specialized agents: planner, coder, tester, reviewer, security.
  3. 03Agents check each other’s work; automated evaluations and scans run in the pipeline.
  4. 04Humans approve at gates, set by risk.
  5. 05Deployment, monitoring and rollback are automated, and every action is logged.

Typical tools: agent orchestration frameworks; deployment pipelines with policy-as-code; security scanning; audit logging for agent actions.

04Profile across the dimensions

How the work is done

Human role
Designs the system, owns the gates
Unit of work
A governed pipeline run
Source of truth
Spec plus policy
Agent autonomy
Acts within guardrails and gates
Lifecycle coverage
Plan through operate, with audit

How you know it’s right

Code read by a person
At gates, according to risk
Verification
Automated evaluations, security scans, human gates
Traceability
Full audit trail
Delivery automation needed
Policy enforced in the pipeline; monitoring and rollback

Fit and risk

Upfront investment
Highest
Durability
Long-lived, regulated
Data and risk ceiling
Restricted data, under audit

Compare all five bands

05Lifecycle coverage

Plan
Design
Build
Test
Deploy
Operate

Solid: AI does core work · Hatched: partial or informal · Dashed: done by people. Compare all bands

06When to use it

Good for: Regulated, long-lived, multi-team systems

One personRarely worth it. The overhead outweighs the risk for one person.
A teamBuilt for teams. Gates, audit trails and named owners exist because many people and agents touch production.

Less suited to

  • Small teams without automated testing and deployment
  • Prototypes, or work where governance costs more than the risk it manages

Signs you aren’t ready

  • No automated tests
  • Manual deployments
  • No clear owner for production systems

07Development environment

Where it runs: Orchestration on top of a full deployment pipeline.

Gates
Policy checks, security scans and required human approvals built into the pipeline
Separation
Development, staging and production, with infrastructure defined in code
Observability
Monitoring, rollback and an audit trail for every agent action
Identity
Each agent has its own scoped service account, never a person’s credentials

Compare environments across bands

08Minimum controls

  • A short written spec: requirements, acceptance criteria, non-goals
  • Traceability from spec to tests to code
  • Security, accessibility and equity review before merge
  • A named owner, separate production and audit logs

Risks

  • Shared blind spots: AI reviewing AI can miss the same errors.
  • Diffused accountability: unclear who answers for a failure.
  • Complexity and cost: more moving parts to secure and pay for.
  • Governance lag: review and policy can’t keep up with the pace of change.

09On campus

No published campus case yet runs agents across the whole lifecycle with measured outcomes. See the case studies for the closest examples.

10Sources

  1. Qian et al., “ChatDev: Communicative Agents for Software Development,” 2023
  2. Hong et al., “MetaGPT: Meta Programming for a Multi-Agent Collaborative Framework,” 2023
  3. NIST, Secure Software Development Framework (SP 800-218), 2022
  4. NIST, AI Risk Management Framework 1.0, 2023
  5. DORA, Accelerate State of DevOps Report, Google Cloud, 2024