Campus AI Development Agentic coding

The continuum Band 03

03

Agentic coding

The agent plans, edits, runs tests and iterates. The human sets goals and reviews.

Also called: Agentic engineering, coding agents

01Definition

A coding agent takes a scoped task, reads the codebase, makes multi-file changes, runs commands and tests, and iterates until it believes the task is done. The human sets the goal, supplies context and conventions, approves risky actions and reviews the diff before it merges.

02Origins

Karpathy later preferred "agentic engineering" as the framing for serious work, as distinct from vibe coding.

Agents that could edit code and run tests appeared in research first, including Princeton’s SWE-agent and Cognition’s Devin in 2024. Terminal-based agents such as Claude Code, OpenAI Codex and Gemini CLI followed in 2025 and brought the approach into everyday use.

Karpathy proposed “agentic engineering” in February 2026 and set out the idea at Sequoia’s AI Ascent in April. Others, including Simon Willison and Addy Osmani, had used the phrase earlier.

03How it works

Person AI
01Scope the task and context
02Plan, edit, run tests
03Approve risky actions
04Iterate until tests pass
05Review the diff, merge

Steps 2–4 repeat inside one task.

  1. 01Write a scoped task with context: the issue, the goal, and a conventions file in the repository (such as CLAUDE.md or AGENTS.md).
  2. 02The agent plans, reads the code, edits files and runs commands and tests.
  3. 03Permission rules decide which actions need human approval.
  4. 04The agent iterates until tests pass, or reports what it couldn’t do.
  5. 05A human reviews the diff and merges.

Typical tools: Claude Code, OpenAI Codex, Gemini CLI, Cursor agent mode, GitHub Copilot coding agent.

04Profile across the dimensions

How the work is done

Human role
Sets goals, reviews diffs
Unit of work
A task handed to an agent
Source of truth
The issue or ticket
Agent autonomy
Acts, with approval at review
Lifecycle coverage
Build, test, debug; sometimes task planning

How you know it’s right

Code read by a person
The diff, at review
Verification
Agent-run tests plus human review
Traceability
Linked issues and pull requests
Delivery automation needed
Automated tests and protected branches

Fit and risk

Upfront investment
Moderate
Durability
Maintainable with review
Data and risk ceiling
Internal data, scoped permissions

Compare all five bands

05Lifecycle coverage

Plan
Design
Build
Test
Deploy
Operate

Solid: AI does core work · Hatched: partial or informal · Dashed: done by people. Compare all bands

06When to use it

Good for: Multi-step internal features with a mostly clear goal

One personWorks well with tests and disciplined review of your own diffs. Reviewing your own work is the weak point.
A teamNeeds a shared conventions file (CLAUDE.md, AGENTS.md), one task per branch and peer review of agent pull requests. Review capacity becomes the bottleneck.

Less suited to

  • Goals you can’t yet state clearly
  • Codebases without tests, where “done” can’t be checked
  • Sensitive environments without sandboxing or permission controls

Signs you’ve outgrown it

  • Diffs are too large to review properly
  • Several people or agents work on the same system
  • Requirements drift from one task to the next

07Development environment

Where it runs: A terminal or editor agent (Claude Code, Codex, Cursor) working in a repository with tests.

Isolation
A container, cloud dev environment or separate branch per task
Permissions
Allow, ask and deny rules for commands; no production credentials; limited network access
Audit
Logs of agent actions
On campus
HPC clusters with agents as installed software (Utah, USC); cloud dev environments run by central IT

Compare environments across bands

08Minimum controls

  • A human reviews every diff before merge
  • Version control and automated tests
  • A separate development environment with scoped credentials

Risks

  • Destructive actions: deleting files or running commands on shared systems.
  • Prompt injection: instructions hidden in files, issues or web pages the agent reads.
  • Credential exposure: agents reading or echoing secrets.
  • Review bottleneck: agents produce changes faster than people can review them.

09On campus

10Sources

  1. Yang et al., “SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering,” 2024
  2. Anthropic, “Claude Code: Best practices for agentic coding,” 2025
  3. Karpathy on “agentic engineering,” Feb. 2026, and Sequoia AI Ascent talk, Apr. 2026
  4. USC CARC, “AI Coding Agents” user guide, 2026