10 — What the evidence says
Faster on a narrow task, mixed on real work, risky for learning
The studies disagree mostly because they measure different things: one well-defined task, a mature codebase, or how well novices learn.
Controlled task
55.8%
faster completion of a bounded JavaScript HTTP-server task by developers using GitHub Copilot, compared with a control group.
Peng et al., 2023 — randomized controlled experiment, 95 developers [cite]
Mature codebases
19%
slower: experienced open-source maintainers took longer on real issues in their own repositories when allowed AI tools, though they believed they had been faster.
METR, 2025 — randomized trial, 16 developers, 246 tasks [cite]
Delivery at scale
−7.2%
estimated drop in delivery stability for every 25% increase in AI adoption, even as respondents reported higher individual productivity.
Google DORA, Accelerate State of DevOps 2024 [cite]
Novice programmers
Mixed
Students with code generators finished more tasks without losing performance on later manual tests; students with more prior knowledge gained the most.
Kazemitabaar et al., CHI 2023 — 69 novices, ages 10–17 [cite]
Metacognition
Gap
AI help sped up students who were already strong, while students with metacognitive difficulties came away with an illusion of competence.
Prather et al., ICER 2024 — "The Widening Gap" [cite]
Scaffolded tutors
At scale
A guardrailed course assistant that gives hints instead of solutions was deployed to CS50 students, with high reported usefulness.
Liu et al., SIGCSE 2024 — "Teaching CS50 with AI" [cite]
Security of generated code
~40%
of programs Copilot generated for security-relevant scenarios contained vulnerabilities.
Pearce et al., IEEE S&P 2022 — "Asleep at the Keyboard?" [cite]
Developer confidence
Less secure
Participants with an AI assistant wrote less secure code than those without, and were more likely to believe it was secure.
Perry et al., ACM CCS 2023 [cite]
Package hallucination
19.7%
of packages referenced across 576,000 generated code samples did not exist, an opening for supply-chain attacks.
Spracklen et al., USENIX Security 2025 [cite]
How to read this: speed-ups on a single bounded task in a lab don't carry over directly to large, unfamiliar or legacy codebases, which describes much campus software.
For teaching: results are mixed. In one study of learners aged 10–17, those with code generators did not lose ground on later manual tests; in another, students with weaker metacognitive skills came away with an illusion of competence. College-level evidence is still thin, so tools that scaffold learning rather than hand out answers are the safer choice.
What the evidence doesn’t cover yet: nearly all controlled studies look at code suggestion and chat-style help (bands 01–02). There is little trial data on spec-driven or multi-agent work. This guide’s case for more structure at higher stakes rests on engineering practice and compliance logic, not experimental results.