Campus AI Development What the evidence says

10 — What the evidence says

Faster on a narrow task, mixed on real work, risky for learning

The studies disagree mostly because they measure different things: one well-defined task, a mature codebase, or how well novices learn.

Developer confidence Less secure

Participants with an AI assistant wrote less secure code than those without, and were more likely to believe it was secure.

Perry et al., ACM CCS 2023 [cite]

How to read this: speed-ups on a single bounded task in a lab don't carry over directly to large, unfamiliar or legacy codebases, which describes much campus software.

For teaching: results are mixed. In one study of learners aged 10–17, those with code generators did not lose ground on later manual tests; in another, students with weaker metacognitive skills came away with an illusion of competence. College-level evidence is still thin, so tools that scaffold learning rather than hand out answers are the safer choice.

What the evidence doesn’t cover yet: nearly all controlled studies look at code suggestion and chat-style help (bands 01–02). There is little trial data on spec-driven or multi-agent work. This guide’s case for more structure at higher stakes rests on engineering practice and compliance logic, not experimental results.