Campus AI Development Use case scenarios

Use case scenarios 02 of 08

Scenario 02

Research software engineer

A postdoc vibe-coded an analysis that now underpins a paper. A research software engineer rebuilds it as a pipeline that other groups will run on the cluster and cite.

Most governedLeast governed

Questions to consider

Six lenses from the scenario framework. Any answer that raises the stakes moves this scenario toward stricter controls; take your answers to the chooser.

01 Data What it touches and how that data is classified
  1. Does any input fall under IRB approval, export control or a data use agreement?
  2. Which model environment is approved for that data?
02 Audience Who relies on it, and what happens if it’s wrong
  1. Who else will run or cite the pipeline?
  2. Will results feed a publication or a policy decision?
03 Lifespan How long it lives and who maintains it
  1. Who maintains it after the grant ends?
  2. Could someone rebuild the environment in five years?
04 Reach What it can read, write or break
  1. What shared cluster resources can the agent touch?
  2. Can it submit jobs or delete files outside the project?
05 Verification How you’ll know it’s right, and keep knowing
  1. Do the tests reproduce the hand-checked results?
  2. How will you notice a change when the model or a dependency updates?
06 Accountability Who owns it, approves it and discloses it
  1. Who is credited as the software author, and how is AI help disclosed?
  2. Does the license allow release?

How the work goes

  1. 01Read the prototypeRun it, capture its outputs and hand-check a small subset.
  2. 02Rebuild with an agentWork on the cluster under the center’s rules file and permission policy.
  3. 03Test and pinWrite tests that reproduce the hand-checked results; pin environments and dependencies.
  4. 04Record provenanceNote models, versions and prompts in the repository and the methods section.
  5. 05ReleaseCheck licenses, add a citation file and publish.

Proportionate controls

T4 · Institutional

Governance should match the risk: enough to protect people and data, no more. Each control area keeps its own tier on the five-tier scale, and those are the controls to apply. The baseline is a label for the project as a whole, taken from its highest area.

  1. Review T4 Human review of every change by someone who can explain it
  2. Documentation T4 A written spec with acceptance criteria and non-goals
  3. Approval T3 The system or data owner approves
  4. Data T4 Confidential records at scale (Level 3, such as FERPA and GLBA data) with steward approval and minimization
  5. Access T4 Scoped service identities; agents never hold production credentials
  6. Testing T4 Automated tests and security scanning on every change
  7. Monitoring T3 Platform logs and periodic usage review
Raises the tier
  • Human-subjects, health or export-controlled data enters the pipeline
  • Results inform clinical, policy or funding decisions
Legal triggers that raise it →
Over-governing looks like
  • Sending analysis notebooks through a change board
  • Banning agents on the cluster outright instead of scoping them
Excess controls push builders toward unsanctioned tools.
Minimum controls
  • Version control and automated tests
  • Pinned environments; reproducible from a clean start
  • Provenance in the repository and methods section
  • Sensitive data only in approved enclaves, with enterprise or locally hosted models
Watch for
  • Agents acting on shared cluster resources
  • License questions before open-source release
  • Researchers losing touch with their own code
When it moves up

When regulated data (IRB-governed, export-controlled) enters, or another institution depends on the pipeline, add a written spec and independent review.