All field guides

Production field guide

What must be true before an AI system ships.

A decision-oriented AI production-readiness checklist covering representative evidence, release authority, rollback, incident response, monitoring, and change control.

Definition

Production readiness is a release decision: the evidence, authority, rollback path, and monitoring required for one consequential AI workflow to operate safely.

01

Define the release surface

Release readiness applies to a specific version, workflow, population, environment, and operating owner—not to an AI initiative in the abstract.

Consequential workflow

Name the user, task, affected system, data, tool permissions, business consequence, and boundary between advice, draft, and autonomous action.

Unacceptable failures

List the errors, misuse, security events, policy violations, and silent quality failures that must block release or require review.

Versioned candidate

Pin the model, prompt, retrieval corpus, tools, policies, thresholds, and environment so the tested system matches the candidate being approved.

02

Build pre-release evidence

A benchmark score becomes a gate only when it represents deployment conditions and has an explicit consequence.

Representative evaluations

Cover ordinary, edge, adversarial, degraded, and recovery cases drawn from the actual workflow. Track baseline, acceptance threshold, severity, and unresolved failures.

Security and permissions

Verify identity, authorization, tool and data scope, isolation, logging, secret handling, abuse cases, and the blast radius of an incorrect action.

Human review load

Measure how often review is required, what reviewers need to decide, how disagreement is handled, and whether the operating model can sustain the burden.

Explore the public proof standard
03

Make release authority explicit

A system is not controlled merely because several stakeholders can object. Someone must own approval and the paths around it.

Approve, condition, or block

Name the release owner and the evidence they require. Record conditions, open actions, expiry dates, and who may accept an exception.

Containment and rollback

Exercise disablement, provider or model fallback, access removal, output quarantine, rollback, and communications before the first consequential incident.

Incident response

Define detection, severity, escalation, evidence preservation, user impact assessment, recovery, and the authority to stop the workflow.

04

Launch the operating evidence loop

Approval is the start of a controlled operating window, not the end of evaluation.

Monitoring

Track quality, safety, drift, cost, latency, review, override, incidents, provider changes, and business outcome at the level of the released workflow.

Change control

Define which model, prompt, data, tool, policy, threshold, or integration changes require reevaluation or a new approval.

First-window review

Set the first review date, evidence owner, thresholds for intervention, and the decision to continue, condition, expand, or roll back.

Sources

Primary references and further reading.

Apply the framework

The decision is live. Build the evidence.

Bring the owner, deadline, current evidence, and the question capable of changing the next move.

Build and deploy one AI workflow