Define the release surface
Release readiness applies to a specific version, workflow, population, environment, and operating owner—not to an AI initiative in the abstract.
Consequential workflow
Name the user, task, affected system, data, tool permissions, business consequence, and boundary between advice, draft, and autonomous action.
Unacceptable failures
List the errors, misuse, security events, policy violations, and silent quality failures that must block release or require review.
Versioned candidate
Pin the model, prompt, retrieval corpus, tools, policies, thresholds, and environment so the tested system matches the candidate being approved.
Build pre-release evidence
A benchmark score becomes a gate only when it represents deployment conditions and has an explicit consequence.
Representative evaluations
Cover ordinary, edge, adversarial, degraded, and recovery cases drawn from the actual workflow. Track baseline, acceptance threshold, severity, and unresolved failures.
Security and permissions
Verify identity, authorization, tool and data scope, isolation, logging, secret handling, abuse cases, and the blast radius of an incorrect action.
Human review load
Measure how often review is required, what reviewers need to decide, how disagreement is handled, and whether the operating model can sustain the burden.
Launch the operating evidence loop
Approval is the start of a controlled operating window, not the end of evaluation.
Monitoring
Track quality, safety, drift, cost, latency, review, override, incidents, provider changes, and business outcome at the level of the released workflow.
Change control
Define which model, prompt, data, tool, policy, threshold, or integration changes require reevaluation or a new approval.
First-window review
Set the first review date, evidence owner, thresholds for intervention, and the decision to continue, condition, expand, or roll back.