Human-in-the-Loop Maturity
Building the gate is the easy part. Articles 4 through 6 cover when gates trigger, what reviewers see and how decisions are logged. Running the program is harder. Reviewer quality drifts. Volumes fluctuate. Backlogs create pressure. And no architecture eliminates the variability that comes with having humans in the loop. That’s the whole point.
Seven dimensions separate a mature HITL program from a well-architected one. Gate conditions need periodic calibration against actual incident data, not just configuration at launch. Reviewer qualification must match the gate type, with domain expertise requirements documented and rosters maintained. Inter-reviewer agreement is tracked via Cohen’s kappa over rolling 30-day windows; target is above 0.7, with quarterly reviews and calibration sessions when disagreement is significant. Downstream quality matters too: when incidents or confabulation findings trace back to HITL-approved outputs, flag those decisions for analysis. The goal isn’t accountability for individual reviewers. It’s catching calibration problems early, before they compound. Throughput limits per reviewer, bias monitoring across decision patterns and a named program lead reporting quarterly to the CAIO round out the requirements.
Set a maximum queue depth per gate type and activate backup reviewers when it’s exceeded. A queue running consistently near its limit is a capacity problem, not an operations problem. It belongs in front of the CAIO with a capacity analysis attached. And gate scope should never expand without corresponding reviewer capacity to match.
Clinical deployments and regulated-professional contexts need more than the seven dimensions above. Reviewers document an independent clinical assessment before seeing the AI recommendation, so the output doesn’t anchor their judgment before they’ve formed one. Every override or rejection gets a brief structured rationale, a record of clinical judgment that lives separately from the AI system record. At least quarterly, reviewers work a sample of cases without AI assistance so the program can measure whether AI-assisted review is improving or eroding independent judgment. Approval rate trends are monitored longitudinally; any reviewer whose rate climbs more than 15 percentage points quarter-over-quarter gets a calibration review, because rising rates in clinical deployments signal habituation, not competence. Reviewer performance is evaluated against clinical outcomes, not consistency with AI recommendations. Annual attestation confirms that licensed reviewers still retain the professional competency to work without AI assistance. The risk tolerance statement (Article 3, Artifact 1) should name this directly: permanent human judgment, not a trajectory toward greater autonomy.