[DSRP Evidence](https://dsrpevidence.org/)

# When Human Review Is Treated as Ground Truth: A Socio-Technical Analysis of Reference Dependence and Measurement Validity in Human-in-the-Loop AI Inspection Systems

## Details

**Authors** Moon-Sup Lee, Seung-Yeon Han

**Year** 2026

**Publisher** Systems

**Kind of work** article

**Discipline** Systems Science

**Secondary disciplines** Information Science, Computer Science & AI

[Read it at the publisher](https://doi.org/10.3390/systems14091126) 
10.3390/systems14091126

## In authors' words

### Abstract

Human-in-the-loop pipelines are widely used in artificial intelligence (AI)-based inspection systems, coupling an automated model, human reviewers, and an organizational reporting layer that drives decisions and, potentially, model retraining. The closed loop this creates is a design feature, not a defect; the pathology lies one level up, where a record produced by editing the model's own output is treated as an independent reference and the automated-versus-reviewed difference is reported as "accuracy"—a closed loop read as open. This paper offers no alternative estimator; it shows that none is available here and that the infeasibility is itself the diagnosis. We examine the loop in a national pavement-management program, matching automated and reviewed outputs for 349 images across 162,520 individual grids. Reviewers accepted 89.06% of grids without change, added 10.25%, removed 0.28%, and reclassified 0.42% (an addition-to-removal ratio of 36.4:1); patching accounted for 64.5% of the added grids (39.8–69.0% across leave-one-route-out folds), added grids largely recording maintenance patches, not cracks. Consistent with the documented review workflow, 24 images were identical even by grid type and 98.06% of automated distress grids survived review (97.6–99.6% across folds); the reviewed record therefore cannot serve as an independent reference, and detection accuracy cannot be estimated. Across 40 route segments, the dependence mechanism and the asymmetry's direction are invariant; the composition and magnitude are not. We reinterpret the reported figure and derive systems-level interventions—reporting items, interface priorities, and independent-reference requirements.

### What they set out to do (purpose)

To determine whether the "accuracy" figures reported for a human-in-the-loop AI pavement-inspection system actually measure model performance against an independent reference, or are an artifact of a closed feedback loop.

### Who or what was studied (sample)

A national (Korean) pavement-management program's automated vs. human-reviewed inspection outputs for 349 road images comprising 162,520 individual grid cells, across 40 route segments.

### How they did it (methods)

Observational socio-technical audit matching automated model outputs to human-reviewed (edited) outputs grid-by-grid, quantifying acceptance/addition/removal/reclassification rates and testing robustness via leave-one-route-out cross-validation folds.

### What they found (results)

98.06% of the automated model's flagged distress grids survived human review unchanged (97.6-99.6% across folds), with an addition-to-removal edit ratio of 36.4:1, showing the "reviewed" record is not independent of the model's own output and that reported accuracy figures for this system are therefore not a valid measurement of detection performance.

## Commentary

### In short

The study shows that a measurement's validity depends on whose perspective is treated as the independent reference point, and that when the reviewer's perspective is causally entangled with the system it is meant to judge within a closed-loop structure, the resulting measurement is invalidated.

**Patterns it shows** S, R, P

**Added** 2026-09-16

**How to cite this** Moon-Sup Lee, Seung-Yeon Han (2026). When Human Review Is Treated as Ground Truth: A Socio-Technical Analysis of Reference Dependence and Measurement Validity in Human-in-the-Loop AI Inspection Systems. Systems.
