Human Intervention in Agentic Work: A Measurement Scheme, Pilot Evidence, and a Research Program

When, if ever, is human intervention required for an AI agent to operate successfully? The question is empirical and the field cannot currently answer it, because the available record registers every correction as one undifferentiated event, at a scale already measured: in the largest published analysis of agentic coding, 91.49% of failure resolutions required explicit user correction against 2.99% agent self-correction. Frequency says nothing about which corrections capability growth can remove. This proposal indexes interventions by epistemic content along two dimensions, what a move supplies and what it operates on. Grounds are absorbable by capability growth, at a rate set by the domain's verifier availability rather than by capability alone. Frames yield to a second reasoner who need not be human, subject to a supply condition: the absorber is a population fact, so monoculture closes the priors route while leaving procedural and representational sources open, and because no frame can be requested by name, their supply turns on a decision to invoke that the executor does not own. Standing cannot be self-conferred, so no capability increment absorbs it. The policy-facing corollary is a compositional prediction: the grounds-shaped fraction of human input declines as capability grows while the standing-shaped fraction does not, so task-level exposure indices can report stability across exactly the period in which the character of human work inverts.

Pilot evidence from one practitioner's seven-month record establishes the scheme's usability and not its findings, and is reported with its failures. Blind rater pairs drawn from different model tiers agree at 0.77 on whether an input is an intervention at all and 0.70 on its family; on ten sessions held out from the scheme's development the gate replicates and slightly improves at 0.82 while family agreement falls to 0.58. One compositional result survives every protocol and replicates out of sample: deliverable-directed correction is the largest family, which measures a blind spot in the field's event ontology rather than a behaviour. A crosswalk to the field's flagship measurement motivates the program. That instrument partitions conversations by division of labour, which is orthogonal to content, and its automation share tracks conversational volume, so as grounds absorb, conversations shed the messages that made them look augmentative while the residual threshold-setting registers as near-silence. An index whose sensitivity to a class runs inversely to that class's verbosity is structurally least able to detect the class this argument says does not absorb. The corresponding longitudinal test was designed, gated, and withdrawn rather than run, on three confounds including an undisclosed classifier that is itself a model improving over the same window in which the effect of improving models is the quantity of interest.

PDF · DOI: 10.5281/zenodo.21719008 · all versions

The full treatment is deposited as three papers, divided along the argument's dependency graph. The mechanism inventory, the discovery corpus, the coding scheme and its reliability measurements, and the literature survey are in An Epistemic-Content Taxonomy of Human Intervention in Agentic Collaboration. The derivation of which classes of intervention are absorbable and which is not is in Grounds, Frames, and Standing. The instruments built to locate one practitioner's contribution on that ladder, and the controls that destroyed most of them, are in What Can Be Established About Human Irreducibility. All three are cited throughout as the technical companions.

← all papers