An Epistemic-Content Taxonomy of Human Intervention in Agentic Collaboration

Under a two-axis coding of one practitioner's full record of 2,205 human inputs to LLM coding agents, the largest single family of interventions is deliverable-directed, operating on the produced artifact rather than on the agent's live reasoning, a result invariant across rater protocols and one that conversational event ontologies are structured to miss. The records behind every large-scale published taxonomy of this traffic register each correction as one undifferentiated event. In the largest published analysis of agentic coding, 91.49% of failure resolutions required explicit user correction, and the finest-grained published taxonomy of that traffic splits it into correction, rejection, and failure report, which is three degrees of disagreement and zero degrees of content. Existing schemes index interventions by timing, authority, granularity, defect class, or interaction state, but none indexes what the input supplies.

From seven months of documented practice, this paper derives nine recurring intervention mechanisms indexed by epistemic content, with a lower-evidence second tier of four, thirteen in all. Three of the nine inject zero domain content, moving only commitment, salience, or attention. Twelve of the thirteen already carry established names in the literatures surveyed, and the recommendation is to adopt those names rather than coin new ones. What is left as new is bounded by the survey that looked for it: the survey is access-filtered and searched on this paper's own coinages, legal cross-examination is a probable prior home for one mechanism, and every novelty claim is therefore scoped to the literatures surveyed.

Re-coding the corpus's full 2,205-input record forced the scheme into two axes: alongside its content class, every intervention carries a target, whether the live reasoning, the deliverable, the rule store, the communication channel, or the agent's self-model. Under that scheme the classes descended from the unmapped residue carry 555 of the 1,094 coded interventions against 407 for the original thirteen, which is not a like-for-like revision of the 14 percent that prompted the re-code but does place a substantial share of the record outside the original inventory.

The measurement that matters for a taxonomy is whether raters other than its author can apply it. Blind rater pairs drawn from different model tiers reach a kappa of 0.77 on whether an input is an intervention at all and 0.70 on its content family once labeling conventions are pinned, against a same-family upper bound of 0.86. The family figure is the third of three passes over one corpus, and the conventions that lifted it were derived from the first pass's disagreements, so the runs are a sequence rather than a design. On sessions held out from the scheme's development the gate replicates at 0.82 while family agreement falls to 0.58. Those two figures rest on different samples, the gate computed over all 93 held-out items, the family figure over the 19 items both raters called an intervention, so the family axis is the weaker measurement. Neither figure bounds the composition shares, which come from a later, author-adjudicated instrument that carries no reliability figure of its own.

The corpus is one practitioner's records, a discovery instrument rather than evidence, and validation runs through re-coding public interaction corpora, which requires no private data. What the inventory entails about durability and absorption is developed in a companion paper.

PDF · DOI: 10.5281/zenodo.21730478 · all versions

Claims are scoped to the literatures surveyed. The discovery corpus is one practitioner's records and is treated as a discovery instrument rather than as evidence, with validation to come from re-coding public interaction corpora.

This paper and its companion were one document through v2.11, and share a mechanism list and nothing else. This half establishes that the inventory exists, that it can be applied by someone other than its author, and what its measurement bounds are. It does not argue durability.

← all papers