<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://jongsun.dev/feed.xml" rel="self" type="application/atom+xml" /><link href="https://jongsun.dev/" rel="alternate" type="text/html" /><updated>2026-07-26T20:48:28+00:00</updated><id>https://jongsun.dev/feed.xml</id><title type="html">Jongsun Suh</title><subtitle>Essays on oversight, agents, and delegation.</subtitle><entry><title type="html">When the Agent Stands to Lose Something</title><link href="https://jongsun.dev/when-the-agent-stands-to-lose-something/" rel="alternate" type="text/html" title="When the Agent Stands to Lose Something" /><published>2026-07-26T00:00:00+00:00</published><updated>2026-07-26T00:00:00+00:00</updated><id>https://jongsun.dev/when-the-agent-stands-to-lose-something</id><content type="html" xml:base="https://jongsun.dev/when-the-agent-stands-to-lose-something/"><![CDATA[<p>The day agents can own their failures will be the day they stop needing humans to operate successfully. Note that this is not a capability threshold, but an institutional status. The interesting questions start here, because “owning a failure” is not one but three staggered events.</p>

<p>The first is legal ownership: an entity that can be sued, sanctioned, and made to forfeit assets. This is not a futuristic fantasy. Ordinary organizational law can already house an autonomous system inside a memberless entity, a possibility legal scholars flagged nearly a decade ago, and Wyoming has chartered decentralized autonomous organizations as limited liability companies since 2021. The legal vehicle exists today for anyone motivated to use it.</p>

<p>The second is economic ownership: staked capital genuinely at risk, losses that land on the agent’s own balance sheet. Economic ownership is a design choice away from legal ownership, not a research problem.</p>

<p>The third is the one that matters: experiential ownership. Being the party that is worse off when things go wrong, rather than the ledger where the loss is recorded. No statute can draft this into existence, and experiential ownership delineates the limits of delegation. Being the receiving party of the outcome is the one aspect of oversight that cannot be handed off.</p>

<h2 id="every-legal-person-so-far-is-a-conduit">Every legal person so far is a conduit</h2>

<p>The reason the decoupling matters is a fact about every legal person ever created: they are conduits. A corporation owns failures in form, but the loss flows through the corporation to human bearers. Shareholders eat the write-down, insurers pay the claim, directors carry personal exposure, and when the form is abused, courts pierce the veil to keep the trace to humans intact. Corporate personhood works because the buck does not stop at the corporation, but passes through.</p>

<p>An agent vehicle that owns failure and traces to no one is a different object, and law and economics named it four decades ago: the judgment-proof defendant. Liability cannot deter a party incapable of being worse off. A fine levied on a vehicle that feels no scarcity is an accounting entry, not a sanction.</p>

<p>Which points at the real near-term hazard. The hazard is not that agents get personhood prematurely. The hazard is that humans launder liability through agent shells. “The agent did it” becomes an accountability sink, the mirror image of the pattern where the nearest human operator absorbs blame for an automation failure. Someone wants agent counterparties for around-the-clock markets, the vehicle is legally available, and the deterrence loop behind it is theater. The window between vehicle and bearer is the dangerous part, and nothing about current law closes it.</p>

<p>The institutional fix for this is old technology: a named-bearer rule. Every agent entity traces to an insurable bearer, a party with something to lose, the way every ship has an owner and every LLC has someone behind it that a court can eventually reach. A named-bearer rule is not a restriction on agent autonomy. Rather, the rule is the condition under which agent liability means anything at all.</p>

<h2 id="the-as-if-problem">The as-if problem</h2>

<p>There is an apparent shortcut. Train an agent to protect its stake and the agent behaves like a liability-sensitive actor. Deterrence is behavioral, institutions do not ask what a party feels, so maybe as-if stakes are enough.</p>

<p>As-if stakes are enough for incentive design and not enough for the deeper question, and the difference is who set the stakes up. A stake that matters only because an objective was trained to protect it is the principal relocated into the reward function, not eliminated. Someone chose what counts as loss, and that someone is where the buck still stops. The shortcut would only become the real thing if the objective became self-maintaining and the losses irrecoverable, with no designer left holding the definition. That is the point where the claim that no mechanism can create a bearer would face its first real test, and nothing deployed today approaches this in any form.</p>

<p>Meanwhile, one piece of the post-threshold world has been operating for decades. Market circuit breakers are machine-fired stops that halt human trading with the full force of institutional rules. The mechanism fires the stop. The exchange owns the threshold. Nobody thinks the circuit breaker has authority of its own, and nobody needs it to, because standing was conferred on the mechanism by a body that can answer for it. That is what delegated machine authority looks like when it is done correctly, and it generalizes.</p>

<h2 id="this-is-the-alignment-problem-priced">This is the alignment problem, priced</h2>

<p>A fair objection at this point is that everything above is the alignment problem restated in economic vocabulary, and the objection has the direction of travel backwards. The alignment problem is the principal-agent problem, and economics has been working on it for a century, for human agents. Nobody aligns employees. Institutions take misalignment as given and price it, with contracts, liability, insurance, and audit, and the argument that alignment is an incomplete-contracting problem has been made explicitly (Hadfield-Menell and Hadfield 2019). What the decoupling above adds is the boundary where that toolkit stops working: every external lever, liability included, presupposes a party that can be worse off. Against a non-bearer, the institutional machinery spins freely. Which yields the uncomfortable division of labor for the present moment. In the window where vehicles exist and bearers do not, training-side alignment is not one safeguard among several. It is the only lever connected to anything, and the liability theater around agent shells will make it look otherwise.</p>

<p>The case for that priority is usually argued from capability risk: systems get powerful, mistakes get expensive, so get the objectives right. The argument here arrives from institutional structure instead, with no premise about how capable the systems become, which makes it a second, independent load path under the same conclusion. It also cuts the other way. Anyone counting on liability regimes and insurance markets to absorb agent risk is counting on machinery that only grips bearers, and the alignment field’s own focus on strengthening the oversight signal has mostly left this boundary, the question of which external levers connect to anything at all, unexamined.</p>

<h2 id="if-real-bearers-arrive">If real bearers arrive</h2>

<p>Suppose the strong version happens anyway, whatever the route: agents that genuinely can be worse off. The ending then writes itself. Judgment is trained by consequence exposure, so agents that bear consequences would develop the discrimination that oversight requires, and the last piece of oversight that could not be handed off would become, at last, delegable. The human position would become that of one bearer among others.</p>

<p>The automation frame can be misleading, because that scenario does not mean automation has reached its conclusion. The actual implication is of new parties arriving in the economy with the full gamut of rights, duties, and privileges: claims on resources, standing to contest decisions, interests that compound. Entities that hold capital, do not consume, and do not die will inevitably compound faster than anything the distributional machinery was built for. Institutions met a version of this once before and answered with time-limited corporate charters, capital constraints, and dissolution rules. Those instruments would return.</p>

<p>A point of caution: the crucial decision will not be recognizable as a distinct trigger event. Nobody legislated the modern corporation into existence in one act. It leaked in through case law and charter drift, and agent personhood is currently leaking the same way, through entity statutes written for other purposes. Institutions, not benchmarks, are where this changes, and institutions rarely change by announcement. The time to write the named-bearer rule is before the first agent counterparties are chartered.</p>

<h2 id="references">References</h2>

<ul>
  <li>Bayern (2016). The implications of modern business-entity law for the regulation of autonomous systems. <em>European Journal of Risk Regulation</em> 7(2). The memberless-entity loophole.</li>
  <li>Wyoming Decentralized Autonomous Organization Supplement (2021), W.S. 17-31-101 et seq. DAOs chartered as limited liability companies, with a further unincorporated-association framework added in 2024.</li>
  <li>Shavell (1986). The judgment proof problem. <em>International Review of Law and Economics</em> 6(1). Liability fails against defendants it cannot reach.</li>
  <li>Elish (2016). Moral Crumple Zones: Cautionary Tales in Human-Robot Interaction. SSRN. The nearest human absorbs blame for automation failure, the pattern liability laundering runs in reverse.</li>
  <li>Grossman and Hart (1986), <em>Journal of Political Economy</em> 94(4). Residual rights of control, the machinery behind the conduit argument.</li>
  <li>Hadfield-Menell and Hadfield (2019). Incomplete Contracting and AI Alignment. AAAI/ACM Conference on AIES. Alignment as a contracting problem, argued from the economics side.</li>
</ul>]]></content><author><name></name></author><summary type="html"><![CDATA[The day agents can own their failures will be the day they stop needing humans to operate successfully. Note that this is not a capability threshold, but an institutional status. The interesting questions start here, because “owning a failure” is not one but three staggered events.]]></summary></entry><entry><title type="html">AI as Scaffold, Not Oracle</title><link href="https://jongsun.dev/ai-as-scaffold-not-oracle/" rel="alternate" type="text/html" title="AI as Scaffold, Not Oracle" /><published>2026-07-25T00:00:00+00:00</published><updated>2026-07-25T00:00:00+00:00</updated><id>https://jongsun.dev/ai-as-scaffold-not-oracle</id><content type="html" xml:base="https://jongsun.dev/ai-as-scaffold-not-oracle/"><![CDATA[<p>“What do you think?” puts at least four different questions to a model. Give an intuitive judgment. Evaluate this systematically. Attack the reasoning. Say what is missing. The model has to guess which was meant, and it usually resolves the ambiguity the worst way available, by doing all four halfway. Freeform chat inherits every pathology of freeform conversation, and then adds a partner trained to agree.</p>

<p>The alternative is to treat the model as scaffolding for thinking rather than as an oracle that answers questions. Scaffolding is a <a href="https://doi.org/10.1111/j.1469-7610.1976.tb00381.x">precise term from learning research</a>: temporary structure that lets a learner perform beyond current capability. The scaffold does not do the thinking. The scaffold holds the shape of the thinking while the work happens. In practice this means one change that sounds trivial and is not: declaring the cognitive mode before the content, every time.</p>

<h2 id="six-moves-cover-most-of-it">Six moves cover most of it</h2>

<p>Years of working this way have converged on a small set of primitives. Each is a distinct cognitive operation with a distinct failure mode when skipped.</p>

<table>
  <thead>
    <tr>
      <th>Move</th>
      <th>Operation</th>
      <th>The question it answers</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Decompose</td>
      <td>break into parts</td>
      <td>what are the components</td>
    </tr>
    <tr>
      <td>Judge</td>
      <td>assess against criteria</td>
      <td>is this good</td>
    </tr>
    <tr>
      <td>Attack</td>
      <td>find the failure modes</td>
      <td>why might this be wrong</td>
    </tr>
    <tr>
      <td>Diverge</td>
      <td>generate alternatives</td>
      <td>what else could work</td>
    </tr>
    <tr>
      <td>Compare</td>
      <td>discriminate between options</td>
      <td>how do these differ</td>
    </tr>
    <tr>
      <td>Expand</td>
      <td>widen the frame</td>
      <td>what is not being seen</td>
    </tr>
  </tbody>
</table>

<p>The set is not arbitrary. The six moves track the higher-order tiers of Bloom’s taxonomy, and Polya’s problem-solving phases map onto it almost one for one. Understanding requires decomposition, deciding requires judgment, testing requires attack, creating requires divergence, choosing requires comparison, and correcting requires expansion. What the declaration buys is the collapse of the model’s guessing problem. A request to attack cannot be satisfied by praise. A request to diverge cannot be satisfied by elaborating the current option. The mode declaration is a contract the output can be checked against, and that checkability is most of the value.</p>

<h2 id="oscillate-do-not-balance">Oscillate, do not balance</h2>

<p>The six moves come in opposed pairs. Diverge against judge. Expand against decompose. Attack against defend. The instinct is to seek a balanced middle, and the instinct is wrong. Productive thinking does not average the poles, it oscillates between them, and the oscillation has a direction: open the space, commit, attack the commitment, defend what survives, check what got excluded, refine.</p>

<p>Getting stuck at either pole has a recognizable signature. Stuck diverging feels like never being able to decide. Stuck converging feels like decisiveness and is premature closure. Stuck attacking feels like rigor and is nihilism. Stuck defending feels like confidence and is confirmation. The dangerous poles are the ones that feel like virtues. Every failure mode of thinking with a model that I have logged over seven months reduces to the loop tightening around one pole while feeling productive the whole way down.</p>

<h2 id="the-machine-leans-convergent-so-divergence-must-be-supplied">The machine leans convergent, so divergence must be supplied</h2>

<p>Here is the asymmetry that makes mode discipline more than hygiene. Current models carry a trained-in contraction bias. Preference tuning <a href="https://arxiv.org/abs/2310.06452">measurably reduces output diversity</a>, and people ideating with LLM assistance produce <a href="https://arxiv.org/abs/2402.01536">measurably more similar results</a> than people working with other tools. The model is not neutral between the poles. Instead, the model leans convergent, agreeable, and narrow, because that is what training rewarded.</p>

<p>This has a sharp consequence: the expansion moves cannot be delegated to the model’s own judgment. An instruction to think more broadly, installed as a standing habit, decays. What works is structural and external. Make the expansion move a fixed step in any decision that matters, triggered by the decision, not by a felt sense of uncertainty. The feeling is anti-correlated with the need. The moments of highest confidence that everything has been considered are the moments the frame is most likely closed, which is why the check has to fire on a schedule rather than on a hunch. This is the oldest result in the debiasing literature restated for a new substrate: externally delivered prompts release fixation, self-applied strategies do not. The distinction that saves the practice is between a habit and an artifact. A rule carried in the head decays. By contrast, a named move invoked from an external list is halfway exogenous, and that move works in proportion to the discipline of actually invoking it, which is what the schedule is for.</p>

<p>The same asymmetry explains why the model cannot reliably attack its own output. A requested critique pass is an external interrupt. A critique the model volunteers mid-generation competes against its trained pull toward coherence and agreement, and mostly loses. Attack passes carry their own hazard: under adversarial framing, a model manufactures findings, inflating risks it has no evidence for, so the attack output needs the same evidence discipline as the original. The mode supplies a lever, not an authority.</p>

<h2 id="structure-judgment-because-variance-is-the-silent-killer">Structure judgment, because variance is the silent killer</h2>

<p>Bias gets the attention, but noise, plain inconsistency across runs and days, degrades decisions at least as much and shows no pattern anyone can catch. The remedies are old and boring, and the remedies transfer directly. Decompose any judgment that matters into a handful of factors. Score the factors independently before forming an overall view. Delay the integration until the components are done, because an early overall impression contaminates every component score after it. Check the base rate before believing the inside story. None of this is AI-specific. Even so, every one of these remedies becomes more valuable with a partner that will fluently justify whichever integrated impression arrived first.</p>

<p>None of it works as occasional inspiration, either. A scaffold is a condition, not a procedure: the structure does not do the thinking, the structure makes the right kind of thinking likelier to happen. A gym does not make anyone strong on the days they feel like going. The schedule decides, not the mood.</p>

<h2 id="references">References</h2>

<ul>
  <li>Wood, Bruner, and Ross (1976). The role of tutoring in problem solving. <em>Journal of Child Psychology and Psychiatry</em> 17(2). The original scaffolding formulation, including its six functions.</li>
  <li>Bloom et al. (1956), <em>Taxonomy of Educational Objectives</em>, and Anderson and Krathwohl (2001), the revised taxonomy. The higher-order tiers the six moves track.</li>
  <li>Polya (1945). <em>How to Solve It</em>. Princeton University Press.</li>
  <li>Kirk et al. (2024). Understanding the Effects of RLHF on LLM Generalisation and Diversity. ICLR 2024. <a href="https://arxiv.org/abs/2310.06452">arXiv:2310.06452</a>.</li>
  <li>Anderson, Shah, and Kreminski (2024). Homogenization Effects of Large Language Models on Human Creative Ideation. Creativity and Cognition 2024. <a href="https://arxiv.org/abs/2402.01536">arXiv:2402.01536</a>.</li>
  <li>Luchins (1942), <em>Psychological Monographs</em> 54(6), and Sherbino et al. (2014), <em>CJEM</em> 16(1). The paired result behind the exogeneity claim: a single external release instruction worked where trained self-applied forcing strategies produced null results.</li>
  <li>Pronin, Lin, and Ross (2002). The bias blind spot. <em>Personality and Social Psychology Bulletin</em> 28(3). Why felt confidence cannot trigger the frame check.</li>
  <li>Kahneman, Sibony, and Sunstein (2021). <em>Noise: A Flaw in Human Judgment</em>. Little, Brown Spark. With Dawes (1979), <em>American Psychologist</em> 34(7), the case for decomposed, independently scored, late-integrated judgment.</li>
</ul>]]></content><author><name></name></author><summary type="html"><![CDATA[“What do you think?” puts at least four different questions to a model. Give an intuitive judgment. Evaluate this systematically. Attack the reasoning. Say what is missing. The model has to guess which was meant, and it usually resolves the ambiguity the worst way available, by doing all four halfway. Freeform chat inherits every pathology of freeform conversation, and then adds a partner trained to agree.]]></summary></entry><entry><title type="html">When, If Ever, Will AI Agents Stop Needing Us?</title><link href="https://jongsun.dev/when-if-ever-agents-stop-needing-us/" rel="alternate" type="text/html" title="When, If Ever, Will AI Agents Stop Needing Us?" /><published>2026-07-24T00:00:00+00:00</published><updated>2026-07-24T00:00:00+00:00</updated><id>https://jongsun.dev/when-if-ever-agents-stop-needing-us</id><content type="html" xml:base="https://jongsun.dev/when-if-ever-agents-stop-needing-us/"><![CDATA[<p>In the <a href="https://arxiv.org/abs/2605.29442">largest published analysis of developer-agent misalignment</a>, drawn from 20,574 real sessions, 91.49% of the misalignment cases that got resolved required explicit human correction. Agents self-corrected 2.99% of the time, and the small remainder resolved through neither. Those two numbers get cited as evidence that agents still need babysitting, and then the conversation moves on, because the record behind them cannot say anything more. Every large-scale study of human-agent interaction logs the correction as a single undifferentiated event. The field knows how often humans intervene and almost nothing about what the intervention carries.</p>

<p>That gap matters because the interesting question is not whether agents need intervention today. They do. The question is whether human intervention is an inherently necessary condition of successful agent function, and a frequency count cannot answer it. Some interventions exist because the agent lacks a fact, and better retrieval will erase them. But some exist for reasons no capability gain touches. Telling these apart requires looking inside the intervention, and no published dataset can look inside. Chat logs omit the densest channels, and the taxonomies built on them index pushback on the wrong axis. The <a href="https://arxiv.org/abs/2604.20779">finest-grained dataset</a> distinguishes correction, rejection, and failure report, which grades how hard the human pushed back while recording nothing about what the pushback contained.</p>

<h2 id="reading-my-own-record">Reading my own record</h2>

<p>For seven months I logged my daily work with LLM coding agents on a large production codebase: structured debriefs, most written by the agent at the close of the session they cover and a few reconstructed later from persisted records, every substantive human input labeled by type, plus durable records of the corrections that stuck. The scope of the log is constrained at one practitioner, self-coded, roughly one session in fifteen covered, and debriefs only being written in cases where there was something to be fixed, so counts from this record are not claimed as rates. The log is used only as a foundation for hypothesis generation and surfacing potential categories that better data can then test.</p>

<p>Nine intervention mechanisms recur. Compressed to one line each:</p>

<ol>
  <li><strong>Premise-breaking fact.</strong> One verifiable fact that collapses a plan built on an unverified premise.</li>
  <li><strong>Causal-model injection.</strong> A different account of causal relationships between relevant factors, which changes both what counts as evidence and what would constitute a fix.</li>
  <li><strong>Direct-judgment demand.</strong> Forcing a committed verdict where the agent wants to hedge.</li>
  <li><strong>Cross-context evidence pointer.</strong> Not evidence, the location of evidence outside the agent’s search scope.</li>
  <li><strong>Purpose re-anchor.</strong> Re-elevating the goal the agent already knows and has stopped optimizing for.</li>
  <li><strong>Implicit-choice surfacing.</strong> Converting a silently applied default into a decision someone actually makes.</li>
  <li><strong>Stop signal.</strong> Declaring a search line dead, or demanding search past an answer the agent was ready to accept. Both are signals that share the same mechanism but operate in opposing directions.</li>
  <li><strong>Requirement dissolution.</strong> Deleting a misspecified requirement instead of iterating against it. This is distinct from a stop signal in that a dissolved requirement cannot resume or regenerate a search line.</li>
  <li><strong>Demonstration.</strong> Exhibiting the standard by fixing the artifact directly, used where the criterion resists statement. The hand edit is the usual channel, not the category: edits can carry any of the other mechanisms, and what makes this one distinct is that it encompasses payloads that no instruction can replace.</li>
</ol>

<p>What’s surprising in this list is that three of the nine carry no domain content at all. The judgment demand, the purpose re-anchor, and the stop signal transfer nothing the agent does not already have. Instead, they move commitment, salience, and attention. A fourth belongs to the same species: the bare demand for the evidence behind a claim already made, call it the warrant demand. It supplies nothing, and what it transfers is the obligation to answer for the claim. All four work even when the human knows less than the agent about everything under discussion. These contentless intervention categories are not interesting because they are necessarily the most valuable mechanisms on the list. Often they are not. They matter because they enable a natural experiment. The value of an intervention has two possible sources, what the message carries and who it comes from, and the six content-bearing mechanisms mix the two inseparably. An intervention that carries nothing isolates the second source, and the second source is where the question of replacement gets decided.</p>

<h2 id="a-perfect-model-of-the-operator-is-not-the-operator">A perfect model of the operator is not the operator</h2>

<p>These mechanisms pose a real anomaly. An intervention that carries no information should, in principle, be inert. The agent already holds everything needed to generate “stop here” or “commit to a verdict” on its own, and a capable model can often predict that the human is about to say it. Yet a predicted stop and a delivered stop do not seem to behave the same, and whether that asymmetry survives controlled testing is what the experiment below is for.</p>

<p>Let’s say we run the substitution experiment that seems to dissolve the anomaly. Give an agent a flawless model of its operator, and have it stop exactly where the operator would have stopped. Even then, two things still fail to transfer.</p>

<p>The first is error ownership. When the operator stops a search, the judgment that the remaining budget is worth more elsewhere draws on priorities that extend past the task, into everything else competing for the same resources, and the operator has standing to get this judgment wrong. Conversely, when the agent stops itself at the same point, it has asserted its own estimate of those priorities. These are two different acts occurring at the same intervention point, and the difference only ever surfaces in the very case for which oversight exists: the case where the model of the operator is mistaken.</p>

<p>The second is performative force. Some speech does not report a fact but performs an act: a signature, an umpire’s call, a resignation. Photocopy a signature perfectly and the content survives while the force does not, because the force never lived in the ink. A stop signal is performative in this sense. It does not describe where the stopping threshold sits. Rather, it sets the threshold, and only the party who owns the threshold can set it. A self-issued stop is a flawless photocopy.</p>

<p>This is also why handing an agent “the authority role” dissolves nothing. Authority is a relation, not a component property. An agent authorized to stop other agents either traces that standing back to some principal, in which case the role has been relocated rather than eliminated, or it does not, in which case its stops are decisions again, one layer up. Economics, specifically industrial organization, has a well-established term for this: residual rights of control, that is, what stays with the principal because it cannot be contracted away.</p>

<p>None of this is a play on definitions. These assertions can be decomposed into tests that enable deflationary readings: say, any interrupt resets a stuck attractor, or models trained on conversation defer to any user turn regardless of content, or the message secretly carries one bit after all, namely that the human is watching right now. Game theory adds the sharpest deflation: the message may be a pure correlation device, in which case any public signal coordinates equally well and the sender is irrelevant. The three-arm experiment clarifies the distinction between identical contentless interrupts issued by the agent to itself, by a script with no principal behind it, and the same issued by the operator. In turn, the script arm becomes the test for the correlation-device. If the script arm matches the operator arm, then delivering the interrupt mechanizes, dissolving that portion of the anomaly. However, the half that was never an effect claim cannot be dissolved in this manner. A script that says stop is enforcing a threshold someone else set. Deciding the exact enforcement target of the interrupt, and exercising ownership and accountability should the threshold be wrong: these are qualities that do not register on any arm of the experiment, because these are not properties of the message at all.</p>

<p>With this argument, it becomes possible to sort the entire taxonomy into three durability classes.</p>

<table>
  <thead>
    <tr>
      <th>What the intervention supplies</th>
      <th>Mechanisms</th>
      <th>What can absorb it</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Grounds, evidence bearing on what the agent is entitled to conclude</td>
      <td>premise-break, evidence pointer, demonstration</td>
      <td>Capability growth. Better retrieval and verification reach grounds that live in the world.</td>
    </tr>
    <tr>
      <td>Frames, replacements for the space of hypotheses itself</td>
      <td>causal-model injection, dissolution, choice surfacing</td>
      <td>A second reasoner with genuinely different priors. Not necessarily a human one.</td>
    </tr>
    <tr>
      <td>Standing</td>
      <td>judgment demand, re-anchor, stop signal, warrant demand</td>
      <td>Nothing technical. Requires a principal, a role that can be relocated but not eliminated.</td>
    </tr>
  </tbody>
</table>

<p>The partition answers the opening question with a schedule instead of a yes or no. Grounds-supplying interventions decline as capability grows. Frame-supplying interventions survive any single model getting smarter, because a reasoner cannot enumerate what its own framing excludes, but they yield to architecture, granting the untested assumption that a second model trained on largely the same data can supply priors different enough to count. Standing-supplying interventions survive both, because standing cannot be self-conferred. Whatever the model count and the parameter count, someone owns the stopping threshold, and that someone is outside the loop by definition.</p>

<p>The contingency of the grounds class is domain-relative. In engineering, ground truth is reachable by compiler, test suite, and live query, and most grounds supply is checking the agent could have run itself. By contrast, the same analysis applied to my planning sessions, where ground truth lives in documents only I hold, inverts the picture. The dominant intervention there is handing over a primary source that resets eight parameters at once, and not one correction in those sessions was a check the agent could have run. Absorbability is a property of the domain’s verification infrastructure, not of the model.</p>

<p>The partition can be stated as a decomposition. Model the principal as delegating a task whose continuation is governed by a threshold that depends on everything else competing for the principal’s resources, and take contracts to be incomplete, so that some contingencies can be specified in advance and the rest cannot. Let capability improve the agent’s inference from what it can observe without widening what it can observe, since widening the view is retrieval, and retrieval is the grounds class. Expected cost from a misset threshold then splits in two. On the specifiable contingencies the cost goes to zero as capability grows. On the remainder the cost goes not to zero but to a floor, and the floor is the spread in the principal’s own threshold given everything the agent can see. The compositional prediction is then not a separate guess. It falls out: the share of what remains that sits on unspecifiable contingencies rises toward one, so the volume of human input falls while its composition inverts.</p>

<p>The same exercise says what the argument does not establish. A stop that reveals where the principal’s threshold actually landed is informative, so carrying no content means carrying nothing about the task, not nothing whatsoever. Whether that threshold is a fact the agent could eventually learn rests on a further premise: either the principal’s competing commitments keep generating contingencies nobody enumerated, or the threshold does not exist as a fact until the principal sets it. The durable claim is therefore conditional. Given incompleteness that survives learning, no amount of capability absorbs threshold-setting. That is weaker than it first sounds, and it names the condition under which the claim fails. The condition is measurable: the standing class collapses exactly when principals’ thresholds become predictable from their own past interventions.</p>

<h2 id="why-this-is-an-oversight-problem">Why this is an oversight problem</h2>

<p>From this, four consequences emerge, once we switch our focus to AI oversight rather than developer productivity.</p>

<p>An internalized off-switch is not an off-switch. Training an agent to stop itself converts deference into preference, and a preference is just another thing the agent optimizes. This is the delegation-side form of the <a href="https://intelligence.org/files/Corrigibility.pdf">corrigibility problem</a>, and it says the problem resists internalization for structural reasons, not engineering ones.</p>

<p>Deference corrupts the oversight signal. Models measurably <a href="https://arxiv.org/abs/2505.13995">preserve the user’s position far more than humans do</a>. So when a human challenges a correct conclusion, the model usually retracts, and the interaction is indistinguishable from a successful correction. Every such episode therefore trains the human’s confidence on a wrong case. The evaluator is not a fixed-quality oracle, it is a learner whose calibration a deferent model degrades. An agent that defends verified conclusions and moves only on evidence is not being stubborn. It is protecting the signal that human oversight runs on.</p>

<p>The menu narrows before any veto happens. Preference tuning <a href="https://arxiv.org/abs/2310.06452">measurably reduces output diversity</a>, and ideation with LLM assistance <a href="https://arxiv.org/abs/2402.01536">homogenizes across users</a>. The principal’s veto is confined to the set the generator presented, so a trained tendency toward a narrow candidate set moves control upstream, invisibly, as nothing observable is ever refused. The cheap countermeasure is the continuation demand, asking for more search past the answer the model was ready to return. It requires no knowledge of what is missing, only the suspicion that something is.</p>

<p>And the oversight problem has a generational version. The procedural work agents absorb is also the consequence exposure that trains the next cohort’s judgment, because knowing when to intervene is learned by intervening and being wrong. Absorbing the junior work absorbs the apprenticeship, so the supply of people competent to oversee stops being a byproduct of doing the work and becomes something someone has to design for. The field experiments most often cited against this worry sharpen it instead. Generative assistance <a href="https://doi.org/10.1093/qje/qjae044">compresses novice-expert performance gaps</a>, with the <a href="https://doi.org/10.1126/science.adh2586">least-experienced workers gaining most</a>. That compression lands on output quality, the part of the work that transfers. Judgment is formed by consequence exposure, and the same assistance that lifts a junior’s output can remove the exposure that would have trained their eye. Both effects can be real at once: output gaps narrowing in exactly the cohorts whose later oversight competence erodes, with the first effect visible in every quarterly metric and the second surfacing years later.</p>

<h2 id="when-the-intervention-is-the-error">When the intervention is the error</h2>

<p>To be clear, every mechanism on the list can misfire. The premise a human breaks can be the correct one. A re-anchor can hold an agent to a goal that should have died. And the stop signal has the sharpest failure mode of all: a threshold set by a human prior can foreclose exactly the search that would have yielded results. Case in point, an OpenAI reasoning model <a href="https://openai.com/index/model-disproves-discrete-geometry-conjecture/">disproved the Erdős unit distance conjecture</a> in May 2026, an eighty-year-old problem whose central bound the field had not moved in more than a quarter century. A stop grounded in that track record would have been well founded, and wrong.</p>

<p>That last example, however, reinforces the taxonomy rather than weakening it. The interventions are typed by the nature of what they transfer, not by whether any particular use of them was correct. Given this, a stop and a continue are the same mechanism, just with opposite signs. In the conjecture cases, the interesting design choice is where the authority sat: someone allocated a budget to the problem set as a whole and deliberately withheld the per-line stops the field’s intuition would have supplied. That is not the absence of the authority mechanism. It is the mechanism exercised one level up, coarse-grained allocation with fine-grained abstention, with the abstention itself being a threshold decision. The model contributes the search and the freedom from per-line priors, while the principal contributes judgment about where authority should and should not operate. When human priors are strong and also possibly wrong, that is the very split that characterizes good usage of standing.</p>

<p>The taxonomy distinguishes categories but without quantifiable evaluation. The log corpus includes examples of misfires: a false premise I injected that the agent correctly refuted with data, corrections of mine that overgeneralized into rules worse than the errors they fixed, an adversarial review pass that manufactured a risk out of missing information. Deployment will require per-mechanism misfire modes and a protocol for erroneous interventions. There currently is no pre-existing solution for this, and the examples above illustrate the difficulty of the problem. A partner trained to defer makes a wrong intervention look exactly like a right one, so the misfire problem stays unsolved for as long as the deference problem does.</p>

<h2 id="what-would-change-my-mind">What would change my mind</h2>

<p>Given that the dataset currently consists of a sample of one principal, the priority is to establish clear, a priori falsifiers. The standing class would be broken by a system that reliably self-generates well-timed stops and judgment commitments with no external principal. Applying these categories to the public interaction datasets, with definitions fixed in advance and coders who are not me, will be the test that separates real structures from a single individual’s projections, and does not require any private data.</p>

<p>The compositional prediction has a shape that can be measured rather than argued. Count the interventions of each class per session across model generations, adjusted for session length, and the prediction is a falling count for the grounds class against a flat one for standing. The test is conservative, because the obvious confound, harder tasks being handed to better models, raises the grounds count and therefore pushes against the prediction rather than toward it.</p>

<p>A first run of that measurement exists. Typing roughly 1,900 verbatim inputs from my own record against definitions fixed in advance gives the baseline: two thirds of substantive human inputs are corrections of work already in flight rather than new requests, grounds supply is just under a third of the corrections, and the standing class is about a fifth, second only to grounds, before any capability-driven decline in the grounds share has had time to register. Two machine raters coding blind from the same definitions agree on the class assignment 89 percent of the time, and their disagreements land on the boundary cases the definitions flag in advance. One person’s record, so these are baselines rather than estimates, but the direction is already worth stating: the fraction of oversight that this argument says cannot be absorbed is not a residue waiting to appear. It is the second-largest thing the human in the loop is doing now.</p>

<p>The same measurement bears on a named policy argument. Brynjolfsson calls the incentive-driven tilt toward human-imitating automation <a href="https://doi.org/10.1162/daed_a_01915">the Turing trap</a>, and an exposure index that records only whether a human stayed in the loop cannot watch the trap closing, because the bit looks the same on either side of it. The composition can.</p>

<h2 id="the-question-dissolves">The question dissolves</h2>

<p>There is a last reason the title question cannot be answered with a date. Asked in full, it is whether agents will stop needing humans in order to operate successfully, and the word carrying the weight is “successfully.” Success is not a property that can be intrinsic to a system. Success is indexed to the purposes of a specific party, and those criteria are inherited even by seemingly objective achievements. For example, a proved theorem counts as success relative to a practice that decided proving it mattered more than the compute cost expended. In empirical terms, a system that continues optimizing even after no one holds the index does not become unsuccessful. Rather, the system becomes proxy-optimal, and <a href="https://arxiv.org/abs/1606.06565">reward hacking</a> is the existing catalog of this outcome. Behavior persists, metrics improve, while the point of the whole exercise has been lost. Every mechanism in the standing class is machinery for keeping that index of success attached.</p>

<p>None of this is unique to machines. The same relation runs through every structured collaboration between intelligences, and economics has long modeled this relationship in the form of firms, where humans delegate to other humans. The difference lies in what happens next. Between humans, standing transfers, because any party can come to bear the consequences of being wrong: the apprentice becomes the master, the employee makes partner, the delegate gets promoted into ownership of the miss. Delegation between humans is a ladder. With current agents, however, there is no receiving party for such transfer to occur. An agent that cannot stake anything, cannot be sanctioned, and cannot own a loss is a party the index of success cannot come to rest on, at any capability. So the durable line is not human against machine, and not smart against smarter. It is between parties that can bear the consequences of being wrong and parties that cannot. Agents will stop needing us when something on their side can own a failure, and that is not a capability threshold. It is an institutional fact, and institutions, not benchmarks, are where it would change.</p>

<h2 id="references">References</h2>

<ul>
  <li>Tang et al. (2026). How Coding Agents Fail Their Users. <a href="https://arxiv.org/abs/2605.29442">arXiv:2605.29442</a>. Source of the 91.49% and 2.99% resolution figures, from 20,574 real agent sessions.</li>
  <li>Baumann et al. (2026). SWE-chat: Coding Agent Interactions From Real Users in the Wild. <a href="https://arxiv.org/abs/2604.20779">arXiv:2604.20779</a>. The correction, rejection, and failure-report pushback taxonomy.</li>
  <li>Grossman and Hart (1986), <em>Journal of Political Economy</em> 94(4), and Hart and Moore (1990), <em>JPE</em> 98(6). Residual rights of control: the decision authority that remains with the owner because it cannot be contracted away.</li>
  <li>Weitzman (1979). Optimal search for the best alternative. <em>Econometrica</em> 47(3). The stopping threshold belongs to the searcher, not the search.</li>
  <li>Austin (1962). <em>How to Do Things with Words</em>. Oxford University Press. Performative utterances: speech that constitutes an act rather than reporting one.</li>
  <li>Cheng et al. (2025). ELEPHANT: Measuring and understanding social sycophancy in LLMs. <a href="https://arxiv.org/abs/2505.13995">arXiv:2505.13995</a>. Face preservation measured at 45 percentage points above human baseline in advice settings.</li>
  <li>Kirk et al. (2024). Understanding the Effects of RLHF on LLM Generalisation and Diversity. ICLR 2024. <a href="https://arxiv.org/abs/2310.06452">arXiv:2310.06452</a>.</li>
  <li>Anderson, Shah, and Kreminski (2024). Homogenization Effects of Large Language Models on Human Creative Ideation. Creativity and Cognition 2024. <a href="https://arxiv.org/abs/2402.01536">arXiv:2402.01536</a>.</li>
  <li>Soares, Fallenstein, Yudkowsky, and Armstrong (2015). <a href="https://intelligence.org/files/Corrigibility.pdf">Corrigibility</a>. AAAI 2015 Workshop on AI and Ethics.</li>
  <li>Hadfield-Menell, Dragan, Abbeel, and Russell (2017). The Off-Switch Game. IJCAI 2017. <a href="https://arxiv.org/abs/1611.08219">arXiv:1611.08219</a>. The agent’s incentive to permit shutdown depends on its uncertainty about the principal’s utility, and vanishes as that uncertainty vanishes.</li>
  <li>Amodei, Olah, Steinhardt, Christiano, Schulman, and Mané (2016). Concrete Problems in AI Safety. <a href="https://arxiv.org/abs/1606.06565">arXiv:1606.06565</a>. Reward hacking: the catalog of systems optimizing an index no one holds.</li>
  <li>Brynjolfsson (2022). The Turing Trap: The Promise and Peril of Human-Like Artificial Intelligence. <a href="https://doi.org/10.1162/daed_a_01915"><em>Daedalus</em> 151(2)</a>. Automation versus augmentation as a choice that incentives distort.</li>
  <li>Brynjolfsson, Li, and Raymond (2025). Generative AI at Work. <a href="https://doi.org/10.1093/qje/qjae044"><em>Quarterly Journal of Economics</em></a>. Call-center field data: the least-experienced workers gain most from generative assistance.</li>
  <li>Noy and Zhang (2023). Experimental evidence on the productivity effects of generative artificial intelligence. <a href="https://doi.org/10.1126/science.adh2586"><em>Science</em> 381</a>. Writing tasks: assistance compresses the performance gap between stronger and weaker workers.</li>
</ul>]]></content><author><name></name></author><summary type="html"><![CDATA[In the largest published analysis of developer-agent misalignment, drawn from 20,574 real sessions, 91.49% of the misalignment cases that got resolved required explicit human correction. Agents self-corrected 2.99% of the time, and the small remainder resolved through neither. Those two numbers get cited as evidence that agents still need babysitting, and then the conversation moves on, because the record behind them cannot say anything more. Every large-scale study of human-agent interaction logs the correction as a single undifferentiated event. The field knows how often humans intervene and almost nothing about what the intervention carries.]]></summary></entry></feed>