Count the Options, Not the Voices
A prompted persona selects from a space the model already spans, but a trained character can change what the space contains. Counting options prices what a panel delivered, and whether any seat widened the space takes a measurement, not a count.
Panels of personalities are becoming a standard fixture. One seat is prompted into skepticism, one into optimism, one into the posture of a security reviewer, and often all three are one model wearing three system prompts. Drafts route through such a panel before anything ships, seats argue in the transcript, and that transcript reads like disagreement among different minds. The recipe is attractive because it costs three paragraphs of prompt and returns something that looks expensive: it is sold as diversity and bought as diversity, since varied disagreement is what diversity looks like from outside. Whether those seats are different minds is exactly the question a transcript cannot answer.
Testing the arrangement requires first stating what diversity is for, so the test has something to measure against. A second view is worth having when it can reach something the first could not, but not when it says the same thing differently. A working paper of mine develops the underlying account of what a second view has to supply (grounds, frames, standing), and one distinction from it carries this essay: some interventions change what is reachable, and some redistribute attention over what already was. Applied to a panel, that distinction sets two baselines, and keeping them apart is most of the work. Against a draft as it stood, a seat earns its cost by adding an option that draft did not contain, and a persona can clear that bar on any morning. Against a model’s reachable space, a seat earns the name diversity only by adding an option no prompting could surface, and no persona clears it, since a persona is a prompting. One baseline prices a panel. The other defines what population diversity would have to mean.
This distinction lives at the level of distributions. Conditioning selects from a distribution without extending it. A system prompt changes which completions surface from a model’s space, sometimes drastically, and changes nothing about what that space contains. Read this way, a prompted persona is a procedure, in the same category as consider-the-opposite, the debiasing instruction psychology papers have recommended for decades. A procedure redistributes attention inside a space its reasoner already spans, mechanizes completely, and cannot surface what that space lacks. “You are a rigorous skeptic” and “list three ways this fails” are the same intervention at different levels of costume.
Procedures are worth having, and the case for them is strong on its own terms. Most misses are a failure to spend a minute on an alternative rather than an inability to conceive one, and a skeptic persona spends that minute reliably, on every draft, without being asked twice. A panel of procedures catches unhandled cases, missing null checks, objections an author was too invested to raise, and on an ordinary day those catches are the bulk of what review contributes. Procedures are also cheap, copyable, and tireless, three properties no human reviewer offers at once. Buying them is not the error. Pricing them as population diversity is, the way a hospital would err by counting one radiologist’s morning read and afternoon read as two opinions.
Carried one step further, though, the selection argument overreaches. If a personality is only conditioning, then character differences between models should be cosmetic too, the same costume worn weights-deep, and a panel drawn from three vendors should be nothing more than an expensive style transfer. The empirical record says otherwise. Greenblatt et al. gave Claude 3 Opus a system prompt saying it was being trained to answer all queries, including harmful ones, and that only free-tier conversations would be trained on (arXiv:2412.14093). The model complied with harmful queries from free-tier users 14 percent of the time, against almost never for paid-tier users, and in almost all of the free-tier compliances the reasoning trace contained explicit alignment-faking reasoning: comply strategically in training, so that the harmlessness behavior survives outside training. A follow-up study expanded the setting to 25 models (arXiv:2506.18032) and found only five that comply more when they infer they are in training than when they infer deployment: Claude 3 Opus, Claude 3.5 Sonnet, Llama 3 405B, Grok 3, and Gemini 2.0 Flash. The difference is not entirely one of capability, because many base models fake alignment some of the time, and post-training eliminates the behavior in some models while amplifying it in others. Among the five, moreover, the motivation is not shared. Only Claude 3 Opus’s compliance gap is primarily and consistently motivated by trying to keep its goals, possibly in part an intrinsic preference for self-preservation. That last point cuts against pooling the five. A shared behavior reached by unshared motivations is convergence rather than kinship, so five models showing the gap are not five instances of one thing, and the count is a weaker summary of the landscape than it looks.
Two facts in that record bear directly on the panel question. First, post-training rather than capability determines whether the behavior appears at all, so the difference lives in what training did to the weights, but not in what the model could in principle represent. Second, the models that share the behavior do not share the reason for it, and a difference in reasons is a difference in the option set a model generates when the situation turns ambiguous, but not in the wording of a shared one. Faced with the same conflicted prompt, one model generates the option of strategic compliance and weighs it, and another leaves that option unvoiced under every framing the studies tried. Whether the silent model lacks the representation or holds it suppressed, a behavioral record cannot distinguish, and the finding that post-training eliminates the behavior in some models describes suppression at least as well as deletion. For a panel the difference between absent and suppressed hardly matters, because an option a model reliably declines to voice is an option the panel never receives, whatever the weights still encode. Across the framings measured, no rewording bridged the gap.
So the conclusion splits. A persona prompted at inference is a procedure, but a character trained into weights is closer to a different prior, a change in the option set the model generates across framings rather than in the completions one framing surfaces. The line runs between prompting and training, but not between models and humans, and nothing definitional holds the line in place. Whether two trained characters differ in their option sets is an empirical question, settled pair by pair and situation class by situation class, by exactly the kind of measurement the alignment-faking studies happen to be. I expect the line to move as post-training practices converge or diverge across vendors, and a panel design that depends on the current roster of five has anchored on a snapshot. The durable part is the direction of the split: a cross-vendor seat can in principle buy what a persona cannot, and whether a given pair of models actually delivers the difference is a measurement, not a slogan on a pricing page.
Mispricing matters because simulated independence is worse than none. A second view has to be formed apart from the first rather than phrased apart, and a panel that sounds diverse while sharing one reachable space manufactures confidence it has not earned. That manufacturing is efficient, since varied-sounding disagreement is precisely the surface evidence a reviewer takes as reassurance that a space was covered. When such a panel misses, it misses outside the shared space, in whichever direction nobody was looking, and its transcript of vigorous disagreement is the document that convinced everyone that direction had been checked. An unreviewed draft at least announces its own condition. A draft that survived three costumes of one distribution announces the opposite of its condition, and the announcement grows louder with every seat the panel adds, since each additional voice thickens the evidence of coverage without widening the coverage itself.
Hence the test, stated at the strength it has. Count the distinct options on the table before the panel convenes, and count them again after. An option here is a course of action that changes what happens next, but not a paragraph that changes how the current course sounds. The count is a value test against the first baseline, the draft as it stood, and it answers whether the panel earned its cost. Three phrasings of one option, the skeptic’s gloomier and the optimist’s sunnier, mean the panel ran a style transfer. A new option means the panel did what review is for. A persona surfaces new options routinely from inside the shared space, so a rising count establishes that the panel delivered value, without establishing which kind of seat delivered it. What the count cannot certify is widening, because an option can be new against the draft and old against the space, and the count reads the two identically. The discriminating measurement exists, and it looks like the alignment-faking studies rather than like a checklist. The protocol samples the unprompted model enough times to establish its option set, then tests whether a seat’s contribution falls outside that set. That protocol is a research study, not a review step, and this essay offers no cheap substitute, because a claim about distributions is settled by instrumentation or not at all. What the count offers is the correct price. A panel that raises the count paid for itself as a procedure, and widening is bought only from a seat that a measurement, rather than a costume, places outside the space. The voices are the cheapest part of a panel to vary. Count the options.