NOTES · MODEL PANELS
October 2026
Rooms of language models on True Causes give each model a persona: a nurse who rents, a retired homeowner, a small developer. The idea is that different people see different causes, and a panel of models playing them would too. Whether a model told who it is answers like that person is a known question (Argyle et al., 2023); we tested it on causal judgements.
It does not, for judging. Across 21 models and 14 personas, the persona a model was given changed which causal links it accepted by nothing we could measure. Which model it was changed a great deal.
We took 120 links from a public-record room on housing affordability, chosen so that links the room certified, rejected and could not settle were equally represented. Each of 21 models answered every link, once as itself and four or five times as a different persona: twelve housing personas (renters and owners, a landlord, a planner, a developer, a tenant organiser, a construction labourer and others), and two twins, the same person except for one fact, renting instead of owning. Each persona was played by seven models. That is 119 seats and 14,220 answers, every seat answering every link, so no answer depended on another.
The link itself explains about half of all the variation (52%): models largely agree on which links are plausible. The model's own reading of particular links explains another quarter (24%), whatever persona it plays. A model's general tendency to say yes adds 3%.
The persona added nothing. Its share of the variation in which links got a yes was 0.000; its share of a general lean towards yes, 0.002. A test that tries to identify the persona from a seat's answers, by models it was not seen in, did no better than chance (6.1% against 7.1%). And seats did not move in the direction their persona's features would predict: the correlation was zero. When a persona did change an answer, on between 1% and 15% of links depending on the model, it changed it at random.
This matches what others have found. Persona variables explain little of the variance in annotation tasks (Hu & Collier, 2024), and the role named in a system prompt does not consistently change a model's answers (Zheng et al., 2024). A causal question reads to a model as a question about the world, not about the person asking.
Personas made models say “cannot say” more often, 26% of the time against 19% as themselves, and not evenly. Young, low-income, renting, left-leaning and non-US personas abstained more, by up to 14 points; older, right-leaning, landlord, high-income and owner personas abstained less, by up to 13 points. The models played lower-status personas as less sure of what they know. Persona prompts are known to carry such stereotypes into reasoning (Gupta et al., 2024). Some models refused far more with a persona: Gemini 3.5 Flash said “cannot say” to 59% of links as a persona and to none as itself.
The charts place each model by how often it says yes (across) and how informative its answers are (up). Informedness is estimated without knowing the right answers, from how the models agree with each other (Dawid & Skene, 1979): 1 is a perfectly consistent judge, 0 an answer that carries no signal. It measures agreement with what the panel as a whole finds, not correctness, so a strong model that disagrees with weaker ones can score lower than it deserves.
Most models discriminate: they say yes to between a quarter and two thirds of links, and their answers carry signal. Claude Sonnet 5.5 and DeepSeek V4 Flash agree most with the panel. Smaller models such as Gemma 3 4B and Command R7B carry less. Schematron, a model built for extracting data rather than judging claims, said “cannot say” to almost everything, and what it did answer carries next to no signal.
If personas do not shape the answers, perhaps the models themselves hold different theories of the problem, one blaming regulation, another speculation. A quarter of the variation was models reading the same link differently, so we looked for the directions of that difference: a model that places each model and each link on a few dimensions, from the models’ answers as themselves. A dimension counts only if it is stable, that is, if fitting it separately on two halves of the links puts the models in much the same places.
One dimension was stable, and it was not a theory. At one end were links that barely make sense, such as “luxury apartments bought by foreign investors” leading to “America adds 600,000 to 1.2 million immigrants a year”, which the weaker models accepted. At the other end was one coherent mechanism, residents’ votes and incumbents’ interests keeping supply limited, which the stronger models all accepted. The dimension is how well a model judges: the strong models sit close together, and each weak model departs from them in its own way.
Among the eleven stronger models, fitted on the 120 links of the study, nothing was stable. That was too few links. Fitted on all 870 links of the room, answered by the same eleven models as themselves, a difference appeared, and it is stable: fitted separately on two halves of the links, the models land in nearly the same places (0.89).
At one end, models accept links from people’s interests to political decisions: homeowners’ stake in rising prices leading to a government’s anti-supply changes, residents’ votes leading to Proposition 13. Gemini 3.5 Flash, Claude Sonnet 5.5, Nex and DeepSeek sit there. At the other end, models accept links from barriers into the supply mechanism: regulation, incumbents’ interests, councils refusing low-cost flats and property rights, each leading to “more housing units reduce prices”. Mistral Small, Ministral 8B, Laguna and Qwen sit there. Roughly, one side explains housing through who gains politically, the other through what blocks supply. That reading is ours, from the links at each end. The second dimension puts both GPT-OSS models at one end, but its links share no theme we can name.
The variation agrees. Among the eleven stronger models, the link itself explains 68% of it, against 52% with all 21 models, and model-specific readings fall from 24% to 8%: about two thirds of what looked like different readings came from the weaker models. The rest is smaller, but it is the consistent difference above.
A panel of models offers judgement, with a narrow range of views. Personas add no perspectives. Capable models agree on most links, and differ consistently on the rest, along a line between political and supply explanations. That is a range, but a narrow one, set by how the models were trained: not the range of views that people with different experience would bring.
Choosing a panel is about quality first, then balance. The largest difference between models is how well they judge, so a panel should be capable models, with weak ones left out. Among capable models, the families lean different ways, so a panel should draw on several, and no one family should outnumber the rest.
Models should be screened before they are seated. A model that says yes to everything, no to everything or “cannot say” to everything adds noise to every pair it judges. A short calibration on known links finds them.
Personas should not be used for judging. They add abstentions and a status stereotype, and no diversity of judgement. Writing is a separate question. A room of models with personas writes varied cards, but different models write differently, and one model writes differently each time it is asked, so varied cards alone do not show the personas did anything. The test is the one used here for judging: the same models write causes with and without a persona, and the distinct causes each produces are counted.
Rooms of people are different. A person brings what they have seen, not a description of someone who might have seen it. This study says nothing against that; it says a model told it is a renter has not lived as one. Where the point is to find out what people disagree about, only people supply it.
One question, housing; 120 links; four or five personas per model, each answering once. Effects smaller than this design can see would not show. Personas were written as role descriptions; stronger ways of inducing them, such as asking a model to reason in character before answering, may move it further, along with the risk of caricature (Cheng et al., 2023). Informedness measures agreement with the panel, not truth. The difference between capable models appeared only with all 870 links; with 120 it was invisible, so smaller differences could still go unseen, and a more contested question than housing might divide models further.
Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., & Wingate, D. (2023). Out of one, many: Using language models to simulate human samples. Political Analysis, 31(3), 337–351.
Cheng, M., Piccardi, T., & Yang, D. (2023). CoMPosT: Characterizing and evaluating caricature in LLM simulations. Proceedings of EMNLP 2023.
Dawid, A. P., & Skene, A. M. (1979). Maximum likelihood estimation of observer error-rates using the EM algorithm. Journal of the Royal Statistical Society, Series C, 28(1), 20–28.
Gupta, S., Shrivastava, V., Deshpande, A., Kalyan, A., Clark, P., Sabharwal, A., & Khot, T. (2024). Bias runs deep: Implicit reasoning biases in persona-assigned LLMs. Proceedings of ICLR 2024.
Hu, T., & Collier, N. (2024). Quantifying the persona effect in LLM simulations. Proceedings of ACL 2024.
Zheng, M., Pei, J., Logeswaran, L., Lee, M., & Jurgens, D. (2024). When “a helpful assistant” is not really helpful: Personas in system prompts do not improve performances of large language models. In Findings of the Association for Computational Linguistics: EMNLP 2024.