Language-model groups overstate consensus when replaying human deliberation on a reasoning task
Tengfei Shao
- Published
- Sep 17, 2026 — 15:11 UTC
Problem
This preprint addresses the gap in understanding how full-consensus rates as indicators of collective cognition are influenced by participation and the operationalization of final states. The authors highlight that existing methods may overstate consensus when evaluating human deliberation through language model (LLM) simulations, particularly in reasoning tasks.
Method
The study employs a replay method where 100 held-out human Wason groups are replayed with matched LLM agent groups. Each agent is initialized based on a belief-anchored model corresponding to each participant's pre-discussion answer. The scoring method ensures that both agents and human participants are evaluated using the same code, allowing for direct comparisons. Sensitivity analyses are conducted through two approaches: a submit-based comparison with 98 instances and a participation-matched comparison with 45 instances. Additionally, a reparameterization step is introduced to remove memorizable answers for groups operating in reasoning mode, aiming to enhance the validity of the reasoning assessments.
Results
The results indicate a full-consensus rate ranging from 24.0% to 57.0% when compared to human participation. In the submit-based comparison, the gaps between chat mode and reasoning mode are significant, with a 34.0 percentage point difference in favor of chat mode and a 43.9 percentage point difference in favor of reasoning mode. Similarly, in the participation-matched comparison, the gaps are 34.1 percentage points (chat mode vs. reasoning mode) and 44.4 percentage points (reasoning mode vs. chat mode). Notably, the convergence of analyses between the two methods is within 0.5 percentage points. However, the reasoning-mode group agreement is reported to be nearly unanimous, predominantly on incorrect answers, raising concerns about the reliability of consensus in this context.
Limitations
The authors acknowledge that agent groups exhibit a higher level of consensus compared to human groups, which may skew the interpretation of collective cognition. Furthermore, the simulated consensus does not align with collective accuracy, indicating a potential disconnect between perceived agreement and actual correctness. The belief-anchored agent groups are also noted to bias estimators of the human group-outcome distribution, which could misrepresent the dynamics of human deliberation.
Why it matters
This work has significant implications for the design and evaluation of AI systems that aim to simulate human reasoning and decision-making processes. By revealing the limitations of current consensus measures, it encourages a reevaluation of how collective cognition is assessed in both human and machine contexts. The findings suggest that reliance on LLMs for understanding human deliberation may lead to misleading conclusions, thereby impacting future research and applications in AI-driven decision support systems.
By Callan Zhang · Sep 17, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: arXiv cs.AI
