Does a model's stated reason for rejecting a candidate do any work?
Archit Rastogi
- Published
- Sep 24, 2026 — 17:13 UTC
Problem
This work addresses a gap in understanding the impact of a model's stated reasons for rejecting candidates, which is crucial for improving interpretability and trust in AI systems. The study is presented as a preprint and has not undergone peer review.
Method
The research utilizes the 2WikiMultihopQA corpus and evaluates six open models. The authors implement control mechanisms to assess the effect of stated reasons for rejection, including a length-matched irrelevant sentence and the same two sentences presented as a third option not mentioned by the model. The decoding method employed is greedy decoding. Statistical analysis is performed using odds ratios and Holm correction to evaluate the significance of the results.
Results
The study reports an odds ratio of 3.57 [1.54, 8.26] when comparing the model's stated reasons against the irrelevant control, indicating a significant influence of the stated reasons on decision-making. The Holm p-value for the named fact at a rival's profile is 0.0210, suggesting a statistically significant result. Conversely, the p-value for the same fact at the unmentioned option is 0.2428, indicating a non-significant result. Additionally, a Holm p-value of 0.0008 is reported for the irrelevant sentence at the named rival compared to the third option, further supporting the influence of stated reasons. Notably, 17.1% of adjudicable responses returned the option that the model had just rejected, highlighting potential defects in decision-making.
Limitations
The authors acknowledge that the content contrasts used in the study bound an effect rather than establish a definitive causal relationship. There is also a noted disagreement between the forced single-token probability read and the free-text choice, which may affect the interpretation of results. Furthermore, validation processes identified eight defects in measurement, indicating potential issues with the data collection or analysis methods employed.
Why it matters
Understanding the role of a model's stated reasons for rejection is critical for enhancing the interpretability and reliability of AI systems. This research has implications for downstream work in model design and evaluation, particularly in contexts where decision-making transparency is essential. By elucidating how stated reasons influence outcomes, this study can inform future efforts to create more robust and trustworthy AI systems.
By Callan Zhang · Sep 24, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: arXiv cs.AI
