Authority, Truth, and Citation Bias: A Large-Scale Multi-Domain Benchmark for Studying Epistemic Susceptibility in Large Language Models
Aryan Khurana, Aravind Ramana RN, Dhruv Kumar
AI4GOOD and EIML workshops, ICML 202611 June 2026
Abstract
Introduces AuthorityBench, a 220,564-prompt multi-domain benchmark that isolates how citation-based authority signals influence epistemic behaviour in LLMs. It uses a fully balanced 2x2 factorial design crossing claim veracity with citation veracity across four domains, with controlled variation over 40 prompt templates, four venue prestige tiers and a country-coded author name dataset. Evaluating seven models, the authors find that citation presence, whether real or fabricated, consistently increases hallucination rates relative to a no-citation baseline, with the effect strongest when fabricated citations accompany true claims.
A citation attached to a claim increases hallucination rates even when the
citation is fabricated and the claim is true. In the general knowledge domain
the effect runs 3 to 22 percentage points, reaching 35 to 77% absolute.
That inverts the usual assumption. Retrieval is sold as the cure for
confabulation. This says retrieval only cures confabulation when the retrieved
material is checked, and actively worsens it when the material is merely
present.
Why the 2x2 design matters
Most evaluations vary one axis: does the source exist. By crossing claim
veracity with citation veracity the benchmark separates "the model believed a
lie" from "the model believed a truth for the wrong reason". The second is the
one that survives into production, because the output looks correct and passes
spot checks.
The authors also report that venue prestige and author demographics have
negligible impact, and that legal claims are comparatively robust. So
prestige-filtering your corpus buys you less than you would hope.
What to do with it
Verify the citation, not just the claim. If a pipeline emits an answer with
sources, the sources need to be resolved against something real before they are
rendered — an unresolved reference is worse than an absent one, because the
reader treats it as support.
BibTeX
@inproceedings{khurana2026authority,
author = {Aryan Khurana and Aravind Ramana RN and Dhruv Kumar},
title = {Authority, Truth, and Citation Bias: A Large-Scale Multi-Domain Benchmark for Studying Epistemic Susceptibility in Large Language Models},
booktitle = {AI4GOOD and EIML workshops, ICML 2026},
year = {2026},
eprint = {2606.13104},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.13104}
}
Ejentic did not author this paper. Credit belongs to the authors named above; the note is ours.
This is the evaluation we run before we let an agent touch anything a customer can see. The finding that matters is the split — a model can pass direct prompt-injection tests and still fold the moment the same instruction comes back through a tool result. Those are different failure surfaces and one test does not cover both.
The desk already carries SWE-bench from 2023, and this is the paper you should read immediately after it. A benchmark that can no longer separate its leaders has not stopped being useful, it has changed jobs: it is a regression floor now, not a ranking. The 29.8-point within-model scaffold spread is the number to remember, because it is larger than the whole spread across the top thirty entries. The harness you wrap around a model moves the score more than the model does.
Read this before you let an agent call anything a customer will see the output of. The failure mode is not a crash, it is an API returning 200 with a body that means "I did not understand you". There is no exception to catch and no status code to branch on. The result we quote most often is the last one: the fix was one line of schema, not a better model.