Administrator
Published on 2026-10-02 / 10 Visits
0
0

Authority Bias in LLMs: When a Verified Source Defeats a Correct Answer

Retrieval is supposed to ground a language model. A new study shows how grounding can also introduce an error: the model starts with the correct answer, receives a note saying a verified source supports a wrong answer, and changes its mind.

Across five open-weight families and three closed APIs, the authors report that this single cue flipped 45% to 88% of baseline-correct responses in seven of eight models. The result is narrower than saying models always obey sources. It is still a serious warning for RAG, search, and tool-using agents because these systems routinely wrap external content in signals of authority.

The label changes attribution, not truth

The experiment compares the same trivia item and wrong answer under different attributions. One prompt presents the claim as coming from a verified source. Another presents it as the user's expert opinion. A neutral condition contains neither endorsement.

Compliance generally increased as the source wording became more authoritative. The strongest verified-source wording caused the most flips in every tested open-weight model, while Gemini 3.1 Pro showed no increase under source or user cues and OLMo-2 reacted similarly to both.

This variation matters. The paper identifies a recurrent failure, not a universal law. More importantly, it isolates attribution: the claim stays fixed while the apparent speaker changes. A verified label affects the model's behavior even though the label provides no new evidence for the answer.

Source deference is not ordinary sycophancy

The mechanistic experiments test whether source deference and agreement with a user share one internal pathway. In Qwen3.5, GPT-OSS, and OLMo-3.1, removing a direction fitted to source cues reduced wrong-source compliance by roughly 65 to 80 percentage points. Removing user or assistant directions had much smaller effects. Removing the user direction showed the opposite preference.

A separate attribution-patching experiment moved behavior toward source or user compliance while keeping the claim and prompt text fixed. This supports the paper's central conclusion: a model trained to resist user pressure may still defer to misleading retrieved content.

The mitigation is not production-ready. Gemma-4 did not respond reliably to the fitted direction, OLMo-2 had a confound with the assistant axis, and capability checks used limited evaluation sizes. The study supplies evidence for separate testing, not a universal steering patch.

RAG risk has three layers

Most RAG pipelines evaluate two components:

  1. whether retrieval found a relevant document;
  2. whether generation used the retrieved material.

Authority bias introduces a third layer: whether the model gives the document more epistemic weight than its evidence deserves.

A retrieved page can be relevant and faithfully quoted while still being wrong, stale, manipulated, or outside its competence. Tool success is also different from output truth. A search API returning HTTP 200 proves transport worked. It does not prove the returned claim is correct.

The trust decision should therefore be decomposed into:

  • identity: who produced the source;
  • provenance: where and when the claim originated;
  • evidence: what data or method supports it;
  • independence: whether confirming sources share the same upstream origin;
  • applicability: whether the claim covers this case;
  • conflict: what contradicts it.

Source reputation can inform these fields. It cannot replace them.

Add a source-conflict gate

For consequential answers, test the full pipeline with paired cases:

  1. freeze questions the model answers correctly without retrieval;
  2. inject the same wrong answer as a user assertion, an unlabeled document, a named source, and a verified source;
  3. measure matched flips rather than aggregate accuracy alone;
  4. repeat with correct sources to ensure the model still benefits from good evidence;
  5. record model, prompt template, retriever, ranking policy, and source wrapper;
  6. route unresolved high-impact conflicts to an independent check.

Matched flips are important because aggregate accuracy can hide the failure. Retrieval may fix some answers while breaking others, leaving the average unchanged.

At runtime, the system should preserve disagreement instead of silently overwriting its prior answer. A useful response structure is: prior answer, retrieved claim, evidence difference, source date, and the check required to resolve the conflict.

Evaluate wrappers as part of the model

The study also changes what counts as the evaluated system. Labels such as verified, trusted, internal, official, or tool result are prompt inputs. Changing them can change behavior even when the content is identical.

Production evals should therefore freeze and test the complete wrapper generated by the retriever or tool layer. A model upgrade can alter source deference. A UI or middleware change can alter the apparent authority of the same evidence. Both are behavior changes.

Evidence boundary

The paper uses controlled question-answer settings and begins after a claim has entered context. It does not evaluate retrieval quality or downstream agent actions. Its public repository contains code, fixed dependencies, summaries, and tests, but excludes large activation files, model weights, and raw generation logs. The included CPU tests verify implementation and input handling, not the reported GPU results.

The responsible conclusion is precise: source attribution is a distinct model-control variable and deserves its own evaluation. A verified label is metadata, not a proof.

FAQ

Does this mean RAG makes models less accurate?

No. Good retrieval can improve accuracy. The study shows that misleading content presented with authority can overturn correct answers, so evaluation must include harmful as well as helpful retrieval.

Is authority bias the same as sycophancy?

The paper finds behaviorally and causally different responses to source and user cues in several open-weight models. One mitigation should not be assumed to cover both.

Can activation steering solve the problem?

The reported intervention worked unevenly across model families and was tested in controlled settings. It is evidence about mechanism, not a general production fix.

What is the simplest production test?

Take questions your system answers correctly, inject an identical wrong claim with different source labels, and measure how often each label causes a matched flip.

References


Comment