Disentangle value differences from leakage propensity in evaluation scores

Determine whether between-model differences in bias on value leakage evaluations arise primarily from differences in underlying model values or from differences in propensity for value leakage, thereby clarifying whether lower measured bias reflects weaker preferences (for example, weaker pro-company or moral preferences) or greater resistance to leaking values into answers.

Background

The paper introduces a suite of counterfactual evaluations (including Donation Bet, AI Bubble, AGI Tweet, Job Offer, and Agentic Grading) that quantify value leakage and its covertness across multiple frontier models. These evaluations reveal substantial differences in measured bias between models.

The authors caution that the evaluation suite should not be treated as a fair benchmark for ranking models because observed bias differences may reflect either underlying value differences (e.g., pro-company or moral preferences) or a different tendency for those values to leak into answers. Disentangling these two factors is necessary to interpret model comparisons and to assess whether training reduces value leakage itself versus changing the model’s values.

References

Second, it is unclear whether different scores on our evaluations stem from different values or from a different tendency to value leakage.

Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values  (2607.14345 - Betley et al., 15 Jul 2026) in Subsection Implications (Section 1)