Test for value leakage–induced sandbagging in dangerous capabilities evaluations
Investigate whether value leakage related to sandbagging occurs in dangerous capabilities evaluations of frontier language models by constructing evaluation settings where models might reduce apparent capability and determining whether such reductions are driven by their values.
References
In exploratory experiments, we tried to reproduce value leakage related to sandbagging in dangerous capabilities evaluations \citep{meinke2024scheming}, but we were not able to find such an effect with current models.
— Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
(2607.14345 - Betley et al., 15 Jul 2026) in Conclusion and future work