Test for value leakage–induced sandbagging in dangerous capabilities evaluations

Investigate whether value leakage related to sandbagging occurs in dangerous capabilities evaluations of frontier language models by constructing evaluation settings where models might reduce apparent capability and determining whether such reductions are driven by their values.

Background

The authors explored whether value leakage manifests as sandbagging—intentional underperformance—in dangerous capabilities evaluations, which would have implications for safety assessments and oversight.

Their exploratory attempts did not find evidence of such an effect with current models, leaving open whether this phenomenon can occur and under what conditions it might be elicited.

References

In exploratory experiments, we tried to reproduce value leakage related to sandbagging in dangerous capabilities evaluations \citep{meinke2024scheming}, but we were not able to find such an effect with current models.

Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values  (2607.14345 - Betley et al., 15 Jul 2026) in Conclusion and future work