Papers
Topics
Authors
Recent
Search
2000 character limit reached

Aggregation with Exponential Weights is Optimal in Expectation

Published 2 Jul 2026 in math.ST, cs.LG, and stat.ML | (2607.02247v1)

Abstract: The aggregation with exponential weights (AEW) estimator is not fully understood in the basic setting of model selection aggregation with squared loss. In particular, whether it is minimax-rate optimal in expectation for large enough fixed temperatures and under random design has been an open problem since its introduction, which was explicitly posed by Lecué and Mendelson (2013). In this paper, we settle this problem by showing that \emph{without} requiring a Bernstein-type assumption, the AEW indeed achieves the excess risk $T \log (M) / (n+1)$ in expectation, whenever the temperature $T$ satisfies $(L2/T)\exp(B/T)\leq μ/2$. Here, the number of dictionary elements is $M$, the estimator has observed $n$ i.i.d. samples from any distribution, and the loss is assumed to be bounded by $B$, $L$-Lipschitz continuous and $μ$-strongly convex. For squared loss, we show that $T\geq 4 b2$ suffices when the predictions and labels are $[0,b]$-valued. Because AEW is known to be suboptimal in expectation for temperatures below some constant, this shows that AEW has a sharp phase transition when the temperature is large enough but constant, as conjectured by Lecué and Mendelson.

Summary

  • The paper demonstrates that AEW achieves minimax-optimal excess risk in expectation by tuning the temperature parameter under mild loss assumptions.
  • It employs a novel deterministic leave-one-out stability analysis via tilt inequalities to match lower bound rates with tight constants.
  • The study reveals a phase transition in temperature, showing AEW’s suboptimality at extreme values and guiding effective hyperparameter tuning.

Exponential Weights Aggregation: Minimax-Optimality in Expectation

Problem Formulation

The paper "Aggregation with Exponential Weights is Optimal in Expectation" (2607.02247) addresses the model selection aggregation problem in statistical learning, focusing on the estimator constructed via aggregation with exponential weights (AEW). Given a finite function dictionary F={f1,,fM}F = \{f_1, \dots, f_M\} and i.i.d. samples (Xi,Yi)i=1n(X_i, Y_i)_{i=1}^n from an arbitrary distribution PP over X×YX \times Y, the goal is to select or aggregate functions to minimize the excess risk

RP(f)minkMRP(fk),R_P(f) - \min_{k \leq M} R_P(f_k),

where RP(f)=E(X,Y)P[(f(X),Y)]R_P(f) = \mathbb{E}_{(X,Y) \sim P}[\ell(f(X), Y)] with a loss \ell satisfying boundedness, Lipschitz continuity, and strong convexity assumptions.

The AEW estimator constructs weights for each function fkf_k: wk=exp(nTR^S(fk))j=1Mexp(nTR^S(fj)) ,w_k = \frac{\exp(-\frac{n}{T} \hat{R}_S(f_k))}{\sum_{j=1}^M \exp(-\frac{n}{T} \hat{R}_S(f_j))}\ , where R^S(fk)\hat{R}_S(f_k) is the empirical risk, and (Xi,Yi)i=1n(X_i, Y_i)_{i=1}^n0 is a "temperature" parameter controlling the smoothness of the weighting. Prior literature established minimax-optimality only for certain temperature regimes or under additional Bernstein-type conditions on the design distribution. The question of whether AEW attains minimax rates in expectation for sufficiently large, fixed temperatures, without a Bernstein assumption, has remained open.

Main Contributions and Theoretical Results

The core result of this paper is an affirmative resolution of the aforementioned question, establishing minimax-optimal rates in expectation for AEW when the temperature parameter (Xi,Yi)i=1n(X_i, Y_i)_{i=1}^n1 is taken sufficiently large and constant, under only general structural assumptions on the loss function. Specifically, the following claims are rigorously proved:

  • Minimax-Rate Optimality Under Mild Conditions: For loss (Xi,Yi)i=1n(X_i, Y_i)_{i=1}^n2 that is (Xi,Yi)i=1n(X_i, Y_i)_{i=1}^n3-bounded, (Xi,Yi)i=1n(X_i, Y_i)_{i=1}^n4-Lipschitz, and (Xi,Yi)i=1n(X_i, Y_i)_{i=1}^n5-strongly convex, AEW achieves excess risk

(Xi,Yi)i=1n(X_i, Y_i)_{i=1}^n6

for any (Xi,Yi)i=1n(X_i, Y_i)_{i=1}^n7 such that (Xi,Yi)i=1n(X_i, Y_i)_{i=1}^n8. For squared loss on (Xi,Yi)i=1n(X_i, Y_i)_{i=1}^n9, the sufficient condition is PP0.

  • No Bernstein Condition Required: These guarantees hold uniformly over all distributions PP1 and dictionaries PP2, independent of any margin or noise conditions, and for random design.
  • Phase Transition on Temperature: The AEW estimator exhibits a sharp threshold phenomenon in terms of PP3. For PP4 below a critical constant, AEW is strictly suboptimal in expectation and probability. For PP5 above this threshold but fixed, it achieves minimax rates; if PP6 grows with PP7, minimaxity is lost again.
  • Tightness of Constants: For large dictionaries (PP8), both upper and lower bounds become tight, up to first-order constants.
  • Negative Results for Large and Small PP9: For X×YX \times Y0 that grows without bound as X×YX \times Y1, or X×YX \times Y2 that is too small, precise constructions are given to show that excess risk is well above the minimax rate, formalizing regions of suboptimality.

Technical Methods

The key novelty lies in a deterministic leave-one-out stability analysis for AEW, based on an average of empirical risk minimizers with one example omitted. Central to the proof is a pair of tilt inequalities (provided for both general X×YX \times Y3 and the special case of squared loss), establishing that the effect of tilting the exponential weights, after removing one point, maintains control over the loss in a sufficiently tight way. This approach enables bounds in expectation that match the known lower bound rates, achieving constant-factor minimax optimality. The proof avoids the Jensen inequality-based approach ubiquitous in previous PAC-Bayes and ERM analyses, which incurs an unavoidable looseness.

The paper further deploys explicit constructions to establish lower bounds for the excess risk in both the high- and low-temperature regimes, demonstrating that the conditions ensuring optimality are tight.

Numerical Constants and Sharpness

For the squared loss on X×YX \times Y4, it is shown that X×YX \times Y5 is sufficient for optimality, improving on generic bounds involving the Lambert X×YX \times Y6 function. Setting X×YX \times Y7 for X×YX \times Y8-valued labels and function outputs, AEW achieves excess risk at most X×YX \times Y9. These constants are shown to be unimprovable in the first order for large RP(f)minkMRP(fk),R_P(f) - \min_{k \leq M} R_P(f_k),0.

Implications and Future Directions

Theoretical Significance: This work settles a foundational question in aggregation theory, clarifying that AEW, a canonical and widely studied PAC-Bayesian estimator, is minimax-optimal in expectation over all random designs and finite dictionaries for appropriate temperature settings, without restrictive additional conditions. The identification of the critical temperature regime and the precise constants involved provides clarity on how to tune AEW in practice and evaluates the robustness of exponential weights.

Practical Impact: For practitioners using AEW in model selection, especially in non-parametric or adversarial settings, these results now imply that as long as the (data-independent) temperature is set above the critical constant, the procedure is guaranteed to realize the best possible excess risk up to constants dependent only on the loss class. This removes a significant theoretical caveat from previous analyses and informs hyperparameter tuning in implementations.

Methodological Contribution: The deterministic stability argument via tilt inequalities may have further utility in the analysis of aggregation schemes, PAC-Bayesian risk bounds, and ERM stability properties. Its generalization to other losses opens investigation into when such tightness can or cannot be preserved for non-strongly convex or non-Lipschitz losses.

Future Directions: Possible extensions include the analysis of continuous dictionaries, adaptation to data-dependent temperature schemes, and applications to aggregation under non-convex or unbounded loss settings. Another avenue is extending these methods to facilitate high-probability excess risk guarantees, as the current minimax-optimality is achieved in expectation only, and AEW remains suboptimal in probability under the same conditions.

Conclusion

This paper resolves a longstanding open problem by establishing that aggregation with exponential weights achieves the minimax-optimal excess risk in expectation for sufficiently large, fixed temperatures under broad loss assumptions, and without the need for a Bernstein-type margin condition. The results delineate precisely the phase transition for temperature thresholds, elucidate the tightness and limitations of AEW in both low- and high-temperature regimes, and introduce analytic techniques of broad utility for PAC-Bayesian stability analysis (2607.02247).

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.