Papers
Topics
Authors
Recent
Search
2000 character limit reached

Algorithmic and Minimax Complexities in Kernel Bandits

Published 9 Jun 2026 in cs.LG, cond-mat.stat-mech, cs.IT, math.OC, and math.ST | (2606.11171v1)

Abstract: Gaussian-process upper confidence bound (GP-UCB) and decision-estimation-coefficient (DEC) methods may appear, at first sight, to belong to different theories. This paper places the two viewpoints in a common algorithmic-information language for frequentist RKHS bandits. GP-UCB fixes an algorithmic, rather than true, Gaussian-process prior and exploits realized-trajectory complexity together with computational tractability, whereas MAMS optimizes a robust class-wide MAIR/DEC envelope. Through the unified MAIR framework and heterogeneous positive-semidefinite algorithmic priors, we generalize both the GP-UCB analysis and the MAMS algorithm, propose a safeguarded master that combines their advantages, and provide a kernel-bandit construction showing that algorithmic complexity can be more informative than class-wide minimax or DEC certificates in overparameterized models. The resulting message is that algorithmic information and class-wide minimax coefficients answer different questions and can lead to different gaps; kernel bandits provide a clean setting in which this distinction becomes mathematically visible.

Authors (1)

Summary

  • The paper establishes that GP-UCB and minimax strategies are projections of the same regret-information identity unified by the MAIR framework.
  • It generalizes kernel bandits to include heterogeneous PSD priors, demonstrating how algorithmic bias shapes posterior variance and information gain.
  • The study proposes a safeguarded master algorithm that balances tractability and robustness, achieving lower regret in structured, overparameterized settings.

Algorithmic and Minimax Complexities in Kernel Bandits

Unified Framework for Kernel Bandits

The paper provides a rigorous synthesis of two major perspectives in sequential kernel-based bandit learning: the algorithmic-information viewpoint, exemplified by Gaussian Process Upper Confidence Bound (GP-UCB) algorithms, and the information-theoretic minimax viewpoint, delineated by Decision-Estimation Coefficient (DEC), Minimax Algorithmic Information Ratio (MAIR), and their algorithmic counterparts such as MAMS. By establishing a unified MAIR algebraic framework, the authors demonstrate that both GP-UCB and minimax algorithms are projections of the same underlying regret-information identity but answer fundamentally different statistical questions. The distinction between algorithmic complexity—actual information acquired during a run—and class-wide minimax complexity—worst-case performance over the full model class—becomes mathematically precise in the RKHS bandit setting.

Kernel Bandits with Heterogeneous Algorithmic Priors

The analysis generalizes kernel bandit models to include heterogeneous positive semi-definite (PSD) algorithmic priors beyond the scalar ridge, thereby exposing how algorithmic bias fundamentally influences the geometry of posterior variance, information gain, and code length. The posterior mean, variance, and realized algorithmic information gain for any design trajectory are computed under the selected precision matrix. The paper formally separates three performance scales: fixed-truth regret (regret for a specific function), Bayesian average risk (expected regret under the algorithmic prior), and class-wide minimax risk. The authors rigorously show that these scales can diverge dramatically in overparameterized kernels, with GP-UCB's realized information often much lower than minimax or DEC certificates.

The MAIR Objective: Algebraic Unification

Central to this unification is the MAIR objective, mathematically structured as

MAIRρ,η(p,μ)=EMμΔM(p)1ηEπpIμ(M;Oπ)1ηKL(μρ)MAIR_{\rho, \eta}(p, \mu) = \mathbb{E}_{M \sim \mu} \Delta_M(p) - \frac{1}{\eta} \mathbb{E}_{\pi \sim p} I_\mu(M; O | \pi) - \frac{1}{\eta} KL(\mu\|\rho)

where ΔM(p)\Delta_M(p) is the regret gap, IμI_\mu quantifies algorithmic information, and KLKL regularizes the belief relative to a reference. The central algebraic result shows that two design principles—calibration and optimism for GP-UCB, robust saddle-point optimization for MAMS—emerge as solutions to different projections of this objective. GP-UCB employs an algorithmic Gaussian prior for computational tractability, calibrating through self-normalized concentration, while MAMS optimizes minimax envelopes, utilizing finite covers and discretization to extend to the RKHS setting.

Algorithmic vs Class-Wide Complexity: Comparison and Separation

A key contribution is the construction and analysis of heterogeneous-prior hub–cloud models, explicitly demonstrating that algorithmic complexity can be substantially smaller than class-wide minimax complexity in overparameterized regimes. For structured truths aligned with the algorithmic prior, GP-UCB achieves strong bounds proportional to realized log determinant information gain, outperforming minimax certificates. Conversely, minimax or DEC certificates are inherently pessimistic in these cases, as shown via rigorous lower bounds and DEC computations for cloud arms. The separation persists even when using robust convex-hull references for DEC, formalizing that practical algorithms can achieve much smaller regret than minimax lower bounds for suitably chosen priors.

Safeguarded and Practical Algorithms

To reconcile practical tractability and robustness, the paper proposes a safeguarded master algorithm that combines a heterogeneous GP-UCB base (targeting structured truths via algorithmic prior) with a robust MAMS benchmarking component (insurance against worst-case truths). Standard bandit corralling protocols guarantee regret no worse than the minimax reference plus a small overhead, but allow adaptation to favorable subfamilies where algorithmic information dominates. The analysis specifies the trade-off between hub geometry (prior-favored direction) and cloud geometry (high-complexity directions), and enumerates the computational overhead and estimation complexity incurred by discretization in minimax approaches.

Implications, Theoretical and Practical

Mathematically, the paper clarifies the longstanding debate regarding GP-UCB's optimality. Previous advances, such as improved calibration via Hilbert-space self-normalized concentration or tight information accounting in Bayesian GP settings, are recast as targeting distinct axes of MAIR: either improving the coefficient (calibration/optimism) or tightening realized information. The analysis establishes that minimax guarantees remain the gold standard for robust, class-wide performance, but algorithmic-information guarantees are equally essential for understanding practical trajectories and inductive bias.

Practically, the results inform the design of kernel bandit algorithms for overparameterized or high-dimensional models, explicitly guiding the choice of prior precision as a vehicle for exploiting structure. Safeguarded algorithms are recommended for applications requiring insurance, while algorithmic-information certificates are preferable in settings where computational tractability and prior alignment are paramount.

Implications for Future AI Developments

Future developments in sequential kernel-based learning should leverage the pluralistic view articulated here. Designing algorithms that adaptively calibrate prior precision, perhaps using learned or meta-optimized priors, will likely enhance practical performance in structured domains. For overparameterized or high-variance settings, the safeguarded master architecture provides an avenue for combining algorithmic and minimax strengths. The mathematical distinctions in complexity measures may drive advancements in theoretical guarantees for robust inference, adaptive sampling, and online learning in AI.

Further, the MAIR identity and its algebraic bracket unify information-directed sampling, posterior sampling, and optimism-based exploration, providing fertile ground for generalizing to RL, structured bandits, and interactive learning.

Conclusion

The paper rigorously establishes that algorithmic and class-wide minimax complexities in kernel bandits are distinct axes of sequential learning, mathematically unified through the MAIR framework but answering fundamentally different statistical and computational questions. Heterogeneous algorithmic priors expose the inductive bias in GP-UCB, enabling strong performance on structured truths. Minimax and DEC certificates guarantee robustness but may overpay computationally in practical regimes. The proposed safeguarded algorithms and theoretical separation underscore the necessity of evaluating algorithmic information alongside minimax risk in kernel bandit optimality debates. Practical and theoretical advances in AI should harness these perspectives to develop tractable, robust, and adaptive learning algorithms.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.