Papers
Topics
Authors
Recent
Search
2000 character limit reached

Empirical Bayes for correlated Gaussian sequence model

Published 3 Jul 2026 in math.ST, stat.ME, and stat.ML | (2607.03596v1)

Abstract: Empirical Bayes methods are among the most widely used statistical methods for large-scale inference. A central paradigm is the NPMLE, whose theoretical guarantees are by now well understood for the independent Gaussian sequence model. In this paper, we study empirical Bayes estimation from dependent observations in the Gaussian sequence model. We show that the maximum Composite Marginal Likelihood (CML) estimator, which ignores all correlations in the likelihood, converges in weighted Hellinger distance at the rate $n_{-1/2}$, where $n_=n/κ0$ is the `effective sample size' determined solely by the number of observations $n$ and the spectral radius $κ_0$ of the correlation matrix of the Gaussian observations. A complementary minimax lower bound shows that $n*$ indeed serves as the right complexity measure, and that the CML estimator is nearly rate optimal under general dependence. We consider two concrete applications. In the first, we consider Bayesian linear regression, where the signal prior is estimated via CML applied to the least squares estimator. In the second, we consider the more challenging Bayesian nonlinear single-index model, where prior is estimated by CML applied to a one-step debiased gradient descent. In both applications, although the full likelihood landscape can be arbitrarily complicated and intractable, our CML method is facilitated by exploiting the high-dimensional distribution of the auxiliary statistics through a correlated Gaussian sequence model. The key ingredient in the proof of our results is a sharp local maximal inequality for the log composite marginal likelihood process under dependent Gaussian observations. In contrast to standard empirical process methods, we prove this inequality by leveraging a recent geometric Brascamp-Lieb inequality for Gaussian measures.

Authors (2)

Summary

  • The paper demonstrates that the CML estimator converges at a rate of n_*⁻¹/² (up to logarithmic factors), where n_* is the effective sample size determined by the spectral norm of the correlation matrix.
  • It leverages a refined geometric Brascamp-Lieb inequality to control dependencies in Gaussian variables, leading to sharp Bernstein-type bounds for the log-likelihood process.
  • Applications in high-dimensional linear and nonlinear regression validate the estimator's computational efficiency and robustness even under complex, dependent data scenarios.

Empirical Bayes Estimation under Correlated Gaussian Sequence Models

Introduction and Context

Empirical Bayes (EB) methodology is widely utilized in large-scale inference tasks, particularly in settings with numerous parallel experiments. The foundational EB setting assumes independent Gaussian observations, under which the nonparametric maximum likelihood estimator (NPMLE) for the prior distribution enjoys well-characterized behavior, converging nearly at a parametric rate in mixture density estimation. However, the independence assumption rarely holds in applications: observational correlations frequently arise in high-dimensional regression, genomics, and other domains.

The paper "Empirical Bayes for correlated Gaussian sequence model" (2607.03596) rigorously addresses EB estimation when the observations are jointly Gaussian with arbitrary covariance structure. The key contribution is a sharp, non-asymptotic analysis of the maximum composite marginal likelihood (CML) estimator—i.e., the estimator that maximizes a likelihood which ignores correlations among coordinates—demonstrating that its convergence depends on the effective sample size, which is controlled by the spectral norm of the correlation matrix.

Formal Model and CML Procedure

The core statistical model is

U=β0+Σ01/2Z,β0ji.i.d.G0,ZN(0,In),U = \beta_0 + \Sigma_0^{1/2} Z, \quad \beta_{0j} \overset{\mathrm{i.i.d.}}{\sim} G_0, \quad Z \sim \mathcal{N}(0, I_n),

where Σ0\Sigma_0 is arbitrary, and G0G_0 is the prior to be estimated. The CML estimator for G0G_0 is defined by maximizing a pseudo-likelihood that treats the UjU_j as if independent:

GnargmaxGG1nj=1nlogφG;σ0,j(Uj),G_n \in \arg\max_{G \in \mathscr{G}} \frac{1}{n} \sum_{j=1}^n \log \varphi_{G;\sigma_{0,j}}(U_j),

where σ0,j2=(Σ0)jj\sigma_{0,j}^2 = (\Sigma_0)_{jj} and φG;σ(x)\varphi_{G;\sigma}(x) is the Gaussian location mixture.

Remarkably, the CML estimator remains computationally tractable even when the correlation structure is complex or high-dimensional. However, the theoretical justification for its efficacy in dependent-data regimes hinges on careful control of concentration and empirical processes with respect to the underlying dependence.

Main Theoretical Results

The central theoretical result is that the CML estimator converges at a rate determined by the "effective sample size" nn_*:

n=nCorΣ0op,n_* = \frac{n}{\| \mathrm{Cor}_{\Sigma_0} \|_{\mathrm{op}}},

where Σ0\Sigma_00 is the correlation matrix corresponding to Σ0\Sigma_01.

Specifically, for suitable classes of priors and moderate heteroscedasticity, the mixture density estimation error in averaged Hellinger distance satisfies:

Σ0\Sigma_02

and under mild conditions, the Wasserstein error converges at the optimal logarithmic rate.

Notably, the paper provides a minimax lower bound showing that Σ0\Sigma_03 is the correct complexity measure: the rate Σ0\Sigma_04 (up to log factors) cannot be improved even in the absence of additional structure.

This demonstrates strong robustness of the CML approach under arbitrary Gaussian dependencies: as correlations increase, information content (i.e., Σ0\Sigma_05) degrades, but the estimator remains rate-optimal for the information available.

Technical Innovations

The analysis departs from traditional empirical process theory for independent data by leveraging a refined geometric Brascamp-Lieb inequality for Gaussian measures [Chen et al., Adv. Math., 2015]. This inequality enables decoupling of exponential moments across correlated Gaussian variables, with the spectral norm of the correlation matrix controlling the decoupling penalty. This leads to a sharp Bernstein-type inequality and ultimately to a local maximal inequality for the log-likelihood process in the presence of arbitrary dependencies.

Applications: Regression and Beyond

Two substantive applications of the CML methodology are demonstrated:

  1. High-dimensional Bayesian Linear Regression: Applying CML to the generalized least squares estimator (GLS) yields practical and theoretical guarantees for prior estimation, even in challenging proportional regimes (Σ0\Sigma_06). The method sidesteps computational and analytic intractabilities of full-likelihood or variational approaches by utilizing the approximately normal distribution of GLS estimators. Figure 1

    Figure 1: Performance of CML in linear regression (left: GLS) compared to debiased GD in nonlinear settings (right), across varying effective sample sizes.

  2. Nonlinear Regression (Single-Index Models): The CML method is used in conjunction with a one-step debiased gradient descent algorithm for nonlinear models, reducing prior estimation in general high-dimensional nonlinear regression to an approximate correlated Gaussian sequence problem.

Both applications showcase that, even when the full likelihood is computationally prohibitive or analytically opaque (as in nonlinear models), the CML-based approach provides non-asymptotic, effective procedures, with convergence determined by the effective sample size.

Practical Implications: Credible Intervals and Regret

The framework supports the construction of empirical Bayes credible intervals and controls their average frequentist coverage, even in the presence of strong dependence. Figure 2

Figure 2: Coverage accuracy for nominal 90% empirical Bayes credible intervals and marginal regret across effective sample sizes Σ0\Sigma_07.

Additionally, the marginal empirical Bayes regret with respect to the best marginal rule converges at the optimal parametric rate in Σ0\Sigma_08, highlighting not just consistency but strong efficiency properties for downstream decision tasks under this general dependence structure.

Numerical Illustration

Extensive simulations confirm the theoretical predictions: average credible interval coverage and marginal regret improve as effective sample size increases, and the convergence rates align with the theoretical Σ0\Sigma_09 and G0G_00 scaling. For linear and nonlinear regression, empirical mixture density estimation errors closely follow predicted rates for varying dependence structures.

Implications and Future Directions

This work establishes that empirical Bayes procedures, particularly the CML estimator, remain effective for large-scale inference under arbitrary Gaussian dependence. The identification of effective sample size as a universal complexity parameter has broad implications for analysis of dependent high-dimensional data and justifies the use of marginal composite likelihoods in myriad applied settings.

The technical approach—leveraging geometric functional inequalities for dependent Gaussian ensembles—opens doors for further research in nonparametric inference under structured dependence, potentially extending to non-Gaussian or non-linear graphical models. Extensions to rapid posterior contraction and adaptation to unknown covariance structure are of particular interest. Additionally, the practical algorithmic simplicity of the method—requiring only one-dimensional optimization per coordinate—makes it attractive for large-scale modern applications in genomics, imaging, and network analysis.

Conclusion

"Empirical Bayes for correlated Gaussian sequence model" (2607.03596) provides a comprehensive, rigorous, and practically impactful framework for nonparametric prior estimation in Gaussian models with arbitrary correlations. Through precise characterizations of convergence rates and careful attention to dependence structure, the CML approach is shown to be robust, rate-optimal, and computationally efficient across a broad spectrum of high-dimensional inference problems.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.