- The paper demonstrates that the CML estimator converges at a rate of n_*⁻¹/² (up to logarithmic factors), where n_* is the effective sample size determined by the spectral norm of the correlation matrix.
- It leverages a refined geometric Brascamp-Lieb inequality to control dependencies in Gaussian variables, leading to sharp Bernstein-type bounds for the log-likelihood process.
- Applications in high-dimensional linear and nonlinear regression validate the estimator's computational efficiency and robustness even under complex, dependent data scenarios.
Introduction and Context
Empirical Bayes (EB) methodology is widely utilized in large-scale inference tasks, particularly in settings with numerous parallel experiments. The foundational EB setting assumes independent Gaussian observations, under which the nonparametric maximum likelihood estimator (NPMLE) for the prior distribution enjoys well-characterized behavior, converging nearly at a parametric rate in mixture density estimation. However, the independence assumption rarely holds in applications: observational correlations frequently arise in high-dimensional regression, genomics, and other domains.
The paper "Empirical Bayes for correlated Gaussian sequence model" (2607.03596) rigorously addresses EB estimation when the observations are jointly Gaussian with arbitrary covariance structure. The key contribution is a sharp, non-asymptotic analysis of the maximum composite marginal likelihood (CML) estimator—i.e., the estimator that maximizes a likelihood which ignores correlations among coordinates—demonstrating that its convergence depends on the effective sample size, which is controlled by the spectral norm of the correlation matrix.
The core statistical model is
U=β0+Σ01/2Z,β0j∼i.i.d.G0,Z∼N(0,In),
where Σ0 is arbitrary, and G0 is the prior to be estimated. The CML estimator for G0 is defined by maximizing a pseudo-likelihood that treats the Uj as if independent:
Gn∈argG∈Gmaxn1j=1∑nlogφG;σ0,j(Uj),
where σ0,j2=(Σ0)jj and φG;σ(x) is the Gaussian location mixture.
Remarkably, the CML estimator remains computationally tractable even when the correlation structure is complex or high-dimensional. However, the theoretical justification for its efficacy in dependent-data regimes hinges on careful control of concentration and empirical processes with respect to the underlying dependence.
Main Theoretical Results
The central theoretical result is that the CML estimator converges at a rate determined by the "effective sample size" n∗:
n∗=∥CorΣ0∥opn,
where Σ00 is the correlation matrix corresponding to Σ01.
Specifically, for suitable classes of priors and moderate heteroscedasticity, the mixture density estimation error in averaged Hellinger distance satisfies:
Σ02
and under mild conditions, the Wasserstein error converges at the optimal logarithmic rate.
Notably, the paper provides a minimax lower bound showing that Σ03 is the correct complexity measure: the rate Σ04 (up to log factors) cannot be improved even in the absence of additional structure.
This demonstrates strong robustness of the CML approach under arbitrary Gaussian dependencies: as correlations increase, information content (i.e., Σ05) degrades, but the estimator remains rate-optimal for the information available.
Technical Innovations
The analysis departs from traditional empirical process theory for independent data by leveraging a refined geometric Brascamp-Lieb inequality for Gaussian measures [Chen et al., Adv. Math., 2015]. This inequality enables decoupling of exponential moments across correlated Gaussian variables, with the spectral norm of the correlation matrix controlling the decoupling penalty. This leads to a sharp Bernstein-type inequality and ultimately to a local maximal inequality for the log-likelihood process in the presence of arbitrary dependencies.
Applications: Regression and Beyond
Two substantive applications of the CML methodology are demonstrated:
- High-dimensional Bayesian Linear Regression: Applying CML to the generalized least squares estimator (GLS) yields practical and theoretical guarantees for prior estimation, even in challenging proportional regimes (Σ06). The method sidesteps computational and analytic intractabilities of full-likelihood or variational approaches by utilizing the approximately normal distribution of GLS estimators.
Figure 1: Performance of CML in linear regression (left: GLS) compared to debiased GD in nonlinear settings (right), across varying effective sample sizes.
- Nonlinear Regression (Single-Index Models): The CML method is used in conjunction with a one-step debiased gradient descent algorithm for nonlinear models, reducing prior estimation in general high-dimensional nonlinear regression to an approximate correlated Gaussian sequence problem.
Both applications showcase that, even when the full likelihood is computationally prohibitive or analytically opaque (as in nonlinear models), the CML-based approach provides non-asymptotic, effective procedures, with convergence determined by the effective sample size.
Practical Implications: Credible Intervals and Regret
The framework supports the construction of empirical Bayes credible intervals and controls their average frequentist coverage, even in the presence of strong dependence.
Figure 2: Coverage accuracy for nominal 90% empirical Bayes credible intervals and marginal regret across effective sample sizes Σ07.
Additionally, the marginal empirical Bayes regret with respect to the best marginal rule converges at the optimal parametric rate in Σ08, highlighting not just consistency but strong efficiency properties for downstream decision tasks under this general dependence structure.
Numerical Illustration
Extensive simulations confirm the theoretical predictions: average credible interval coverage and marginal regret improve as effective sample size increases, and the convergence rates align with the theoretical Σ09 and G00 scaling. For linear and nonlinear regression, empirical mixture density estimation errors closely follow predicted rates for varying dependence structures.
Implications and Future Directions
This work establishes that empirical Bayes procedures, particularly the CML estimator, remain effective for large-scale inference under arbitrary Gaussian dependence. The identification of effective sample size as a universal complexity parameter has broad implications for analysis of dependent high-dimensional data and justifies the use of marginal composite likelihoods in myriad applied settings.
The technical approach—leveraging geometric functional inequalities for dependent Gaussian ensembles—opens doors for further research in nonparametric inference under structured dependence, potentially extending to non-Gaussian or non-linear graphical models. Extensions to rapid posterior contraction and adaptation to unknown covariance structure are of particular interest. Additionally, the practical algorithmic simplicity of the method—requiring only one-dimensional optimization per coordinate—makes it attractive for large-scale modern applications in genomics, imaging, and network analysis.
Conclusion
"Empirical Bayes for correlated Gaussian sequence model" (2607.03596) provides a comprehensive, rigorous, and practically impactful framework for nonparametric prior estimation in Gaussian models with arbitrary correlations. Through precise characterizations of convergence rates and careful attention to dependence structure, the CML approach is shown to be robust, rate-optimal, and computationally efficient across a broad spectrum of high-dimensional inference problems.