Papers
Topics
Authors
Recent
Search
2000 character limit reached

Heavy-Tailed Priors in Bayesian Modeling

Updated 10 July 2026
  • Heavy-tailed priors are probability distributions with polynomially decaying tails used to retain significant mass for large coefficients and extreme latent states.
  • They balance aggressive shrinkage near the origin with robustness to outliers, making them ideal for sparse estimation, multiple testing, and adaptive modeling.
  • Hierarchical and tail-adaptive constructions, such as global-local mixtures and horseshoe variants, enable both precise regularization and computational tractability.

Heavy-tailed priors are prior distributions whose tails decay polynomially, or more generally more slowly than Gaussian or exponential tails, and they are used when Bayesian procedures must retain substantial probability mass on large coefficients, large local scales, or extreme latent states. In the recent literature they appear in several mathematically distinct roles: as stable laws and infinite-dimensional Cauchy priors in inverse problems, as global-local shrinkage priors for sparse estimation and multiple testing, as heavy-tailed series priors for nonparametric adaptation, and as Student-tt priors in modern generative modeling (Sullivan, 2016, Polson et al., 2010, Agapiou et al., 2023, Pandey et al., 2024). A recurring theme is that heavy tails are valuable not simply because they are diffuse, but because they combine aggressive shrinkage or regularization near the origin with robustness to genuinely large signals, abrupt breaks, or rare events.

1. Tail behavior and probabilistic meaning

A precise formulation of heavy-tailedness in this literature is regular variation. One representative definition is

P(Z>x)=xαL(x),x>0,P(Z>x)=x^{-\alpha}L(x), \qquad x>0,

where L()L(\cdot) is slowly varying (Ghosh et al., 2024). In stable-prior formulations, if uStable(α,β,γ,0;0)u\sim \mathrm{Stable}(\alpha,\beta,\gamma,0;0) with 0<α<20<\alpha<2, then

E[up]={Cα,βγα<,0<p<α, ,pα,\mathbb E[|u|^p] = \begin{cases} C_{\alpha,\beta}\,\gamma^\alpha<\infty, & 0<p<\alpha,\ \infty, & p\ge \alpha, \end{cases}

and the tail satisfies

P[u>x]cαγα(1+β)xα,ρu(x)cααγα(1+β)x(α+1),\mathbb P[u>x]\sim c_\alpha \gamma^\alpha(1+\beta)x^{-\alpha}, \qquad \rho_u(x)\sim c_\alpha \alpha \gamma^\alpha(1+\beta)x^{-(\alpha+1)},

so a Cauchy prior is the α=1\alpha=1 case and has no finite first moment (Sullivan, 2016). This makes explicit that heavy-tailed priors are not merely “vague”: they alter moment existence, posterior integrability arguments, and the geometry of shrinkage.

Several families recur across the literature. Student-tt, Cauchy, half-Cauchy, generalized Pareto, Fréchet-class local-scale laws, hypergeometric inverted-beta laws, and stable distributions all appear as heavy-tailed priors or heavy-tailed mixing laws (Schmidt et al., 2018, Lee et al., 2020, Polson et al., 2010). In generative modeling, the multivariate Student-tt density is written as

P(Z>x)=xαL(x),x>0,P(Z>x)=x^{-\alpha}L(x), \qquad x>0,0

with P(Z>x)=xαL(x),x>0,P(Z>x)=x^{-\alpha}L(x), \qquad x>0,1 controlling tail heaviness and P(Z>x)=xαL(x),x>0,P(Z>x)=x^{-\alpha}L(x), \qquad x>0,2 recovering the Gaussian limit (Pandey et al., 2024). In shrinkage settings, the same basic tail mechanism is often studied indirectly through local scales or shrinkage coefficients rather than directly through the coefficient prior.

A common misconception is that heavy-tailed priors are interchangeable because they all “protect large signals.” The literature instead distinguishes between fixed-tail and tail-adaptive constructions, between priors on coefficients and priors on local scales, and between heavy tails used for robustness in likelihoods and heavy tails used for sparse regularization. This suggests that the phrase “heavy-tailed prior” names a design principle rather than a single canonical distribution.

2. Hierarchical constructions and canonical families

Much of the modern theory is built on hierarchical Gaussian mixtures. In global-local shrinkage form,

P(Z>x)=xαL(x),x>0,P(Z>x)=x^{-\alpha}L(x), \qquad x>0,3

with P(Z>x)=xαL(x),x>0,P(Z>x)=x^{-\alpha}L(x), \qquad x>0,4 global and P(Z>x)=xαL(x),x>0,P(Z>x)=x^{-\alpha}L(x), \qquad x>0,5 local (Schmidt et al., 2018). The heavy-tailed behavior is then induced by the law of P(Z>x)=xαL(x),x>0,P(Z>x)=x^{-\alpha}L(x), \qquad x>0,6, or equivalently by a prior on the shrinkage factor

P(Z>x)=xαL(x),x>0,P(Z>x)=x^{-\alpha}L(x), \qquad x>0,7

A particularly influential example is the hypergeometric-beta prior

P(Z>x)=xαL(x),x>0,P(Z>x)=x^{-\alpha}L(x), \qquad x>0,8

whose induced hypergeometric inverted-beta family includes Strawderman, half-Cauchy, and uniform-shrinkage special cases (Polson et al., 2010). This construction was proposed for large-scale simultaneous testing, where heavy tails and computational tractability were both central.

A second line of work uses the log local scale P(Z>x)=xαL(x),x>0,P(Z>x)=x^{-\alpha}L(x), \qquad x>0,9 as the primitive object. In that parameterization, the log-L()L(\cdot)0 prior is

L()L(\cdot)1

so the prior is essentially L()L(\cdot)2 multiplied by a slowly varying factor (Schmidt et al., 2018). Its importance lies in two invariant properties established in that work: irrespective of L()L(\cdot)3, the induced marginal prior on L()L(\cdot)4 diverges at zero and has super-Cauchy tails.

A third construction sharpens horseshoe-type behavior by mixing over shape parameters. The heavy-tailed horseshoe prior is defined through

L()L(\cdot)5

which yields the marginal

L()L(\cdot)6

(Womack et al., 2019). The paper’s central mathematical point is that this produces a prior that is more singular at zero and more robust in the tails than the standard horseshoe and horseshoe+.

Tail-adaptive constructions make the tail index itself random. The global-local-tail prior uses

L()L(\cdot)7

with generalized Pareto density

L()L(\cdot)8

(Lee et al., 2020). Here L()L(\cdot)9 directly controls regular variation. This is the clearest formalization of the idea that tail-heaviness itself can be learned rather than fixed.

3. Sparse estimation, multiple testing, and adaptive shrinkage

The strongest empirical and theoretical case for heavy-tailed priors has been made in sparse estimation. In large-scale simultaneous testing, the hypergeometric inverted-beta family was introduced precisely because heavy tails induce a mild rate of tail decay in the marginal likelihood uStable(α,β,γ,0;0)u\sim \mathrm{Stable}(\alpha,\beta,\gamma,0;0)0, a property that the paper identifies as important in testing (Polson et al., 2010). In the application to return on assets for uStable(α,β,γ,0;0)u\sim \mathrm{Stable}(\alpha,\beta,\gamma,0;0)1 publicly traded firms across uStable(α,β,γ,0;0)u\sim \mathrm{Stable}(\alpha,\beta,\gamma,0;0)2 countries, the model concluded that demonstrably superior performance was rare after multiplicity adjustment and longitudinal dependence correction (Polson et al., 2010). This is a prototypical example of heavy-tailed priors being used to distinguish large genuine signals from many null effects.

The heavy-tailed horseshoe develops the same theme in normal means estimation. For the induced marginal prior on uStable(α,β,γ,0;0)u\sim \mathrm{Stable}(\alpha,\beta,\gamma,0;0)3, the paper derives the tail comparison

uStable(α,β,γ,0;0)u\sim \mathrm{Stable}(\alpha,\beta,\gamma,0;0)4

showing that shape-parameter mixing changes the asymptotic tail order rather than merely adding a slowly varying correction (Womack et al., 2019). In simulations with uStable(α,β,γ,0;0)u\sim \mathrm{Stable}(\alpha,\beta,\gamma,0;0)5, the heavy-tailed horseshoe family improved both mean absolute error and distance to the oracle estimator, especially as sparsity became more severe (Womack et al., 2019).

The global-local-tail prior pushes this argument further by questioning fixed-tail defaults. In the sparse normal means model, the paper shows that the posterior contracts at the minimax optimal rate, and it derives the heuristic relation

uStable(α,β,γ,0;0)u\sim \mathrm{Stable}(\alpha,\beta,\gamma,0;0)6

so the tail-heaviness should increase as the regime becomes less sparse (Lee et al., 2020). This is one of the main conceptual corrections to the earlier “always use the horseshoe” viewpoint: a fixed tail rule may be appropriate in ultra-sparse regimes but can be undesirable under moderate sparsity.

Heavy-tailed priors also support feature selection in generalized linear models. A fully Bayesian Robit regression with a small-scale Cauchy prior

uStable(α,β,γ,0;0)u\sim \mathrm{Stable}(\alpha,\beta,\gamma,0;0)7

was designed to select high-dimensional features with grouping structure, while allowing automatic within-group selection without a pre-specified grouping structure (Jiang et al., 2016). The posterior is explicitly multimodal in correlated designs, and the paper interprets these modes as alternative sparse representative subsets rather than as a pathology. A related logistic-regression study reported that heavy-tailed priors with moderately small degree freedom and very small scale, combined with Hamiltonian Monte Carlo, identified only uStable(α,β,γ,0;0)u\sim \mathrm{Stable}(\alpha,\beta,\gamma,0;0)8 non-redundant genes out of uStable(α,β,γ,0;0)u\sim \mathrm{Stable}(\alpha,\beta,\gamma,0;0)9 candidates while achieving better leave-one-out cross-validated prediction accuracy than many other methods (Li et al., 2013).

These examples clarify the central shrinkage mechanism. Heavy-tailed priors do not merely reduce bias on large coefficients; they also reshape the posterior so that one or a few coordinates can escape while many others remain near zero. This suggests that the practical effect of heavy tails depends as much on the near-zero geometry and the hierarchy of local scales as on the upper-tail decay rate itself.

4. Infinite-dimensional, nonparametric, and deep-network formulations

In inverse problems and function-space Bayes, heavy-tailed priors are used to replace Gaussian-type regularity assumptions. One construction on a quasi-Banach space 0<α<20<\alpha<20 with unconditional basis 0<α<20<\alpha<21 defines

0<α<20<\alpha<22

and proves almost-sure convergence under 0<α<20<\alpha<23, 0<α<20<\alpha<24, and the borderline condition

0<α<20<\alpha<25

(Sullivan, 2016). The same paper shows that the posterior depends Lipschitz continuously in the Hellinger metric upon perturbations of the misfit function and observed data, under assumptions adapted to heavy tails rather than Gaussian exponential integrability (Sullivan, 2016).

A different nonparametric line uses heavy-tailed series priors for adaptation to smoothness. The oversmoothed heavy-tailed prior is defined by

0<α<20<\alpha<26

or more generally 0<α<20<\alpha<27, while the coordinate law 0<α<20<\alpha<28 satisfies a heavy-tail condition such as

0<α<20<\alpha<29

(Agapiou et al., 2023). The paper interprets the resulting mechanism as “soft selection through the prior tail”: there is no random truncation and no sampled smoothness hyperparameter, but the posterior still achieves adaptive rates in Gaussian regression, inverse problems, Besov settings, and tempered posterior formulations (Agapiou et al., 2023). The subsequent nonparametric analysis of oversmoothed heavy-tailed priors and horseshoe priors shows that

E[up]={Cα,βγα<,0<p<α, ,pα,\mathbb E[|u|^p] = \begin{cases} C_{\alpha,\beta}\,\gamma^\alpha<\infty, & 0<p<\alpha,\ \infty, & p\ge \alpha, \end{cases}0

is not merely convenient but necessary for full adaptation to smoothness, and it derives minimax posterior contraction rates even in the sparse Besov zone (Agapiou et al., 21 May 2025).

The same principle has now been transferred to Bayesian deep learning. In a ReLU-network prior with deterministic overparameterized architecture, each parameter is drawn as

E[up]={Cα,βγα<,0<p<α, ,pα,\mathbb E[|u|^p] = \begin{cases} C_{\alpha,\beta}\,\gamma^\alpha<\infty, & 0<p<\alpha,\ \infty, & p\ge \alpha, \end{cases}1

with heavy-tailed E[up]={Cα,βγα<,0<p<α, ,pα,\mathbb E[|u|^p] = \begin{cases} C_{\alpha,\beta}\,\gamma^\alpha<\infty, & 0<p<\alpha,\ \infty, & p\ge \alpha, \end{cases}2 and very small scales E[up]={Cα,βγα<,0<p<α, ,pα,\mathbb E[|u|^p] = \begin{cases} C_{\alpha,\beta}\,\gamma^\alpha<\infty, & 0<p<\alpha,\ \infty, & p\ge \alpha, \end{cases}3 (Castillo et al., 2024). The resulting posterior achieves near-optimal minimax contraction rates, simultaneously adaptive to both intrinsic dimension and smoothness, and the paper proves variational Bayes analogues for a mean-field heavy-tailed family (Castillo et al., 2024). A plausible implication is that heavy-tailed priors can replace hard architecture selection by a deterministic overfitting architecture plus soft sparsity at the weight level.

5. Dynamic, inverse, and generative modeling

Heavy-tailed priors are also used outside classical sparse regression. In time-varying parameter state-space models, one construction imposes a Dirichlet-Laplace prior on

E[up]={Cα,βγα<,0<p<α, ,pα,\mathbb E[|u|^p] = \begin{cases} C_{\alpha,\beta}\,\gamma^\alpha<\infty, & 0<p<\alpha,\ \infty, & p\ge \alpha, \end{cases}4

for both static coefficients and E[up]={Cα,βγα<,0<p<α, ,pα,\mathbb E[|u|^p] = \begin{cases} C_{\alpha,\beta}\,\gamma^\alpha<\infty, & 0<p<\alpha,\ \infty, & p\ge \alpha, \end{cases}5, the transformed state innovation standard deviations (Huber et al., 2018). The same model combines that heavy-tailed shrinkage prior with Student-E[up]={Cα,βγα<,0<p<α, ,pα,\mathbb E[|u|^p] = \begin{cases} C_{\alpha,\beta}\,\gamma^\alpha<\infty, & 0<p<\alpha,\ \infty, & p\ge \alpha, \end{cases}6 state innovations and Student-E[up]={Cα,βγα<,0<p<α, ,pα,\mathbb E[|u|^p] = \begin{cases} C_{\alpha,\beta}\,\gamma^\alpha<\infty, & 0<p<\alpha,\ \infty, & p\ge \alpha, \end{cases}7 observation disturbances, so that shrinkage favors constant coefficients while local heavy-tailed innovations still allow abrupt breaks (Huber et al., 2018). This shows a second major role for heavy tails: not only shrinkage, but also robust latent dynamics.

In inverse-mean problems, heavy-tailed priors are used to slow posterior-risk deterioration under nonlinear inversion. For parallax-based distance estimation, the paper studies reciprocal invariant priors such as the Half-Cauchy and Product Half-Cauchy, with the latter

E[up]={Cα,βγα<,0<p<α, ,pα,\mathbb E[|u|^p] = \begin{cases} C_{\alpha,\beta}\,\gamma^\alpha<\infty, & 0<p<\alpha,\ \infty, & p\ge \alpha, \end{cases}8

having tail E[up]={Cα,βγα<,0<p<α, ,pα,\mathbb E[|u|^p] = \begin{cases} C_{\alpha,\beta}\,\gamma^\alpha<\infty, & 0<p<\alpha,\ \infty, & p\ge \alpha, \end{cases}9 (Ghosh et al., 2024). The main conclusion is deliberately limited: heavy-tailed priors can delay the explosion of posterior risk but cannot eliminate the “curse of a single observation” (Ghosh et al., 2024). This is an important reminder that heavy tails are robustifiers, not universal cures.

In deep generative modeling, prior choice becomes part of the forward noising mechanism. Heavy-tailed diffusion replaces Gaussian corruption by

P[u>x]cαγα(1+β)xα,ρu(x)cααγα(1+β)x(α+1),\mathbb P[u>x]\sim c_\alpha \gamma^\alpha(1+\beta)x^{-\alpha}, \qquad \rho_u(x)\sim c_\alpha \alpha \gamma^\alpha(1+\beta)x^{-(\alpha+1)},0

so the terminal prior is Student-P[u>x]cαγα(1+β)xα,ρu(x)cααγα(1+β)x(α+1),\mathbb P[u>x]\sim c_\alpha \gamma^\alpha(1+\beta)x^{-\alpha}, \qquad \rho_u(x)\sim c_\alpha \alpha \gamma^\alpha(1+\beta)x^{-(\alpha+1)},1 rather than Gaussian (Pandey et al., 2024). The resulting P[u>x]cαγα(1+β)xα,ρu(x)cααγα(1+β)x(α+1),\mathbb P[u>x]\sim c_\alpha \gamma^\alpha(1+\beta)x^{-\alpha}, \qquad \rho_u(x)\sim c_\alpha \alpha \gamma^\alpha(1+\beta)x^{-(\alpha+1)},2-EDM and P[u>x]cαγα(1+β)xα,ρu(x)cααγα(1+β)x(α+1),\mathbb P[u>x]\sim c_\alpha \gamma^\alpha(1+\beta)x^{-\alpha}, \qquad \rho_u(x)\sim c_\alpha \alpha \gamma^\alpha(1+\beta)x^{-(\alpha+1)},3-Flow models were reported to outperform standard diffusion models in heavy-tail estimation on high-resolution weather datasets where rare and extreme events are crucial (Pandey et al., 2024). A closely related constrained-generation paper uses a Student-P[u>x]cαγα(1+β)xα,ρu(x)cααγα(1+β)x(α+1),\mathbb P[u>x]\sim c_\alpha \gamma^\alpha(1+\beta)x^{-\alpha}, \qquad \rho_u(x)\sim c_\alpha \alpha \gamma^\alpha(1+\beta)x^{-(\alpha+1)},4 prior in the dual space of mirror flow matching,

P[u>x]cαγα(1+β)xα,ρu(x)cααγα(1+β)x(α+1),\mathbb P[u>x]\sim c_\alpha \gamma^\alpha(1+\beta)x^{-\alpha}, \qquad \rho_u(x)\sim c_\alpha \alpha \gamma^\alpha(1+\beta)x^{-(\alpha+1)},5

paired with a regularized mirror map, and proves spatial Lipschitzness, temporal regularity, Wasserstein convergence rates, and primal-space guarantees for constrained generation (Guan et al., 10 Oct 2025).

Heavy-tailed priors have even been imported into spectral regularization for deep neural networks. A Bayesian version of heavy-tailed regularization places a power-law prior on the spectrum,

P[u>x]cαγα(1+β)xα,ρu(x)cααγα(1+β)x(α+1),\mathbb P[u>x]\sim c_\alpha \gamma^\alpha(1+\beta)x^{-\alpha}, \qquad \rho_u(x)\sim c_\alpha \alpha \gamma^\alpha(1+\beta)x^{-(\alpha+1)},6

or a Fréchet-inspired prior on the maximum eigenvalue,

P[u>x]cαγα(1+β)xα,ρu(x)cααγα(1+β)x(α+1),\mathbb P[u>x]\sim c_\alpha \gamma^\alpha(1+\beta)x^{-\alpha}, \qquad \rho_u(x)\sim c_\alpha \alpha \gamma^\alpha(1+\beta)x^{-(\alpha+1)},7

to promote heavy-tailed weight spectra associated with good generalization (Xiao et al., 2023). This extends the concept of heavy-tailed priors from parameter shrinkage to spectral-shape regularization.

6. Limitations, controversies, and design principles

Several limitations recur across the literature. First, heavy-tailed priors rarely produce exact sparsity unless combined with discrete model indicators or thresholding rules. In fully Bayesian feature selection with Cauchy priors, selected subsets are extracted by post-processing posterior draws rather than read off as exact zeros (Jiang et al., 2016). Second, heavy tails often create posterior geometries with multimodality, long ridges, or difficult local-scale dependencies, so computation typically requires carefully designed MCMC or variational procedures rather than naive Gibbs updates (Jiang et al., 2016, Schmidt et al., 2018, Womack et al., 2019).

Third, the strongest adaptation results often rely on tempered or fractional posteriors. Heavy-tailed series priors in nonparametric models and heavy-tailed Bayesian deep networks both obtain their cleanest contraction theory under P[u>x]cαγα(1+β)xα,ρu(x)cααγα(1+β)x(α+1),\mathbb P[u>x]\sim c_\alpha \gamma^\alpha(1+\beta)x^{-\alpha}, \qquad \rho_u(x)\sim c_\alpha \alpha \gamma^\alpha(1+\beta)x^{-(\alpha+1)},8-posteriors or fractional posteriors rather than only under the ordinary posterior (Agapiou et al., 2023, Castillo et al., 2024). This does not diminish the statistical value of the priors, but it does delimit the current scope of the theory.

Fourth, there is no single optimal tail specification across regimes. The global-local-tail analysis argues explicitly that the amount of tail-heaviness should adapt to the unknown sparsity level, while fixed-tail priors such as the horseshoe can be excellent in ultra-sparse problems yet undesirable in more moderate sparse regimes (Lee et al., 2020). A common misconception is therefore that “heavier tails are always better.” The more accurate statement is that heavy-tailed priors are useful when their tail geometry matches the scientific structure of the problem: sparse large signals, boundary-induced dual heavy tails, extreme latent states, or polynomial-tail targets.

A final design principle emerges from the combined evidence. Heavy-tailed priors can be placed on coefficients, local scales, shrinkage coefficients, variance components, state innovations, basis coefficients, spectral summaries, or generative latent variables. Their role is best understood through the balance they induce between concentration near zero, accommodation of extremes, and the computational tractability of the resulting posterior. This suggests that heavy-tailed priors are not a niche alternative to Gaussian defaults, but a broad modeling strategy for settings in which the relevant Bayesian geometry is sparse, abrupt, edge-preserving, or genuinely heavy-tailed.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Heavy-Tailed Priors.