Generality of the p=1/2 stability boundary

Determine whether the empirically observed training-stability boundary at exponent p=1/2 for the Post-LN looped Transformer residual scaling family α=(2N)^p and β=(8N)^{-p} persists when (i) increasing the number of physical blocks K, (ii) changing normalization placement (e.g., Pre-LN versus Post-LN), and (iii) substantially extending the training schedule length beyond the short runs used in the reported p-sweep.

Background

The paper derives a loop-aware perturbation bound showing that, under aligned visits at fixed physical depth, stability requires increasing the residual scaling exponent from DeepNorm’s p=1/4 to p=1/2. An empirical p-sweep at loop count R=3 on a GPT-2 small backbone finds a training-stability boundary near p=1/2, consistent with the theoretical prediction.

However, this sweep is conducted at a single model scale, a single loop count, and a limited training budget. The authors explicitly leave open whether the same p=1/2 stability boundary holds when scaling the number of physical blocks K, changing normalization placement, or training substantially longer, indicating the need to verify the boundary’s robustness across these dimensions.

References

Whether the same boundary at p=1/2 applies at larger K, at different normalization placements, or at substantially longer training, is left to future work and discussed in the Limitations paragraph (§\ref{sec:related}).

DeepLoop: Depth Scaling for Looped Transformers  (2607.13491 - Li et al., 15 Jul 2026) in Appendix, Section 'Empirical p-sweep at fixed loop count'