Papers
Topics
Authors
Recent
Search
2000 character limit reached

Path-Measure Dynamics of Attention-Driven World Models: A Nonlocal Onsager--Machlup Approach

Published 2 Jul 2026 in cond-mat.stat-mech | (2607.02154v1)

Abstract: Attention enables a world model to condition on its entire history, providing long-term memory that facilitates long-range predictions. While the local Onsager--Machlup theory in our companion paper assumes a temporally local predictive action, we investigate the conditions under which this locality holds. We derive the predictive path measure for latent dynamics that become non-Markovian due to attention-induced memory, demonstrating that this measure is the projection of a hidden linear Markov augmentation. Eliminating the auxiliary field results in a nonlocal Onsager--Machlup action, where memory manifests as a nonlocal quadratic form rather than a force. These kernels are completely monotone and exactly match a hidden Markov embedding with a finite relaxation spectrum; otherwise, the dynamics remain fundamentally nonlocal. By expanding the action in terms of the scale-separation parameter $ε=τ{\text{mem}}/τ{\text{dyn}}$, we show that the leading order recovers the local action of the companion paper, establishing locality as the short-memory limit of a nonlocal theory. We verify the reversible sector of this expansion term by term against an exactly solvable vector linear model.

Authors (1)

Summary

  • The paper establishes that attention mechanisms induce nonlocal memory effects, reformulating the predictive path measure as a nonlocal Onsager–Machlup action.
  • It employs a hidden linear Markovian augmentation to derive completely monotone memory kernels, bridging short-memory and long-memory regimes.
  • A controlled expansion using a scale-separation parameter quantifies deviations from locality while preserving entropy-producing dynamics.

Path-Measure Dynamics in Attention-Driven World Models: A Nonlocal Onsager–Machlup Framework

Overview

This paper constructs an analytic framework connecting attention-based world models to the formalism of nonlocal Onsager–Machlup actions. It establishes that attention mechanisms, by introducing history-dependent dynamics, fundamentally render the predictive path measure nonlocal in time. These nonlocal dynamics are characterized as the projection of an underlying linear Markovian augmentation, with the memory effects entering as completely monotone kernels. The conventional local Onsager–Machlup theory emerges only in the regime of fast memory decay (short-memory or Markovian limit), revealing that temporal locality is a derived property rather than a foundational assumption.

Attention-Induced Nonlocality in World Models

The core premise is that attention mechanisms, such as those used in Transformers and state-of-the-art world models, inject effective memory into the latent dynamics. The resulting transition probabilities p(zt+1zt)p(z_{t+1} | z_{\leq t}) depend on the entire past trajectory rather than solely the current state, invalidating the Markov property. The paper demonstrates that under plausible conditions, such attention-induced memory admits a formal representation as a memory kernel acting on past latent states.

This insight recasts the predictive object: rather than a sequence of one-step conditionals, the world model defines a measure over entire future paths P[]eA[]P[\cdot] \propto e^{-A[\cdot]}, with A[]A[\cdot] (the action) now generically nonlocal in time. The result is a path-integral formalism for the latent trajectories, where prediction, planning, and uncertainty quantification are unified.

Markov Augmentation and Memory Kernels

To formalize the connection between the observed non-Markovian process and an underlying Markov evolution, the paper introduces a hidden (auxiliary) linear dynamics, in which the observed process zz is coupled to an unobserved fast variable ww. This augmented system is Markovian, and integrating out the auxiliary field produces an effective description for zz only, with the crucial result that the resulting memory kernel K(u)K(u) is not assumed but mathematically generated by this projection.

The authors prove that the class of kernels K(u)K(u) arising from such linear auxiliary augmentations corresponds precisely to the set of completely monotone functions. In operational terms, any kernel of this form can be decomposed as an exponential mixture (Bernstein–Widder representation), and any such kernel can be realized by appropriate Markovian embedding. This establishes a sharp characterization of the history-dependence induced by attention mechanisms within this analytic class.

Nonlocal Onsager–Machlup Functional and Its Expansion

Eliminating the auxiliary variables, the path measure for the observed zz is a nonlocal Onsager–Machlup action,

A[z]=14D0Tz˙(t)0K(u)f[z(tu)]du2dt+A[z] = \frac{1}{4D} \int_0^T \left| \dot{z}(t) - \int_0^\infty K(u) f[z(t-u)] du \right|^2 dt + \cdots

where P[]eA[]P[\cdot] \propto e^{-A[\cdot]}0 is the drift induced by the learned dynamics. The action includes memory effects as nonlocal quadratic forms rather than as additional forces.

The main technical advance is a controlled expansion of this action in a scale-separation parameter P[]eA[]P[\cdot] \propto e^{-A[\cdot]}1, corresponding to the ratio of the memory timescale to the intrinsic dynamical timescale. At leading order in P[]eA[]P[\cdot] \propto e^{-A[\cdot]}2, the local Onsager–Machlup functional is recovered. Subsequent corrections, captured systematically, quantify the deviation from locality and have clear operational interpretations:

  • The P[]eA[]P[\cdot] \propto e^{-A[\cdot]}3 correction renormalizes the kinetic metric via the symmetric part of the drift Jacobian, without introducing new dynamical forces.
  • The P[]eA[]P[\cdot] \propto e^{-A[\cdot]}4 term introduces explicit history-dependent (inertial) effects, formally as higher-derivative corrections.

The regime of validity of the local theory is thus precisely and operationally defined via measurable parameters of the latent dynamics.

Preservation of Irreversibility and Sector Decomposition

A key structural claim is that the sector of the action responsible for time-irreversible (entropy-producing) behavior, identified as the antisymmetric part of the drift, passes through the Markov projection and nonlocal embedding unaltered. The memory kernel only modifies the reversible metric sector, not the irreversible channel. This ensures that signatures of non-equilibrium behavior present in the local formulation remain diagnostic after including nonlocality.

This claim is proven using the spatial scalar structure of the memory kernel. If generalizations to matrix-valued kernels are considered, the preservation result can fail, as the sector decomposition may no longer commute with the convolution structure of the kernel.

Implications, Diagnostic Methods, and Research Directions

The theoretic framework constructed here reframes attention-based world models as stochastic processes over path measures with structurally derived memory. This has several notable implications:

  • Model Comparison: Since the scale-separation parameter and the kernel class are dynamical, not architectural, properties, they provide a principled basis for comparing memory effects across a range of architectures and implementations, independent of superficial design choices.
  • Empirical Verification: The completely monotone nature of the kernel in the linear passive case yields concrete, falsifiable diagnostics. Kernel estimation from observed latent trajectories allows practitioners to test whether a trained model operates within or beyond the analytic class considered.
  • Irreversibility Quantification: The analytic separation of reversible and irreversible contributions, preserved by the nonlocal transition, allows for targeted measurement of entropy production and thermodynamic bounds in learned models.
  • Measurement and Generalization: The path-measure point of view suggests a shift in model evaluation from predicting next-step statistics to measuring full-path distributions, which can capture richer aspects of long-range memory and uncertainty.

Several concrete research directions are identified:

  1. Experimental deconvolution to extract memory kernels from trained models and assess kernel-class membership.
  2. Direct measurement of the action parameters and irreversibility diagnostics using trajectories.
  3. Investigation of models with kernel spectra outside the completely monotone class, i.e., active or oscillatory memory, which would fall outside the scope of the current analysis.
  4. Extending the field-theoretic and renormalization-group treatment to explore universal properties of nonlocal world models.

A caveat is acknowledged: residual, normalization, or feedforward components could also contribute to the effective nonlocal memory, beyond the attention mechanism per se. Disentangling these experimentally remains an open challenge.

Conclusion

This paper rigorously formulates the predictive path measure in attention-driven world models as a nonlocal Onsager–Machlup action, arising from the projection of a hidden linear Markov augmentation. The memory kernel, fully characterized by complete monotonicity, encodes the memory effects imparted by attention. The local Markovian form used in earlier studies is shown to be the leading-order limit in a controlled expansion, making locality a derived limit rather than an axiom. Empirical kernel estimation will determine the scope of applicability of this analytic framework, and the separation of reversible and irreversible sectors sharpens future investigations into memory, irreversibility, and non-equilibrium learning in AI.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.