Papers
Topics
Authors
Recent
Search
2000 character limit reached

When do trajectories matter? Identifiability analysis for stochastic transport phenomena

Published 17 Apr 2026 in nlin.CG, q-bio.QM, and stat.AP | (2604.15598v1)

Abstract: Stochastic models of diffusion are routinely used to study dispersal of populations, including populations of animals, plants, seeds and cells. Advances in imaging and field measurement technologies mean that data are often collected across a range of scales, including count data collected across a series of fixed sampling regions to characterize population-level dispersal, as well as individual trajectory data to examine at the motion of individuals within a diffusive population. In this work we consider a lattice-based random walk model and examine the extent to which model parameters can be determined by collecting count data and/or trajectory data. Our analysis combines agent-based stochastic simulations, mean-field partial differential equation approximations, likelihood-based estimation, identifiability analysis, and model-based prediction. These combined tools reveal that working with count data alone can sometimes lead to challenges involving structural non-identifiability that can be alleviated by collecting trajectory data. Furthermore, these tools allow us to explore how different experimental designs impact inferential precision by comparing how different trajectory data collection protocols affects practical identifiability. Open source implementations of all algorithms used in this work are available on GitHub.

Summary

  • The paper demonstrates that trajectory data resolves structural non-identifiability in key parameters like carrying capacity compared to count data.
  • It employs stochastic simulations and mean-field PDE approximations to accurately estimate motility, drift, and exclusion effects in random walk models.
  • The results highlight that optimal experimental design, such as leading edge tagging, enhances inference precision and narrows prediction intervals.

Identifiability Analysis of Stochastic Transport via Trajectories

Introduction

The paper "When do trajectories matter? Identifiability analysis for stochastic transport phenomena" (2604.15598) systematically analyzes the extent to which the parameters of random walk models describing diffusive population dispersal can be determined from count data and individual trajectory data. The authors employ a lattice-based random walk model with carrying capacity, bridging exclusion, partially occupied, and fully occupied lattice limits. Through combined stochastic simulations and mean-field PDE approximations, likelihood-based estimation, and identifiability analysis, the study quantifies how data type and experimental design impact parameter identifiability, particularly the structural and practical identifiability of carrying capacity, motility, and drift parameters.

Stochastic Transport Model Framework

A discrete-time random walk model is defined on a 2D lattice with per-site carrying capacity κ\kappa, motility probability MM, and bias parameters (ρx,ρy)(\rho_x, \rho_y). The model generalizes from the exclusion process (κ=1\kappa=1) to unbiased and biased Brownian motion (κ\kappa \to \infty). Agents move probabilistically, subject to site occupancy constraints and potential motion bias.

Mean-field PDE surrogates are derived for both the evolution of mean site occupancy N(x,t)N(x,t) (count data) and the probability density of tagged agent positions P(x,t)P(x,t) (trajectory data). For unbiased motility, the PDE for N(x,t)N(x,t) is independent of KK (continuum carrying capacity), introducing structural non-identifiability of KK. In contrast, MM0 retains dependence on MM1 in all cases, making trajectory data informative about carrying capacity even when count data is not. Figure 1

Figure 1: Schematic snapshot of a random walk population; count data is derived from region-wise tallies, while trajectory data arises from tracking tagged individuals.

Likelihood-Based Parameter Estimation and Identifiability

Loglikelihood functions are formulated for both data types: a binomial likelihood for count data, explicitly incorporating carrying capacity, and a likelihood from the PDE-driven PDF for trajectory data. Structural identifiability is analytically characterized, distinguishing settings where unique parameter recovery is impossible even with ideal data. Practical identifiability is assessed numerically via loglikelihood profiles and Wilks' confidence threshold, reflecting uncertainty from finite and noisy observations.

The approach enables efficient parameter searches in low-dimensional settings and constructs profile likelihoods to visualize pairwise parameter correlations and confidence sets:

  • For unbiased random walks, count data yields broad confidence regions for MM2 — MM3 is poorly-identified (structural non-identifiability). Trajectory data sharply restricts MM4 and MM5, confirming the necessity of trajectory measurements in such cases.
  • For biased random walks, both count and trajectory data are informative. Negative correlation between drift and carrying capacity is observed for count-derived estimates. Figure 2

    Figure 2: Visualization of the random walk process, illustrating unbiased and biased motility; tagged agents are visibly dispersed, facilitating trajectory-based inference.

    Figure 3

    Figure 3: Comparison of stochastic simulation outcomes with mean-field PDE predictions for count profiles and trajectory histograms.

Experimental Design and Combined Data Analysis

The practical utility of trajectory data is evaluated with respect to tagging protocols and spatial arrangement. Tagging at the leading edge produces more informative trajectory data for parameter recovery than tagging within the bulk population. Combining count and trajectory data further constrains parameter confidence sets and enhances inference precision. Figure 4

Figure 4: Heatmaps of loglikelihoods for MM6 and MM7 from count, trajectory, and combined data—trajectory data resolves parameter uncertainty left by count data.

Figure 5

Figure 5: Bivariate profile loglikelihoods for MM8, MM9, and (ρx,ρy)(\rho_x, \rho_y)0: count and combined data yield tightly constrained confidence regions, trajectory-only regions are broader when sampling is suboptimal.

Figure 6

Figure 6: Impact of tagging protocols on identifiability: leading edge tagging produces tightly peaked loglikelihoods, while tagging in central population columns yields broader confidence intervals.

Likelihood-Based Prediction Intervals

Model-based prediction intervals are constructed from parameter confidence sets, propagating inferred uncertainty into predicted observables. For count-derived prediction, parameter uncertainty is dominated by the poorly identified (ρx,ρy)(\rho_x, \rho_y)1; for trajectory and combined data, prediction intervals are much narrower and match the data with high empirical coverage. Figure 7

Figure 7: Visualization of prediction intervals for count and trajectory observables, demonstrating improved predictive accuracy with trajectory and combined data.

Implications and Future Directions

The results demonstrate that count data alone may be insufficient for reliable inference in common stochastic transport scenarios, particularly when crowding and exclusion effects are modeled. Trajectory data significantly enhances structural and practical identifiability, especially for parameters such as carrying capacity that are invisible to traditional count-based inference. This underscores the importance of quantitative modeling in guiding experimental design — the location and quantity of tagged individuals directly influence inferential precision.

Practically, these findings are relevant in ecology and cell biology, where the investment in trajectory-tracking technology and protocol must be justified by improved inference and predictive accuracy. The computational approach (efficient mean-field surrogates) allows rapid analysis across settings, and the modular likelihood framework is extendable to higher-dimensional models, alternative stochastic processes, and more complex experimental arrangements.

Theoretically, this work sets a foundation for further studies on optimal data collection design, extension beyond mean-field models via pair or higher-order approximations, and rigorous treatment of measurement error and observation noise. Applications in the broader AI context include the statistical design of agent-based learning protocols and accurate inference in systems with hidden or latent space structure.

Conclusion

This paper offers a rigorous identifiability analysis for stochastic transport models, demonstrating that collecting and analyzing individual trajectory data is essential in resolving structural and practical non-identifiability inherent in count-only datasets. The computational framework is efficient and widely applicable, and the results have direct implications for experimental design in fields utilizing stochastic models of population dispersal, such as cell migration and animal movement ecology. Future work includes expansion to complex system geometries, refinement of trajectory data processing, richer noise and measurement models, and application of these concepts to additional domains in statistical inference and agent-based AI.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.