Papers
Topics
Authors
Recent
Search
2000 character limit reached

Simultaneous Inference for Partially Observed Functional Time Series

Published 30 Jun 2026 in stat.ME and math.ST | (2606.31269v1)

Abstract: Functional data analysis (FDA) provides statistical methods for analyzing samples of time-continuous stochastic processes. Measurements often arise in the form of sensor data for a key scientific variable. The practical problem of irregular sensor disruptions has fostered interest in analyzing partially observed random functions. Specifically, this paper is motivated by a time series of intermittently missing pollution data with dependence along pollution paths and missingness patterns. To allow statistical analysis, we develop the first inference methods for dependent, partially observed functional time series. Existing methods were not appropriate for this task, because they heavily rely on the independence of the data functions. Mathematically, we model data on the space of bounded functions equipped with the supremum norm. This allows simultaneous inference across the entire functional domain, including simultaneous confidence bands -- something existing Hilbert-space-based methods cannot provide. To study non-stationary trends along the time series, we extend state-of-the-art multiscale inference methods (originally developed for scalar data) to partially observed functions. The key application of the latter methods is testing for excessive pollution levels in inner cities. Our approach combines state-of-the-art Gaussian approximations with stochastic process theory. Interestingly, it also improves existing results for fully observed functional time series by avoiding a functional CLT.

Authors (2)

Summary

  • The paper introduces a novel framework for constructing simultaneous confidence bands under partial observation and serial dependence using supremum-norm geometry.
  • It extends multiscale statistics to test stationarity, change points, and threshold exceedances, achieving optimal convergence rates.
  • Empirical results on air pollution data validate the method's ability to handle real-world sensor dropout and complex temporal trends.

Simultaneous Inference for Partially Observed Functional Time Series

Introduction and Context

This work addresses the challenge of conducting simultaneous statistical inference for functional time series where observations are partially missing. Motivated by practical settings such as air pollution monitoring, the authors focus on time-ordered collections of real-valued functions Xi(â‹…)X_i(\cdot) observed only on random subdomains, with both functions and missingness exhibiting serial dependence. The classical functional data analysis (FDA) framework primarily assumes independence and full observation; however, modern sensor data exhibits none of these properties, necessitating new methodologies.

The methodological contributions are threefold:

  • Simultaneous inference for dependent, partially observed functional time series under general serial dependence, including in the missingness patterns.
  • A supremum-norm framework using the Banach space ℓ∞[0,1]\ell^\infty[0,1], enabling construction of simultaneous confidence bands (SCBs) for the mean function, even in the presence of function- and missingness-driven dependence.
  • Extension of multiscale statistics—originally scalar data tools—to local inference tasks such as stationarity testing, change point analysis, and threshold exceedance in non-stationary, partially observed functional time series.

A running application is the detection and quantification of air pollution exceedances, where both the pollutant profiles and the missing data mechanisms are strongly temporally dependent.

Mathematical Framework and Assumptions

The data are modeled as bounded functions in ℓ∞[0,1]\ell^\infty[0,1], with partial observation formalized via indicator functions Oi(⋅)O_i(\cdot). Serial dependence is incorporated by a physical dependence framework, generalizing classical mixing conditions and enabling tractable high-dimensional Gaussian approximations. The mean function μ(⋅)\mu(\cdot) and the probability π(⋅)\pi(\cdot) of observation are assumed to be Hölder-continuous, with long-run variance conditions ensuring statistical identifiability.

A novel aspect is the modeling of missingness not as an L2L^2 process, but as a pointwise (indicator) process, capturing realistic sensor outage scenarios that create complex dependence both within and across functions. The authors construct stylized models (e.g., double waiting-time models with exponential failure/repair times) to justify the wide applicability of their structural and regularity assumptions. Figure 1

Figure 1: Example of dependent missingness in sensor data, with functional observations exhibiting overlapping periods of dropout due to serially dependent outages.

Simultaneous Confidence Bands: Theory and Construction

The first main result is the construction of SCBs for the mean function μ(⋅)\mu(\cdot) in the presence of serial dependence and partial observation. The methodology consists of:

  1. Grid-based Simultaneous Inference: Estimating μ(t)\mu(t) on a polynomially dense grid {tj}\{t_j\} by pointwise empirical means, normalizing by the effective sample size at each ℓ∞[0,1]\ell^\infty[0,1]0.
  2. High-dimensional Gaussian Approximations: Using recent results for weakly dependent processes, the distribution of the supremum of appropriately normalized errors across the grid is approximated by that of a tight Gaussian process ℓ∞[0,1]\ell^\infty[0,1]1. This bypasses the need for a functional central limit theorem (CLT), which is often intractable in the presence of partial observation and dependence.
  3. Confidence Band Construction: Individual pointwise confidence intervals are interpolated to yield a confidence corridor, with anti-concentration inequalities ensuring that the resulting SCB covers ℓ∞[0,1]\ell^\infty[0,1]2 at the desired nominal level, up to negligible error rates.

Notably, the SCB width achieves the optimal rate ℓ∞[0,1]\ell^\infty[0,1]3, avoiding logarithmic inflation common in multiple testing. Figure 2

Figure 2

Figure 2: Illustration of pointwise confidence intervals at grid points and their interpolation to yield a confidence corridor; coverage is illustrated both in cases of success and rare "escape" failure.

A crucial finding is the necessity of grid-based discretization: for dependent, partially observed data, uniform-in-ℓ∞[0,1]\ell^\infty[0,1]4 results are provably impossible for generic processes, as explicit counterexamples demonstrate divergence of ℓ∞[0,1]\ell^\infty[0,1]5 even when the grid-based maximum converges. This stands in contrast to fully observed, independent FDA, where process-level CLTs are more attainable.

Multiscale Local Inference: Nonstationarity, Change Points, and Threshold Exceedance

The methodology is extended to non-stationary regimes, where the mean ℓ∞[0,1]\ell^\infty[0,1]6 may vary across time. The multiscale approach considers statistics over all intervals (in ℓ∞[0,1]\ell^\infty[0,1]7) and at all points (ℓ∞[0,1]\ell^\infty[0,1]8), normalized with log-factors to balance sensitivity between local and global alternatives.

Key statistics include:

  • ℓ∞[0,1]\ell^\infty[0,1]9: For stationarity and change point detection, combines local differences of means across intervals at multiple scales.
  • ℓ∞[0,1]\ell^\infty[0,1]0: For threshold exceedance, searches for intervals and locations where the mean significantly exceeds a prespecified level.

Limiting distributions are characterized using the modulus of continuity of a Gaussian process-valued Brownian motion ℓ∞[0,1]\ell^\infty[0,1]1, enabling calibration of quantile thresholds and test levels even under complex dependence.

Power analysis under local alternatives establishes sharp detection boundaries for both "spiky" and "smooth" deviations, with almost parametric rates for persistent or intense changes and near-optimal adaptation for short, localized signals. Figure 3

Figure 3

Figure 3: Synthetic data with a short, localized change in mean, and multiscale test statistics at various scales, demonstrating scale adaptivity and localization.

Empirical Results and Application

Simulations confirm nominal coverage and error control of the SCBs, matching the performance of methods designed for the much more restrictive setting of independent, fully observed data. Empirical power is demonstrated to be high against both uniform and highly localized alternatives, with the multiscale method comparing favorably to and in many settings outperforming contemporaries in the literature, particularly for the detection of short-lived, spatially/functally concentrated changes.

A real-world application to air pollution data demonstrates the practical utility of the methods. The threshold exceedance statistic is applied to six months of Mumbai ℓ∞[0,1]\ell^\infty[0,1]2 data, featuring significant serially dependent missingness and temporal trends. Testing for exceedances of multiple thresholds, the method detects substantial periods of hazardous pollution, localized both in time of day and season. Figure 4

Figure 4: Proportion of observed samples at each time under two different missingness schemes (interval-based and serially dependent exponential outages), showing realistic dropout patterns analogous to those observed in air quality sensor deployments.

Figure 5

Figure 5

Figure 5

Figure 5

Figure 5: Heatmaps of threshold exceedances over time and within the day, with color indicating the smallest scale at which significant exceedance is detected. Peaks are seen overnight and in early morning, with periods of hazard localized by season.

Implications and Future Directions

The theoretical advances in this paper are significant for both practical and theoretical reasons:

  • Practical Implications: The construction of valid, simultaneous inference tools for partially observed, dependent functional data enables robust environmental, biomedical, and engineering analyses under highly realistic data incompleteness typical for large-scale sensor deployments.
  • Theoretical Implications: The methodology breaks from Hilbert-space-based FDA and overcomes longstanding obstacles in constructing nonparametric SCBs under dependence and missingness, introducing grid-based, supremum-norm approaches as an alternative foundation.
  • Improved Generality: All results for partial observation also improve the state-of-the-art for fully observed dependent functional time series, as function-valued CLTs and tightness conditions can be replaced with grid-based Gaussian approximations with minimal smoothness.

Future directions may include the extension of these frameworks to:

  • High-dimensional (vector- or tensor-valued) functional observations with structured missingness,
  • Adaptive grid selection schemes, balancing computational and statistical efficiency for very large ℓ∞[0,1]\ell^\infty[0,1]3,
  • Nonparametric modeling of the missingness process—relaxing independence or Markovian assumptions—and robust, model-free inference procedures,
  • Integration with downstream learning pipelines (e.g., prediction or classification with incomplete functional covariates), leveraging the developed uncertainty quantification tools.

Conclusion

This paper presents a comprehensive framework for simultaneous inference on partially observed, serially dependent functional time series. By exploiting supremum-norm geometry and multiscale statistics, it achieves strong inferential guarantees for both global (mean estimation, SCBs) and local (change point, threshold exceedance) tasks, even under conditions precluding classical functional CLTs. The techniques are validated both theoretically and empirically, with broad applicability to complex, incomplete functional datasets.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.