Papers
Topics
Authors
Recent
Search
2000 character limit reached

Markov Information Processes

Published 5 Jul 2026 in math.OC and econ.TH | (2607.04308v1)

Abstract: We study information design when a designer with commitment shapes the information of strategically interacting, far-sighted agents whose actions drive a persistent, controlled Markov state. We introduce the Markov Bayes correlated equilibrium (Markov BCE), the controlled-Markov generalisation of the BCE of Bergemann and Morris (2016), characterised by a dynamic obedience condition that adds a continuation-value term to the static one and reduces to it when actions cannot move the state. Recommending actions is without loss; the designer's problem is recursive in the agents' promised continuation utilities and is solved by a set-valued backward-induction algorithm whose optimum exists and lies between the no-disclosure and first-best values. For linear-quadratic-Gaussian payoffs the obedience condition becomes a covariance condition with a modified interaction matrix, and the stationary case reduces to an algebraic Riccati equation. When agents instead learn the transition, we identify the rent an agent earns from a model of the dynamics sharper than the designer anticipates: it is non-negative, zero at the known-dynamics benchmark, and deterred only by building slack into obedience. Under persistent excitation the cumulative rent grows logarithmically as heterogeneous agents' estimates converge. Two worked examples, in congestion and resource coordination, together with a numerical study illustrate the theory.

Authors (1)

Summary

  • The paper introduces a rigorous framework for dynamic obedience in controlled Markov games by extending static BCE to a Markov setting.
  • It formulates the designer's problem as a recursive program with explicit algorithms, especially in structured LQG settings using Riccati recursion.
  • It quantifies strategic rents from agent learning and model uncertainty, demonstrating welfare gains over benchmarks with limited cost.

Markov Information Processes: Dynamic Obedience in Controlled Markov Games

Overview

This paper introduces a rigorous framework for information design in dynamic games where a designer shapes the information available to strategically interacting, far-sighted agents whose actions evolve a persistent, controlled Markovian state. The key innovation is the notion of Markov Bayes Correlated Equilibrium (Markov BCE), which extends static Bayes Correlated Equilibrium (BCE) by incorporating the effect of continuation utilities. The framework addresses both the case of known and unknown transition dynamics, analyzing incentive compatibility, recursion, and learning-driven rents in linear-quadratic-Gaussian (LQG) settings. The analysis is supported by explicit recursive algorithms and illustrative applications in congestion and resource coordination.

Markov Bayes Correlated Equilibrium: From Static to Dynamic Settings

The BCE concept [bergemann2016bayes] is foundational in static information design, specifying obedience constraints that require each agent to prefer its recommended action, considering its posterior belief induced by the designer. The present work generalizes this to controlled Markov environments where the state evolves dynamically in response to agents' actions.

The central contribution is the dynamic obedience condition: agents are not only stage-utility maximizers but are also forward-looking, evaluating deviations by the entire continuation value—the expected utility from future stages under the transition kernel shaped by their deviations. The dynamic obedience constraint (Definition \texttt{def:dynobd}) captures both stage-wise incentives and the differentiated expected future utility of deviating versus obedience.

The main technical insight is that, under Markov BCE, there is a dynamic revelation principle: recommendations (as opposed to abstract signals) suffice for incentive compatibility. The designer's problem—optimizing over policy recommendations—can thus be cast as a recursive program in promised continuation utilities, tractable via backward induction.

Recursive Information Design and Algorithmic Solution

The designer's problem is formulated recursively, carrying continuation utilities as state variables [abreu1990toward, sannikov2008continuous]. At each stage, the designer selects a policy maximizing expected payoff, subject to dynamic obedience, with each agent’s continuation utility maintained via a promise-keeping constraint. This induces a Bellman-type recursion, solved via a set-valued Abreu–Pearce–Stacchetti (APS) backward induction (Algorithm 1).

The framework establishes that, for structured classes (notably LQG), the dynamic obedience constraints admit an explicit, finite-dimensional parametrization, and the set-valued recursion simplifies to a tractable Riccati recursion. Existence of an optimal Markov BCE is guaranteed under mild regularity and continuity (Propositions \texttt{prop:fixedpoint}, \texttt{prop:optexist}).

LQG Case: Covariance Characterization and Riccati Recursion

For LQG payoffs and transitions, dynamic obedience collapses to a covariance condition on second moments (Theorem \texttt{thm:lqg}), generalizing [Bergemann2013]. The effective strategic interaction matrix is renormalized to account for state-contingent payoffs. The recursive structure admits an explicit Riccati recursion on continuation value matrices (Lemma \texttt{lem:quadratic}), with the stationary case solved by a fixed point in the algebraic Riccati equation (Proposition \texttt{prop:stationary}).

Numerical results demonstrate that, in practical settings, the value of sophisticated information design can be considerable compared to no-disclosure benchmarks, while the cost of enforcing obedience (the price of incentives) can be relatively modest. For example, in a congestion instance, utilitarian optimal design attains a value of $3.43$ compared to a first-best of $3.58$ and no-disclosure at zero, with the maximin design increasing welfare for the weaker agent at a modest total cost.

Learning, Model Uncertainty, and Strategic Rents

The second part of the paper investigates strategic learning: agents observe the persistent state and learn the Markov transition kernel. When agents acquire superior models, they can potentially steer the state distribution to increase future payoffs, circumventing the designer's intention. The rent extracted by an agent from model superiority is explicitly quantified; it is non-negative, vanishes when the designer's model is accurate, and is fundamentally quadratic in estimation error in the LQG case (Proposition \texttt{prop:rent}).

Logarithmic rent bounds are established under persistent excitation and regularized on-policy learning: the cumulative rent grows as O~(logT)\tilde{O}(\log T) as parameter estimates concentrate (Theorem \texttt{thm:regret}). This is supported by self-normalized least-squares estimation theory [abbasi2011improved, dean2020sample]. The designer can mitigate rent extraction only by over-satisfying obedience, building in slack proportional to the maximal anticipated model misspecification (Proposition \texttt{prop:robust}).

Illustrative Examples: Congestion and Power Coordination

Two explicit examples demonstrate the framework:

  • Evacuation/Congestion: Actions are complements; dynamic obedience penalizes synchronization and induces staggering.
  • Power Coordination: Actions are substitutes; strategic withholding emerges as the dominant deviation when agents possess superior system knowledge.

In both cases, the Markov BCE and the associated rent formula reduce to analytically tractable forms, and the sign of the cross-coupling parameter in payoffs inverts the interpretation.

Implications and Open Problems

The paper clarifies that information design in dynamic games fundamentally intertwines with incentive design, recursive utility promises, and learning. The Markov BCE formulation systematically extends BCE to dynamic, controlled-state environments, allowing tractable analysis in both general and LQG settings.

The findings suggest several implications:

  • Robust information design must account for agents' learning and model uncertainty; static benchmarks are fragile to dynamic mis-specification.
  • Algorithmic tractability is attainable in structured classes, enabling applications in complex environments (e.g., power systems, dynamic resource allocation).
  • On a theoretical level, incentive-compatible information design intertwines with recursive contracts, dynamic mechanism design, and identification of strategic system parameters.

A central open problem is the unconditional behavior of agent rents under fully endogenous, mutually entangled learning, wherein designer and agents co-evolve their models and obedience constraints recursively—a genuinely dynamic information arms race.

Conclusion

Markov information processes unify dynamic information design, recursive incentives, and learning in Markovian dynamic games. This framework extends static Bayesian persuasion and correlated equilibrium to settings of persistent, controlled state evolution and forward-looking agents, enabling both rigorous analysis and computationally viable algorithms in domains where information, learning, and incentives are deeply intertwined.

References:

  • [bergemann2016bayes]
  • [Bergemann2013]
  • [abreu1990toward]
  • [sannikov2008continuous]
  • [abbasi2011improved]
  • [dean2020sample]
  • [makris2023information]

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.