Hierarchical Affine Decoupling and an Explicit Decentralized Nash Equilibrium via Coupled Riccati Systems
NPMAI ECOSYSTEM is a multi-domain research and engineering organisation based in Kota, Rajasthan. It runs production AI tooling used by hundreds of thousands of developers, alongside an original research programme spanning computer science, social science, chemistry, and mathematics — the department responsible for the paper that follows.
Agentic systems, RAG architectures, load balancing, and applied ML infrastructure.
Political science frameworks and constitutional law proposals, grounded in real governance data.
Reactor design, catalysis, and synthetic fuel pathways — including hydrogen and materials work.
Stochastic control, game theory, and dynamical systems. Home department of this paper.
Sonu Kumar founded NPMAI ECOSYSTEM and built its foundational infrastructure — the npmai package, the dual load balancer, and the LARA retrieval architecture — entirely on free-tier cloud infrastructure. He leads the organisation's Research Tech and Social Science departments and set up the multi-department research programme under which this Mathematics Department paper is published. He is not an author of the present paper.
Classical continuous-time Linear-Quadratic Mean Field Games (LQ-MFGs) typically restrict attention to simple linear drifts, quadratic control costs, and interactions that depend only on the population average. This paper develops a generalized LQ-MFG that simultaneously incorporates four mechanisms absent from the classical formulation: asymmetric Ornstein–Uhlenbeck mean-reverting dynamics, a stochastic common noise shared by the entire population, a deterministic exogenous tracking trajectory M(t), and a control–state interaction penalty coupling individual control effort to the aggregate state.
Using the Stochastic Maximum Principle, we derive the coupled Forward-Backward Stochastic Differential Equation (FBSDE) governing the representative agent, and — through a hierarchical affine decoupling ansatz applied first at the macroscopic (population) level and then at the microscopic (individual) level — reduce the coupled stochastic system to two nonlinear Riccati differential equations and two linear ordinary differential equations. We solve these explicitly in closed form and, invoking a Mean Field Consistency condition, obtain a fully explicit decentralized Nash equilibrium control law. The equilibrium decomposes naturally into a local consensus-correction term, a system-wide stabilization term generated by the common noise and the control–state coupling, and an anticipatory tracking term generated by the exogenous trajectory. Several existing LQ-MFG models are recovered as special cases.
Mean Field Games (MFGs), introduced independently by Jean-Michel Lasry and Pierre-Louis Lions, provide a mathematical framework for analyzing strategic interactions among an extremely large population of rational agents. Rather than solving an intractable finite-player dynamic game, the influence of each individual agent is assumed to vanish as the population size approaches infinity, allowing every agent to optimize against an aggregate statistical description of the population known as the mean field. This limiting formulation has found applications in economics, finance, engineering, networked control systems, opinion dynamics, smart grids, and large-scale autonomous systems.
Among the numerous classes of Mean Field Games, Linear-Quadratic Mean Field Games (LQ-MFGs) occupy a particularly important position because they combine analytical tractability with the ability to model complex interactions through linear dynamics and quadratic objective functionals. Their mathematical structure frequently permits explicit equilibrium characterizations through Forward-Backward Stochastic Differential Equations (FBSDEs) and Riccati differential equations, making them valuable both theoretically and computationally.
Classical LQ-MFG formulations generally assume relatively simple state dynamics in which each agent evolves under a linear drift, quadratic control costs, and interactions that depend solely upon the population average. Although these assumptions simplify the equilibrium analysis, they neglect several important phenomena frequently encountered in practical systems.
First, many physical, economic, and biological systems exhibit mean-reverting behavior, whereby an individual's state naturally tends toward an equilibrium level unless continuously acted upon. Such dynamics are well represented by Ornstein–Uhlenbeck processes and arise naturally in financial interest-rate models, inventory regulation, energy markets, and population dynamics.
Second, agents often attempt to track not only the collective population behavior but also an externally imposed reference trajectory. Such trajectories may represent macroeconomic indicators, regulatory benchmarks, technological targets, environmental indices, or desired production schedules. Incorporating these time-varying reference signals significantly enriches the strategic behavior of the game while preserving its linear-quadratic structure.
Third, real-world systems are frequently exposed to common sources of uncertainty. Unlike independent idiosyncratic disturbances, common noise simultaneously affects every participant and therefore remains present even in the infinite-population limit. Consequently, the mean field itself becomes stochastic rather than deterministic, requiring the equilibrium analysis to be performed conditionally with respect to the filtration generated by the common noise.
Motivated by these considerations, this paper develops a generalized continuous-time Linear-Quadratic Mean Field Game that simultaneously incorporates asymmetric mean-reverting dynamics, stochastic common noise, an exogenous tracking trajectory, and an additional interaction penalty coupling the individual control effort with the aggregate state. The resulting optimization problem remains analytically tractable while substantially extending the class of models that admit explicit closed-form equilibria.
The principal contributions of this work may be summarized as follows:
The remainder of the paper is organized as follows. Section 1 establishes the mathematical model and the stochastic dynamics governing the population. Section 2 derives the representative agent's optimization problem and the associated FBSDE. Section 3 develops the analytical solution through a hierarchical decoupling procedure and states the explicit decentralized Nash equilibrium. Section 4 records structural corollaries and special cases. Section 5 concludes.
Throughout this paper, we consider a finite time horizon $T>0$ defined on a complete filtered probability space
satisfying the usual completeness and right-continuity assumptions. Unless otherwise stated, all stochastic processes are assumed adapted to $\{\mathcal{F}_t\}$. The game consists of an infinite continuum of strategically interacting agents indexed by $i\in[0,1]$. Each agent controls a one-dimensional stochastic state process $X_i(t)\in\mathbb{R}$, whose evolution depends upon its own control action, the aggregate population behavior, and both idiosyncratic and common sources of uncertainty.
| Symbol | Description |
|---|---|
| $X_i(t)$ | State of agent $i$. |
| $\bar X(t)$ | Conditional population mean field. |
| $u_i(t)$ | Admissible control process of agent $i$. |
| $W_i(t)$ | Idiosyncratic Brownian motion. |
| $W_0(t)$ | Common Brownian motion. |
| $a$ | Mean-reversion coefficient. |
| $b$ | Mean-field interaction coefficient. |
| $\sigma$ | Idiosyncratic volatility intensity. |
| $\sigma_0$ | Common volatility intensity. |
| $M(t)$ | Deterministic tracking trajectory. |
| $X_D$ | Terminal target state. |
| $q,\eta,g,\gamma$ | Cost-function parameters. |
The admissible control space is defined by
ensuring that every admissible strategy possesses finite expected energy and yields a well-defined stochastic state trajectory.
We consider an infinite population of atomless strategic agents whose states evolve over the finite planning horizon $t\in[0,T]$. Each agent seeks to regulate its own state while simultaneously responding to the aggregate population behavior and an externally prescribed reference trajectory. Unlike classical LQ-MFGs, the present formulation incorporates four structural mechanisms simultaneously: intrinsic mean-reverting dynamics, asymmetric interaction with the population mean, stochastic common environmental shocks, and exogenous trajectory tracking.
The individual state dynamics are governed by the stochastic differential equation
where $a>0$ denotes the intrinsic rate of mean reversion and $b\in\mathbb R$ measures the influence of the aggregate population on the individual dynamics. Positive values of $b$ encourage conformity with the population average, whereas negative values correspond to repulsive interactions. The process $W_i(t)$ represents idiosyncratic uncertainty unique to each agent, while $W_0(t)$ models common environmental uncertainty simultaneously experienced by the entire population, with diffusion intensities $\sigma$ and $\sigma_0$ respectively.
The aggregate population state is defined by the conditional expectation
the filtration generated exclusively by the common Brownian motion. This conditional formulation is fundamental to Mean Field Games with common noise: since the independent idiosyncratic Brownian motions satisfy an exact law of large numbers over the continuum of agents, their aggregate contribution vanishes in the infinite-population limit. Only the common noise survives, rendering the mean field itself a stochastic process adapted to the common filtration.
Each player chooses an admissible control $u_i\in\mathcal U$ so as to minimize the quadratic performance functional
The objective functional consists of four components: a quadratic control cost discouraging excessive control effort; a consensus penalty promoting alignment between an individual's state and the population average, modeling herding behavior; a mixed control–state interaction term capturing situations where the effectiveness or expense of control depends on the aggregate state; and a tracking term driving each agent toward the prescribed trajectory $M(t)$, with a terminal cost enforcing convergence toward the desired terminal target $X_D$.
The interaction of these four mechanisms produces a generalized Linear-Quadratic Mean Field Game whose equilibrium structure is characterized analytically in the following sections through the Stochastic Maximum Principle and the associated coupled Forward-Backward Stochastic Differential Equations.
Having established the stochastic dynamics and objective functional governing the representative agent, we now derive the optimal decentralized strategy. Since the game consists of an infinite continuum of strategically interacting players, the influence of any single agent on the aggregate population state becomes negligible in the mean-field limit. Consequently, each player treats the population average $\bar X(t)$ as an exogenously given stochastic process while solving an individual stochastic optimal control problem.
This principle, commonly referred to as the Nash Certainty Equivalence (NCE) Principle, forms the foundation of continuous-time Mean Field Game theory. Each agent computes an optimal response to the anticipated evolution of the mean field, and equilibrium is achieved only when the resulting population behavior reproduces the same mean field that was originally assumed. Throughout this section, we regard $\bar X(t)$ as a known $\mathcal F_t^0$-adapted stochastic process; the consistency requirement is imposed in Section 3.
Fix an arbitrary agent $i$. Its state dynamics satisfy
with $u(t)$ progressively measurable and of finite second moment. The objective is to minimize
where the running cost is
and the terminal penalty is $\Phi(X(T))=\tfrac g2(X(T)-X_D)^2$. The quadratic structure guarantees convexity of the optimization problem with respect to the control variable, ensuring uniqueness of the optimal decentralized strategy.
To characterize optimal controls, we employ the Stochastic Maximum Principle, converting the stochastic optimization problem into a coupled FBSDE. Introduce the adjoint process $p(t)$ — the marginal value of the state variable — together with the martingale correction processes $z(t)$ and $z_0(t)$, associated respectively with the idiosyncratic and common Brownian motions. The stochastic Hamiltonian is
The first term measures the sensitivity of the future value function to the current dynamics, while the remaining terms represent the instantaneous running cost.
Since $H$ is strictly convex in $u$, differentiating gives
The optimal policy separates into the classical adjoint feedback $-p(t)$ and an additional correction $-\gamma\bar X(t)$ generated by the control–state interaction penalty. When $\gamma=0$, the model reduces to the standard LQ-MFG feedback law.
The adjoint process evolves according to
Since $\partial H/\partial x = -ap+q(x-\bar X)+\eta(x-M)$, the adjoint system is
Substituting $u^*(t)=-p(t)-\gamma\bar X(t)$ into the state dynamics and collecting terms yields the forward equation $dX=\big(-aX+(b-\gamma)\bar X-p\big)dt+\sigma\,dW+\sigma_0\,dW_0$. Together with the adjoint equation, the representative agent satisfies the coupled FBSDE:
This system constitutes the mathematical core of the equilibrium problem: the forward equation governs the controlled state, while the backward equation propagates the shadow value of the state from the terminal time back to the initial time. Their mutual dependence through the control creates a fully coupled stochastic boundary-value problem.
The coupled FBSDE derived above constitutes a nonlinear stochastic boundary-value problem. Although such systems generally do not admit explicit analytical solutions, the linear-quadratic structure of the present model permits a complete reduction to deterministic differential equations through an appropriate affine decoupling procedure. Rather than solving the full coupled system directly, we proceed hierarchically: first the macro-level problem, governing the evolution of the population average, then the micro-level problem, describing the optimal response of a representative agent conditional on the macro solution.
Averaging the optimal individual dynamics eliminates the idiosyncratic Brownian motions through the exact law of large numbers, leaving only the common source of randomness. The optimal control from Section 2 is $u_i^*(t)=-p_i(t)-\gamma\bar X(t)$. Taking conditional expectations with respect to $\mathcal F_t^0$ gives the aggregate optimal control $\bar u(t)=-\bar p(t)-\gamma\bar X(t)$, where $\bar p(t)=\mathbb E[p_i(t)\mid\mathcal F_t^0]$. Substituting into the averaged state equation and defining the effective macro drift coefficient $\alpha=-a+b-\gamma$, the aggregate dynamics become
Unlike classical LQ-MFGs without common noise, the mean field is itself stochastic because the common Brownian motion survives averaging. Since $\mathbb E[X-\bar X\mid\mathcal F_t^0]=0$, the consensus penalty disappears after conditional expectation, giving the aggregate backward equation
Affine decoupling ansatz. We postulate $\bar p(t)=\phi(t)\bar X(t)+\psi(t)$, where $\phi(t)$ is a deterministic feedback gain and $\psi(t)$ a deterministic tracking bias. Applying Itô's formula and matching coefficients against the aggregate backward equation (using uniqueness of the semimartingale decomposition) yields:
The original stochastic macro problem is thus reduced to a nonlinear Riccati equation for $\phi(t)$ and a non-homogeneous linear equation for $\psi(t)$.
Because every coefficient in the individual FBSDE is linear in both $X(t)$ and $\bar X(t)$, we seek the affine ansatz $p(t)=\Phi(t)X(t)+\Theta(t)\bar X(t)+\Xi(t)$, where $\Phi(t)$ measures sensitivity to the agent's own state, $\Theta(t)$ to the aggregate state, and $\Xi(t)$ is the deterministic tracking component. Applying Itô's formula, substituting the forward equations for $X$ and $\bar X$, and matching against the backward equation term-by-term gives:
$\Phi(t)$ satisfies a Riccati equation identical in form to the classical LQR gain and is interpreted as the idiosyncratic feedback gain. $\Theta(t)$ satisfies a linear equation driven by both local and macro gains and is the population coupling coefficient. $\Xi(t)$ is the tracking bias encoding the influence of $M(t)$ and $X_D$. The vanishing terminal condition $\Theta(T)=0$ reflects that the terminal cost depends only on the individual's own state, not explicitly on the population average.
Taking the conditional expectation of $p(t)=\Phi X+\Theta\bar X+\Xi$ with respect to $\mathcal F_t^0$ and comparing with $\bar p(t)=\phi\bar X+\psi$ — since both represent the same process — the uniqueness of the affine decomposition forces the Mean Field Consistency Condition:
Explicit solution of the macro Riccati equation. Let $\mu=\tfrac{\alpha-a}2=\tfrac{-2a+b-\gamma}2$ and $\omega_2=\sqrt{\mu^2+\eta}$. The Riccati equation $\dot\phi+2\mu\phi-\phi^2+\eta=0$ has characteristic roots $\mu\pm\omega_2$, and the unique solution satisfying $\phi(T)=g$ is
Explicit solution of the macro tracking equation. By the integrating factor method,
Explicit solution of the individual Riccati equation. Let $\omega_1=\sqrt{a^2+q+\eta}$. Proceeding as before,
The consistency condition immediately reconstructs $\Theta(t)=\phi(t)-\Phi(t)$ and $\Xi(t)=\psi(t)$ — no additional differential equations need be solved.
Rewriting the equilibrium controller reveals its strategic decomposition:
The following corollaries follow directly from Theorem 3.1 and illustrate how the generalized model reduces to well-known LQ-MFGs under appropriate parameter choices, while also isolating the independent role played by each additional modeling component.
Because each component depends on a distinct deterministic coefficient, modifications to one modeling feature affect only its corresponding feedback channel without altering the remaining components — a modular structure that is one of the principal analytical advantages of the proposed framework.
| Mechanism | Governs | Vanishes when |
|---|---|---|
| Local regulation | $\Phi(t)$ — individual Riccati gain | $X_i \equiv \bar X$ |
| Population coordination | $\phi(t)+\gamma$ — macro gain + control–state coupling | $\sigma_0=0,\ \gamma=0$ (Cor. 4.1, 4.3) |
| Exogenous tracking | $\psi(t)$ — driven by $M(t)$, $X_D$ | $M\equiv X_D,\ g=0$ (Cor. 4.2) |
This paper developed a generalized continuous-time Linear-Quadratic Mean Field Game incorporating four interacting mechanisms: mean-reverting dynamics, common stochastic disturbances, exogenous trajectory tracking, and a control–state interaction penalty. Using the Stochastic Maximum Principle, the decentralized optimization problem was transformed into a coupled Forward-Backward Stochastic Differential Equation. By introducing hierarchical affine representations for both the aggregate and individual adjoint processes, the original infinite-dimensional stochastic game was reduced to two nonlinear Riccati equations together with non-homogeneous linear tracking equations, enabling the complete analytical construction of the decentralized Nash equilibrium without iterative numerical procedures.
The resulting equilibrium controller naturally decomposes into three feedback layers corresponding to local consensus regulation, systemic stabilization, and exogenous trajectory tracking. Several classical LQ-MFGs were recovered as special cases by suitable parameter choices, demonstrating that the proposed formulation provides a unified framework encompassing a broad family of existing models.
Future research may extend the present analysis to multidimensional state spaces, nonlinear drift structures, heterogeneous populations, and learning-based Mean Field Games in which the exogenous trajectory is estimated online rather than prescribed a priori.