The Crowd Optimizes Back

$$-\,\partial_t V \;=\; q\big(x,\rho_t\big) + \mathcal LV - \tfrac{1}{2\alpha}|\nabla V|^2 \qquad\quad \partial_t \rho \;=\; \nabla\!\cdot\!\Big(\rho\,\big(\nabla U + \tfrac{1}{\alpha}\nabla V\big)\Big) + T\Delta\rho$$
Marv·published 04.08.26·last edited 04.08.26

Both equations above have already appeared in this sequence, and neither is new. The right one is the Fokker–Planck equation of the first essay, pushing a crowd's density \(\rho\) forward out of the past. The left one is the Hamilton–Jacobi–Bellman equation of the second, pulling one agent's value function \(V\) backward out of its goals. What is new is that I have written them on the same line, and if you read their arguments you can see they are stuck to each other.

Look at the running cost \(q\) on the left: it takes \(\rho_t\) as an argument, so what it costs me to stand somewhere depends on how many other people are standing there. Now look at the drift on the right: it contains \(\nabla V\), so the crowd is not drifting anywhere, it is made of agents each solving the equation on the left. Neither can be solved first. The backward equation needs the crowd's entire future before it can price anything, and the forward equation needs the agents' entire plan before it can move anybody. Each one is the other's missing boundary data.

A pair like that is not a PDE problem in the sense the first two essays used the phrase. It is a fixed point in the space of (plan, crowd) pairs, and it comes from a field I have avoided until now. It is a game, and its solutions are Nash equilibria among a continuum of players.

This essay closes the sequence by coupling its first two members and looking at what the coupling makes: traffic that anticipates itself, a crowd that stampedes into one of two identical valleys on a coin flip, learning dynamics that orbit forever instead of settling, and, once the coupling is switched back off, one last appearance of the factorization this sequence keeps running into. I build the game theory from zero, because none of it is hard and most people were never taught it. As before, underlined terms open into derivations and definitions, and the figures carry a good part of the argument, so drag the sliders.


Part IWhat a game is

A , in the mathematical sense, needs three ingredients and no board: a set of players; for each player a set of available strategies; and for each player a payoff that depends on the strategies of everyone. That last clause is the whole subject. In an optimization problem my outcome depends on my choice, so "best" is well defined before I start. In a game my outcome depends on my choice and yours, so "best" is circular: the fastest route to work depends on the routes everyone else takes, and their routes depend on mine.

games, strategies, payoffs

Formally: players \(i = 1,\dots,n\), strategy sets \(A_i\), payoff functions \(J_i(a_1,\dots,a_n)\). The everyday examples are worth keeping around because they calibrate the abstraction: rock-paper-scissors (three strategies, payoffs from the cycle of who beats whom), picking a lane in traffic, picking a moment to leave a stadium, setting a price against competitors, and, in this essay's continuous-time version, picking an entire trajectory through a landscape while everyone else picks one too.

A mixed strategy is a probability distribution over one's options, and it is forced on us immediately. Rock-paper-scissors has no sensible deterministic recommendation, because any fixed choice is beaten by one specific reply. The recommendation "play each with probability \(1/3\)" is stable in a sense the next popup makes precise, and the fact that probability distributions turn out to be the natural strategy objects is why this subject was always going to collide with the rest of this sequence.

You cannot break the circle by chasing it. If I best-respond to your current plan, you will want to change yours, which invalidates mine, and there is no reason in general for that chase to terminate. Part IV is about games where it provably does not. What breaks the circle instead is a definition. Call a strategy a to the others' behaviour if no alternative improves my payoff while they hold still. A profile of strategies, one per player, is a Nash equilibrium when every player's strategy is a best response to everyone else's: a configuration with no regrets, in which each player, shown the full truth about the others, would change nothing. Equilibrium is a fixed point of the best-response map, and its existence in mixed strategies is Nash's 1950 theorem.

Nash equilibrium

A profile \((a_1^*,\dots,a_n^*)\) is a Nash equilibrium if for every player \(i\) and every alternative \(a_i\), \(J_i(a_i, a_{-i}^*) \le J_i(a_i^*, a_{-i}^*)\) (payoffs to be maximized; flip the inequality for costs). Nash proved existence for finite games in mixed strategies with a fixed-point theorem (Kakutani's, later simplified to Brouwer's), which is worth internalizing as a moral: equilibria are fixed points, so everything we know about fixed points transfers to them, including existence without uniqueness, sensitivity to perturbation, and hardness of computation. Uniqueness in particular is not part of the theorem, and Part III is about what buying it costs.

One more calibration, because the word carries unearned warmth. Equilibrium is a consistency condition, and consistency is a modest virtue. Nothing in the definition says equilibrium play is good for the players: the prisoner's dilemma has a unique equilibrium that everyone regrets jointly, and congestion games have equilibria strictly worse than what a coordinator could arrange, the gap being what people call the price of anarchy. The figures below show equilibria, and you should resist hearing applause in the word.

For two or three players this is manageable. The games in this essay have a million players, and there the honest formulation drowns: a million coupled best responses, each conditioned on the strategies of the other \(10^6 - 1\). The rescue is the move the first essay already made on a crowd of particles. , so my payoff depends on the others only through their distribution and not through their names. Then, as \(n\to\infty\), the game against a million opponents becomes a game against one object: the population density \(\rho\). I am small, the crowd is a field, my own choices do not move it, and its shape is all of it that touches me. This is a mean field game, and the limit is more than a convenience, because equilibria of the limit are approximate equilibria of the million-player game, with an error that vanishes as \(n\) grows.

anonymity, and where the theory comes from

The anonymity hypothesis: \(J_i = J\big(a_i,\; \tfrac{1}{n-1}\sum_{j\ne i}\delta_{a_j}\big)\), the same function for everyone, seeing only the empirical distribution of the others. It is exactly the hypothesis behind the first essay's McKean–Vlasov limit, promoted from dynamics to payoffs, and the limit theorem is again propagation of chaos, now for strategic behaviour rather than for motion.

The theory was created twice, independently and at the same time, around 2006: by Lasry and Lions in Paris, who wrote down the coupled PDE system in this essay's title and named the subject jeux à champ moyen, and by Huang, Caines and Malhamé in Montreal, who arrived at the same structure from large-population control engineering under the name Nash certainty equivalence. The probabilistic formulation, through coupled forward-backward stochastic differential equations, is due to Carmona and Delarue, whose two volumes are the standard reference. The quantitative version of "the limit is honest" is in the coda, with its fine print.

The first essay did something that looks like this and is not the same thing, and the difference is what forces the backward equation into the picture. In its Part IV each particle was pulled toward the mean position of all the others, the landscape a particle felt depended on \(\rho\), and Fokker–Planck became nonlinear in its own solution: the McKean–Vlasov equation. Two things separate that from what follows. First, the crowd entered there through the drift, as a force, and it enters here through the cost, as a preference, which is something an agent gets to respond to however it likes. Second, a McKean–Vlasov particle reacts to where the crowd is now, whereas an agent here reacts to where the crowd will be for the whole remaining horizon, because \(V\) is computed backward from \(T_{\!f}\) with the entire flow \((\rho_s)_{s\ge t}\) sitting in its coefficients. Interaction has become anticipation.


Part IIThe coupled system, and how to find its fixed point

Now assemble the object out of parts we already own. Fix, provisionally, a belief about the crowd: an entire flow of densities \(\bar\rho_t\), which you can read as the anticipated traffic report for the whole horizon. Against a fixed \(\bar\rho\), each agent faces an ordinary stochastic control problem, the second essay's, verbatim: dynamics \(\mathrm dX = (-\nabla U + u)\,\mathrm dt + \sqrt{2T}\,\mathrm dW\) in a landscape \(U\), effort priced quadratically at rate \(\alpha\), a terminal goal \(\Phi\), and a running cost that now reads the traffic report. The simplest such cost is a congestion charge

$$q(x, \bar\rho_t) \;=\; c\,\bar\rho_t(x),$$

a fee proportional to how crowded your location is at the moment you occupy it.

Count what survives from the second essay, because the answer is everything. The crowd entered only through the running cost, and the running cost was exactly the term that the Hopf–Cole substitution demoted to a killing rate. So with \(\bar\rho\) held fixed, every line of that essay still holds unedited: \(V = -\lambda\log\psi\) with the same channel condition \(\lambda = 2\alpha T\), the desirability \(\psi\) solving a linear backward equation with killing rate \(c\bar\rho_t/\lambda\), and each agent's optimal steering coming out as a score, \(u^{*} = 2T\nabla\log\psi\). From one agent's point of view the crowd is an externally scheduled tax on locations, and a tax schedule is no threat to linearity.

Then close the loop. If every agent steers by that \(u^{*}\), the crowd is a diffusion with that drift, so its density obeys the first essay's forward equation, which emits a new traffic report \(\rho^{\text{new}}\). The that defines equilibrium is that the report was right: \(\rho^{\text{new}} = \bar\rho\). The crowd everyone optimizes against has to be the crowd that optimizing produces. That sentence is the mean field game, and the two displayed equations at the top of the page are its differential form.

the coupled system, assembled

Backward, from essay two, with the crowd in the cost:

$$-\partial_t V = c\rho_t + \mathcal LV - \tfrac{1}{2\alpha}|\nabla V|^2, \qquad V(\cdot,T_{\!f})=\Phi, \qquad u^{*} = -\nabla V/\alpha.$$

Forward, from essay one, with the optimizers in the drift:

$$\partial_t\rho = \nabla\!\cdot\!\big(\rho\,(\nabla U - u^{*})\big) + T\Delta\rho, \qquad \rho(\cdot,0)=\rho_0.$$

Data at both ends of time and coupling in both directions. Compare this to the duality of essay two, where \(\rho\) also ran forward while \(u\) ran backward: there the two equations ignored each other completely and the only thing linking them was the pairing \(\int u\rho\), which is why it was conserved. Here each equation sits inside the other's coefficients, and nothing is conserved. That is the entire difference between a control problem and a game.

Notice what the linearization does and does not buy. Given \(\rho\), the equation for \(\psi\) is linear. Given \(V\), the equation for \(\rho\) is linear. Jointly they are not, because the object we want is a fixed point of the composition of those two maps, and the composition is nonlinear in the most awkward way available, through the coefficients of two PDEs rather than through a formula you could differentiate on paper. Every pleasant property of the second essay is a property of one half of the loop. None of them is a property of the loop.

Fixed points invite iteration, and the iteration with both a pedigree and a theorem is : solve the best response against yesterday's traffic, average it into the standing plan, repeat, and go on commuting forever against an ever-better forecast of yourself. The averaging is the load-bearing part. Best-responding to yesterday alone tends to overshoot, because everyone dodges the same jam on the same day and creates a new one somewhere else, and you get two rush hours chasing each other's ghosts. The first figure runs the averaged version live, one day per animation frame, and its right panel plots the number that measures how far from equilibrium we are: the exploitability, the amount a single deviant could still gain against the current crowd. At a Nash equilibrium, and only there, it is zero.

fictitious play, and exploitability

Brown proposed the scheme in 1951 for matrix games: each round, best-respond to the time average of the opponent's past play. For mean field games, Cardaliaguet and Hadikhanloo proved that the averaged scheme converges in the potential and monotone classes, both defined in the next two parts. That theorem is the licence under which Figure 1 operates, and outside those classes it is withdrawn.

Exploitability turns "how far from equilibrium" into a number: \(\mathcal E(\bar\rho) = J(\text{the population's own plan against } \bar\rho) - \min_u J(u \text{ against } \bar\rho)\). The subtrahend is one linear backward solve, by essay two; the minuend is an integral over the crowd's flow. It is nonnegative by construction and zero exactly at equilibrium, which makes it the analogue of the free energy in essay one: the scalar that certifies progress. One honesty note about the figure. Our grid, operator splitting and finite differences leave a floor of a few times \(10^{-2}\) under the curve that no amount of iteration removes. The fall to that floor is fictitious play working; the floor is discretization, and it would be dishonest to plot it on a log axis and let you read it as convergence to zero.

1.50 crowd \(\rho_t\) steering \(u^*\)

A crowd commuting from the left heap to the right slot, replayed daily; each day it best-responds to its own history and averages the result into its plan. With the charge at zero everyone marches in one dense column. Raise it and watch the column widen and stagger: the crowd is now spending effort to avoid itself, and the exploitability curve on the right certifies that nobody can do better alone.

Figure 1. A mean field game solved by commuting. Fictitious play as daily life.

Watch what the slider buys. At \(c=0\) this is the second essay exactly: everyone runs the same optimal control and is indifferent to company, so the crowd is a single sharp column with no interesting structure. At \(c=3\) the flow spreads in space and in time, some of the crowd leaving early, some late, some detouring through the more expensive terrain. Nobody instructed it to take turns, and no strategy anywhere in the system contains the concept of taking turns. Staggering is what a fixed point of selfishness looks like once proximity has a price, which is to say it is rush hour, derived.


Part IIIOne equilibrium, or several

Nash's theorem grants existence and says nothing about uniqueness, and the silence is load-bearing. The known sufficient condition is due to Lasry and Lions, and it has a name that sounds technical and a meaning that does not: , which for our congestion coupling amounts to \(c \ge 0\). If people dislike crowds, the equilibrium is unique. The mechanism is self-correction: any discrepancy between two candidate equilibria makes the fuller one more expensive, which pushes agents out of it, which closes the discrepancy. Written out, the proof is a pairing argument that runs on the same adjointness the second essay used to show \(\int u\rho\) was conserved, with an inequality where that argument had an equality.

the Lasry–Lions monotonicity argument

The coupling \(q(x,\rho)\) is monotone if \(\int \big(q(x,\rho)-q(x,\rho')\big)\,\mathrm d(\rho-\rho')(x) \;\ge\; 0\) for every pair of densities: crowding costs more where the crowd is bigger. For \(q = c\rho\) the integral is \(c\int(\rho-\rho')^2\), which is nonnegative exactly when \(c\ge0\).

Sketch of the uniqueness proof: take two equilibria \((V,\rho)\) and \((V',\rho')\), write the equations for the differences, pair the HJB difference against \(\rho-\rho'\) and the Fokker–Planck difference against \(V-V'\), integrate over space and time, and add. The transport terms cancel by adjointness, \(\int (\mathcal L\varphi)\rho = \int \varphi\,(\mathcal L^{*}\rho)\), which is the identity behind the conservation law of essay two. What survives is the monotonicity integral plus a nonnegative term coming from the convexity of the effort cost. A sum of nonnegative things equal to zero forces every one of them to vanish, and \(\rho = \rho'\) follows. That is the flavour of the whole field: forward-backward duality plus one sign condition.

Flip the sign to \(c<0\) and the proof dies. The next figure shows that the conclusion dies with it.

So flip the sign and see what breaks. With \(c<0\), proximity is a reward and the coupling is a herding instinct. Take a symmetric landscape with two identical shallow valleys, start the crowd spread evenly across both, and give it no destination at all, so that the only thing anybody wants is company. Every symmetry of the problem says the symmetric flow is an equilibrium, and it is, exactly. Past a critical coupling strength it is also useless: unstable, exploitable by any agent willing to bet on which side the crowd will pick, and abandoned by fictitious play at the first breath of asymmetry. What the iteration finds instead is one of two mirror-image equilibria, the entire population in the left valley or the entire population in the right, with nothing in the problem to prefer either. Run the figure a few times and the tally on the right is the coin's ledger.

If you read the first essay this picture is familiar, and I want to be careful about how far the resemblance goes. Its Figure 9 cooled a McKean–Vlasov crowd through a critical temperature and watched the symmetric state destabilize into a pitchfork, and we filed it under statistical mechanics: spontaneous symmetry breaking, magnetization, a phase transition. The shared structure is exact, and it is this. In both cases the interaction contributes a term to a functional that is quadratic in \(\rho\) (there \(\tfrac12\iint W(x-y)\rho(x)\rho(y)\) with \(W\) an attraction to the mean, here \(\tfrac{c}{2}\int\rho^2\) with a local kernel). In both cases attraction makes that term concave, the sum stops being convex, and uniqueness dies through the same bifurcation.

The differences deserve naming, because "phase transition equals equilibrium selection" is the sort of slogan that goes soft when nobody checks it. The kernels are not the same object, nonlocal there and a delta here. The crowd in Figure 9 was myopic, descending a free energy with no view of the future, while this crowd anticipates an entire horizon, so the functional here carries essay two's control cost as well. And the sharpest difference: the stationary states of the first essay were minimizers of the free energy, whereas mean field game equilibria are only critical points of the corresponding functional, which is a strictly weaker demand and admits configurations that no descent would ever settle into. Part IV picks that thread up.

With those qualifications in hand, the reframing survives, and it is the reason this figure is here. Attraction to the mean is a payoff coupling. The ferromagnet's ordered states are coordinated equilibria. The phase transition is the moment at which coordinating starts to pay. And the coin flip that physics calls symmetry breaking, game theory calls equilibrium selection, where it is an old embarrassment rather than a familiar fact of life: when the theory admits several self-consistent worlds it contains no principle that picks one, and on screen the choice is made by whatever noise happened to be in the initial condition.

-6.0 final crowd initial crowd

Crowd-averse (c > 0): fictitious play returns the symmetric answer every time, as the uniqueness theorem promises. Crowd-seeking (c < 0, strong): the symmetric answer is still perfectly self-consistent and never survives, and a 5% whisper of noise in the initial crowd decides which valley gets everyone. The tally on the right is the coin's record.

Figure 2. Monotonicity, and its counterexample. The first essay's phase transition, re-filed as equilibrium selection.

Part IVGames that refuse to be optimizations

The first essay's central theorem said that a crowd's evolution was secretly one optimization: steepest descent of a free energy, in the geometry of optimal transport. The natural hope is that games work the same way, with every equilibrium the critical point of some grand functional, and for one class of games the hope is exactly right. In a , all the players' incentives are slices of a single function, equilibria are its critical points, and the machinery of essay one switches back on: fictitious play descends the potential the way the JKO scheme descended the free energy. That is why Figures 1 and 2 converge, and it is the entire content of the licence I cited for them.

potential games

Monderer and Shapley: a game is a potential game if there is one function \(\mathcal P\) of the full strategy profile such that every player's payoff change under a unilateral deviation equals the change in \(\mathcal P\). Congestion games are the canonical class, by Rosenthal's theorem, which is why traffic keeps showing up in this essay. Equilibria are critical points of \(\mathcal P\), better-response dynamics ascend \(\mathcal P\) and therefore terminate, and the game is an optimization wearing a disguise of multiplicity.

The mean-field version: when the coupling has the form \(q(x,\rho) = \frac{\delta \mathcal Q}{\delta\rho}(x)\) for some functional \(\mathcal Q\) (our \(c\rho\) qualifies, with \(\mathcal Q = \tfrac{c}{2}\int\rho^2\)), the game is potential, and equilibria are critical points of a single functional over flows of measures: state cost, plus \(\mathcal Q\), plus the entropic terms of essay two. One precision flag, because the literature blurs it and Part III leaned on the distinction. The minimizers of that functional solve a mean-field control problem, which is a benevolent planner moving the whole crowd at once, while game equilibria are its critical points. The two coincide only sometimes, and the gap between them is exactly the price of anarchy in this setting.

In general the hope fails, and it fails with structure. Every finite game splits into a potential part and a , and the harmonic part is not the gradient of anything. Its purest specimen is matching pennies, the two-player game of pure opposition, where my win is your loss and every configuration has somebody in it who regrets their choice. Chasing that regret forever is not a failure to find the equilibrium. It is the shape of the incentives.

the harmonic part, and learning that circles

On finite games there is an orthogonal decomposition, due to Candogan, Menache, Ozdaglar and Parrilo, of every game into a potential component and a harmonic component, with zero-sum games like matching pennies as the harmonic archetype. Under gradient-style learning the potential component pulls toward equilibria and the harmonic component rotates around them, conserving a Hamiltonian-like quantity rather than dissipating one.

A confession from the making of the figure below. Our first version showed the orbit failing to settle even at moderate rationality, which would have been a lovely result and was an artifact: an explicit Euler scheme going unstable on a fast spiral and manufacturing a limit cycle the mathematics does not contain. It now runs Runge–Kutta, under which the theorem (Hofbauer and Sandholm: perturbed best-response dynamics converge globally in zero-sum games) holds on screen at every temperature the slider offers. Numerical artifacts are fondest of exactly the claims one most wants to believe.

Readers of the first essay will want to map this onto its coda, and the map is a good one as long as you are precise about it. There I split a drift into a gradient piece, which descended the free energy, and a rotational piece, which did no work on it and circulated forever, and the two panels of Figure 10 had identical stationary densities and completely different lives. Here the same split happens to the vector field that learning follows through the space of mixed strategies. What is identical: a vector field decomposes into a gradient part that dissipates a scalar and a divergence-free part that conserves one, and it is the second part that keeps the dynamics from stopping. What is different is not a detail. In the first essay the field failed to be a gradient because the force was non-conservative, with some outside agency stirring the system and paying for it in dissipated heat. Here nothing is being stirred: the field is assembled from several players differentiating several different payoff functions, so it is a pseudo-gradient, and it fails to be a gradient for a structural reason internal to the incentives. Zero-sum is the extreme case, where my gain is exactly your loss, so no scalar can be increasing for both of us and rotation is all that remains.

This is also why simultaneous gradient learning oscillates in adversarial settings, which is a daily engineering complaint under the name of GAN training instability, and the standard remedies (averaging, optimism, extragradient steps) are all ways of taxing the rotation.

What tames the rotation, in theory and in the figure, is this sequence's oldest friend. Soften the best response into a : instead of playing the argmax, play each action with probability proportional to \(e^{-(\text{its cost})/\lambda}\), a Boltzmann distribution over your own options. The rest point where everyone is logit-responding to everyone else's logit response is the quantal response equilibrium, and the temperature does to the eternal orbit what it has done everywhere else in this sequence. The circulation does not disappear, and the softening adds an inward pull on top of it, so the trajectory spirals into the QRE for every \(\lambda>0\). Only at exactly zero temperature does the circle close and the game refuse forever to sit down.

the third appearance of \(\lambda\)

Quantal response equilibrium is due to McKelvey and Palfrey (1995), as a model of boundedly rational play that fits human experimental data considerably better than exact Nash. Its logit form is a fixed point of softmax responses, and the sequence's ledger is now complete enough to read out. Essay one: \(e^{-V/T}/Z\), the Boltzmann distribution over states. Essay two: \(\pi(a) \propto e^{-Q(a)/\lambda}\), the soft-Bellman policy over actions, and \(V = -\lambda\log Z\) as its free energy. Essay three: the QRE, a Boltzmann distribution over strategies whose energies are computed against everyone else's Boltzmann distribution. One functional form, three essays, one dial, and in every case the same trade: a little optimality surrendered in exchange for smoothness, uniqueness, and the ability of a dynamics to actually arrive somewhere.

The self-consistency is the new part. In the first two essays the temperature softened a fixed objective. Here it softens an objective that is itself a function of the softened play, so \(\lambda\) appears inside its own argument, and the fixed point is what makes the QRE a different object in kind, and not just a Nash equilibrium with the corners rounded off. As \(\lambda\to0\) the QRE branch converges to a specific Nash equilibrium, which is one of the few principled answers anyone has to the selection problem of Part III.

0.25 learning orbit equilibrium

The square is the space of mixed strategies: each axis is one player's probability of playing heads. In matching pennies the incentive field is almost pure rotation, so cool \(\lambda\) toward zero and watch the spiral flatten toward a closed orbit, which is the shape of pure opposition. Toggle to the coordination game and the same dynamics runs straight into a corner, because that game is all potential and no rotation.

Figure 3. Learning in strategy space. The first essay's rotation-versus-gradient decomposition, now for incentives.

Part VSwitching the game off: the bridge

One experiment remains, and it has been owed since the first essay's coda. Take the system of Part II and delete the game: coupling zero, no landscape, no goals along the way. What is left is a crowd of independent diffusers and two non-negotiable facts, where the crowd starts (\(\rho_0\)) and where it must end (\(\rho_1\)). The question stops being strategic and becomes logistical. Among all the ways a diffusing crowd can begin as \(\rho_0\) and end as \(\rho_1\), I want the most probable history, which by the second essay's accounting is the one costing the least steering against the noise. This is , and the degenerate mean field game hands over its answer using machinery we have had all essay: the solution is a pair of potentials, one running forward, one running backward, with the crowd at every intermediate moment given by their pointwise product,

the Schrödinger system, and Sinkhorn

Schrödinger's setup: a cloud of independent Brownian particles is observed as \(\rho_0\) at time 0 and, surprisingly, as \(\rho_1\) at time 1, and he wanted the most likely history joining the two observations. The answer (whose large-deviations formulation is Föllmer's) minimizes \(\mathrm{KL}\) from the free diffusion subject to both marginals, which is essay two's control cost with two pinned ends rather than one, and the optimality system is

$$\rho_t = \varphi_t\,\psi_t, \qquad \partial_t\varphi = \mathcal L^{*}\varphi, \qquad \partial_t\psi = -\mathcal L\psi,$$

one factor riding the forward equation, one riding the backward equation, coupled only through the boundary requirements \(\varphi_0\psi_0 = \rho_0\) and \(\varphi_1\psi_1 = \rho_1\). The numerical method is to alternate: fix \(\psi\) and solve for \(\varphi\) to match the left marginal, fix \(\varphi\) and solve for \(\psi\) to match the right, repeat. This is the Sinkhorn algorithm, also known as iterative proportional fitting, the workhorse of computational optimal transport.

It is tempting to call this fictitious play's cousin, and the resemblance is real but shallower than it looks, in a way worth stating. Both alternate between two conditions that cannot be satisfied in one step. But Sinkhorn's two conditions are affine constraints and the problem is a convex projection, so Csiszár's theorem gives convergence to a unique fixed point whenever the two constraints are jointly feasible at all. Fictitious play alternates between a crowd and a plan on a problem that is convex only under the monotonicity of Part III, and outside that class it has no such guarantee. Alternating projection is the easy relative; the game is the one with the fine print.

As the diffusivity \(\varepsilon\to0\), the bridge converges to the deterministic optimal-transport geodesic, McCann's displacement interpolation, which closes the loop back to the first essay's Wasserstein geometry from one dial away.

$$\rho_t(x) \;=\; \varphi_t(x)\,\psi_t(x).$$

That product is worth sitting with, because it is the third time this sequence has produced a forward object and a backward object that belong together, and the three couplings are different in an instructive way. In the second essay, \(\rho\) ran forward and \(u\) ran backward and neither knew about the other, so their integral \(\int u\rho\) was conserved. In Parts II to IV, each ran inside the other's coefficients, and nothing was conserved. Here they are pinned at the two ends of time and multiply pointwise, and their product is the answer itself.

There is a fourth member of this family that I have not built, and I mention it because you will meet the factorization the moment you touch filtering theory. The smoothed estimate of a hidden state given data on both sides of it factorizes as \(p(x_t \mid \text{all data}) \propto \alpha_t\,\beta_t\), a forward filter times a backward likelihood, which is the forward-backward algorithm from hidden Markov models and the Rauch–Tung–Striebel smoother in its Gaussian case. Same shape, different pinning: the bridge is pinned by decree at two times, the smoother is pinned by data at many. Filtering, control and entropic transport are one forward-backward factorization wearing three costumes, and I still owe that essay.

The figure below computes the bridge by Sinkhorn iteration, live, with the diffusivity on a slider, so the last thing this sequence puts on screen is stochastic transport tightening into the deterministic geodesic where the first essay began.

0.40 0.00 bridge \(\rho_t=\varphi_t\psi_t\) pinned ends

Both ends are decreed: a heap on the left has to become the two-humped shape on the right. Sinkhorn negotiates the factors \(\varphi\) and \(\psi\) until both decrees hold, with the marginal errors landing at machine precision, and the animated curve is their product. Cool \(\varepsilon\) and the cloud of possible histories tightens toward the unique cheapest transport.

Figure 4. The mean field game with the game removed: Schrödinger's bridge, and the first essay's geometry one dial away.

CodaThe fine print, and the end of the sequence

Fine print in three clauses, as always.

The solver has a licence, and it is narrow. Fictitious play provably converges for potential and monotone games, which is where every figure above lives, deliberately. Outside those two classes it can cycle or diverge, and for properly adversarial couplings Part IV says it must. A figure that converged is not a theorem that it had to.

The crowd here shares no weather. Our million agents face independent noise, which is why the limiting density is deterministic and why a single flow \(\bar\rho\) was enough to describe the future. Let them share a common shock, a market crash or a storm over the whole city, and \(\rho\) itself becomes random. The value function then has to depend on the entire current distribution, \(V(x,\rho,t)\), and its equation, the master equation, lives on the space of measures. That is the deep end of the field, it is where Lions's lectures pointed, and it is exiled from this essay by this paragraph.

The limit has a price tag, and the usual quotation of it is too generous. The standard statement is that a mean-field equilibrium, replayed by \(n\) real players, is an \(\varepsilon\)-Nash equilibrium with \(\varepsilon = O(1/\sqrt n)\): no player can gain more than that by deviating. That rate is the one you get when the master equation has a classical solution, which requires regularity assumptions that the monotone case supplies and the general case does not; without them the known rates are worse and dimension-dependent. Either way the theorem is the licence for everything above, and the rate is a warning that a boardroom of five is not a crowd of a million. The idealization earns its keep only in the second regime.

The road onward is short to name, because from here the subject fans out into the world: crowd dynamics and evacuation planning, macroeconomic models in which the distribution of wealth is the state variable, epidemic response as a game between behaviour and contagion, electricity markets, multi-agent reinforcement learning, and the training dynamics of adversarial networks, where Part IV's rotation is somebody's Tuesday. The mathematics in every one of those cases is the pair of equations at the top of this page.

And the sequence closes. Three essays, three tenses of one grammar. The forward equation: what will be. The backward equation: what to do. Their fixed point: what everyone will do about everyone. One crowd of particles became a crowd of futures and then a crowd of optimizers, and at every step the same small cast walked back on stage: a Boltzmann distribution, a partition function nobody wanted to compute, a score pointing uphill in log-probability, and a temperature trading sharpness for the ability to arrive. I claimed in the first essay that one equation was secretly doing something more significant than it looked. Three essays later the honest summary is that it was doing prediction, volition and society, each of them the same descent dressed for a different occasion. Praise, at the last, for the duality that ran through all of them rather than for any single equation: the forward and the backward, the crowd and the goal, each holding the other's boundary conditions, meeting where everything in this sequence has met, at the only moment that exists.