Speaker
Dr Alan King,
Room “Sala Seminari” - Abacus Building (U14)
State, Horizon, and Value: Comparing Multistage
Stochastic Programming and Reinforcement Learning
Abstract
Multistage stochastic programming (MSP) and reinforcement learning (RL) both solve sequential decisions under uncertainty. Both organize the problem around a state. Both, in the end, reduce to approximating a value function. But the two fields grew up almost without speaking to each other. A student trained in one often cannot read a paper in the other. Why is that?
We will take Chapter 6 of King and Wallace’s Modeling with Stochastic Programming as our guide. We will unfold a capacity inventory model to develop the core concepts: information states built from forecast/surprise decompositions, nonanticipativity on a scenario tree, and Grinold's dual equilibrium device for terminating an infinite-horizon problem. We then turn to King’s 2016 tutorial to see the same problem from the algorithmic side: risk measures, time consistency, and stochastic dual dynamic programming (SDDP), a value-function approximation scheme based on nested decomposition.
We then compare MSP scenario-tree/value-function formalism with the Markov decision process (MDP) formalism underlying RL. To answer this properly we derive the finite- and infinite-horizon Bellman equations. It turns out the MSP nested decomposition is a Bellman equation, just indexed by scenario-tree nodes rather than Markov states — and dual equilibrium and Bellman-equation bootstrapping turn out to be two structurally different answers to the same question: what value do we assign beyond a truncation point? The talk closes by asking which questions about state, about what a decision is allowed to depend on, and about dynamic structure survive the translation between the two languages.
Short bio
Alan King had a long and distinguished career at IBM Research in Yorktown Heights, New York.
At IBM, Alan participated in many research programs, including neural network proxy models, cryptocurrencies, and massively parallel computing, as manager, senior software engineer, lead consultant, university relations, and research staff advisor to senior leadership.
His research interest concerns the modeling and solution of stochastic programs, a branch of decision-making under uncertainty that applies to planning and operations. His contributions in this field include analysis of variance for solutions, software for solvers and tools, and modeling methodology.
Today, Alan is exploring how time series foundation models can be used to represent probability distributions in stochastic programs, with applications in energy and finance.
contact person for this Seminar: enza.messina@unimib.it