Inverting the Bellman Equation: From $Q$-Values to World Models
By Alistair Letcher, Mattie Fellows, Alexander D. Goldie, Jonathan Richens, Jakob N. Foerster, Oliver Richardson
This paper proves that value-based RL agents trained over a sufficiently rich set of reward functions implicitly encode a unique, accurate world model, and introduces P-learning to extract that model as an inverse of Q-learning. It bridges the model-free versus model-based dichotomy with theoretical conditions on goal/reward diversity.