This paper describes a novel machine learning framework for solving sequential decision problems called Markov decision processes (MDPs) by iteratively computing low-dimensional representations and approx imately optimal policies. A unified mathematical framework for learn ing representation and optimal control in MDPs is presented based on a class of singular operators called Laplacians, whose matrix repre sentations have nonpositive off-diagonal elements and zero row sums. Exact solutions of discounted and average-reward MDPs are expressed in terms of a generalized spectral inverse of the Laplacian called the Drazin inverse. A generic algorithm called representation policy iter ation (RPI) is presented which interleaves computing low-dimensional representations and approximately optimal policies. Two approaches for dimensionality reduction of MDPs are described based on geometric and reward-sensitive regularization, whereby low-dimensional represen tations are formed by diagonalization or dilation of Laplacian opera tors. Model-based and model-free variants of the RPI algorithm are presented; they are also compared experimentally on discrete and continuous MDPs. Some directions for future work are finally outlined.
Article navigation
2 June 2009
Research Article|
June 02 2009
Learning Representation and Control in Markov Decision Processes: New Frontiers
Mahadevan Sridhar
Mahadevan Sridhar
Department of Computer Science, University of Massachusetts — Amherst, 140 Governor’s Drive
, Amherst, MA 01003, USA
Search for other works by this author on:
Online ISSN: 1935-8245
Print ISSN: 1935-8237
© 2010 S. Mahadevan
2010
S. Mahadevan
Licensed re-use rights only
Foundations and Trends in Machine Learning (2009) 1 (4): 403–565.
Citation
Sridhar M (2009), "Learning Representation and Control in Markov Decision Processes: New Frontiers". Foundations and Trends in Machine Learning, Vol. 1 No. 4 pp. 403–565, doi: https://doi.org/10.1561/2200000003
Download citation file:
New and popular articles
Suggested Reading
From management accountant networking to firm performance: the mediating role of innovative management accounting tool and the moderating role of innovation-oriented culture
Asia-Pacific Journal of Business Administration (August,2025)
PROJECTIVE METHODS FOR LAPLACIAN FIELDS
COMPEL (April,1985)
Eigensolution of Laplacian matrices for graph partitioning and domain decomposition: Approximate algebraic method
Engineering Computations (October,2009)
Related Chapters
Markov Switching Rationality
Essays in Honor of Joon Y. Park: Econometric Methodology in Empirical Applications
Option Pricing with Markov Switching Stochastic Volatility Models
Advances in Pacific Basin Business, Economics and Finance
Markov Switching Models in Empirical Finance
Missing Data Methods: Time-Series Methods and Applications
Recommended for you
These recommendations are informed by your reading behaviors and indicated interests.
