Algorithms and Bounds for Rollout Sampling Approximate Policy Iteration

Dimitrakakis, Christos; Lagoudakis, Michail G.

doi:10.1007/978-3-540-89722-4_3

Christos Dimitrakakis³ &
Michail G. Lagoudakis⁴

Part of the book series: Lecture Notes in Computer Science ((LNAI,volume 5323))

Included in the following conference series:

European Workshop on Reinforcement Learning

1115 Accesses
5 Citations

Abstract

Several approximate policy iteration schemes without value functions, which focus on policy representation using classifiers and address policy learning as a supervised learning problem, have been proposed recently. Finding good policies with such methods requires not only an appropriate classifier, but also reliable examples of best actions, covering the state space sufficiently. Up to this time, little work has been done on appropriate covering schemes and on methods for reducing the sample complexity of such methods, especially in continuous state spaces. This paper focuses on the simplest possible covering scheme (a discretized grid over the state space) and performs a sample-complexity comparison between the simplest (and previously commonly used) rollout sampling allocation strategy, which allocates samples equally at each state under consideration, and an almost as simple method, which allocates samples only as needed and requires significantly fewer samples.

This project was partially supported by the ICIS-IAS project and the Marie Curie International Reintegration Grant MCIRG-CT-2006-044980 awarded to Michail G. Lagoudakis within the 6th European Framework Programme.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Institutional subscriptions

Preview

Unable to display preview. Download preview PDF.

Importance sampling in reinforcement learning with an estimated behavior policy

Article Open access 07 May 2021

A Survey on Constraining Policy Updates Using the KL Divergence

Sample complexity of variance-reduced policy gradient: weaker assumptions and lower bounds

Article 27 June 2024

References

Auer, P., Cesa-Bianchi, N., Fischer, P.: Finite-time analysis of the multiarmed bandit problem. Machine Learning Journal 47(2-3), 235–256 (2002)
Article MATH Google Scholar
Auer, P., Ortner, R., Szepesvari, C.: Improved Rates for the Stochastic Continuum-Armed Bandit Problem. In: Bshouty, N.H., Gentile, C. (eds.) COLT 2007. LNCS, vol. 4539, pp. 454–468. Springer, Heidelberg (2007)
Google Scholar
Bertsekas, D.: Dynamic programming and suboptimal control: From ADP to MPC. Fundamental Issues in Control, European Journal of Control 11(4-5) (2005); From 2005 CDC, Seville, Spain
Google Scholar
Dimitrakakis, C., Lagoudakis, M.: Rollout sampling approximate policy iteration. Machine Learning 72(3) (September 2008)
Google Scholar
Even-Dar, E., Mannor, S., Mansour, Y.: Action elimination and stopping conditions for the multi-armed bandit and reinforcement learning problems. Journal of Machine Learning Research 7, 1079–1105 (2006)
MathSciNet MATH Google Scholar
Fern, A., Yoon, S., Givan, R.: Approximate policy iteration with a policy language bias. Advances in Neural Information Processing Systems 16(3) (2004)
Google Scholar
Fern, A., Yoon, S., Givan, R.: Approximate policy iteration with a policy language bias: Solving relational Markov decision processes. Journal of Artificial Intelligence Research 25, 75–118 (2006)
MathSciNet MATH Google Scholar
Kocsis, L., Szepesvári, C.: Bandit based Monte-Carlo planning. In: Fürnkranz, J., Scheffer, T., Spiliopoulou, M. (eds.) ECML 2006. LNCS, vol. 4212, pp. 282–293. Springer, Heidelberg (2006)
Chapter Google Scholar
Lagoudakis, M.G., Parr, R.: Reinforcement learning as classification: Leveraging modern classifiers. In: Proceedings of the 20th International Conference on Machine Learning (ICML), Washington, DC, USA, pp. 424–431 (August 2003)
Google Scholar
Langford, J., Zadrozny, B.: Relating reinforcement learning performance to classification performance. In: Proceedings of the 22nd International Conference on Machine learning (ICML), Bonn, Germany, pp. 473–480 (2005)
Google Scholar

Download references

Author information

Authors and Affiliations

Informatics Institute, University of Amsterdam, Amsterdam, The Netherlands
Christos Dimitrakakis
Department of ECE, Technical University of Crete, Chania, 73100, Greece
Michail G. Lagoudakis

Authors

Christos Dimitrakakis
View author publications
You can also search for this author in PubMed Google Scholar
Michail G. Lagoudakis
View author publications
You can also search for this author in PubMed Google Scholar

Editor information

Editors and Affiliations

INRIA Lille-Nord Europe, 59650, Villeneuve d’Ascq, France
Sertan Girgin
INRIA, LIFL, CNRS, Université de Lille, Villeneuve d’Ascq, France
Manuel Loth , Rémi Munos , Philippe Preux & Daniil Ryabko , , &

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Dimitrakakis, C., Lagoudakis, M.G. (2008). Algorithms and Bounds for Rollout Sampling Approximate Policy Iteration . In: Girgin, S., Loth, M., Munos, R., Preux, P., Ryabko, D. (eds) Recent Advances in Reinforcement Learning. EWRL 2008. Lecture Notes in Computer Science(), vol 5323. Springer, Berlin, Heidelberg. https://doi.org/10.1007/978-3-540-89722-4_3

Download citation

DOI: https://doi.org/10.1007/978-3-540-89722-4_3
Publisher Name: Springer, Berlin, Heidelberg
Print ISBN: 978-3-540-89721-7
Online ISBN: 978-3-540-89722-4
eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics

Algorithms and Bounds for Rollout Sampling Approximate Policy Iteration

Abstract

Access this chapter

Preview

Similar content being viewed by others

Importance sampling in reinforcement learning with an estimated behavior policy

A Survey on Constraining Policy Updates Using the KL Divergence

Sample complexity of variance-reduced policy gradient: weaker assumptions and lower bounds

References

Author information

Authors and Affiliations

Editor information

Editors and Affiliations

Rights and permissions

Copyright information

About this paper

Cite this paper

Download citation

Publish with us

Navigation

Algorithms and Bounds for Rollout Sampling Approximate Policy Iteration

Abstract

Access this chapter

Preview

Similar content being viewed by others

Importance sampling in reinforcement learning with an estimated behavior policy

A Survey on Constraining Policy Updates Using the KL Divergence

Sample complexity of variance-reduced policy gradient: weaker assumptions and lower bounds

References

Author information

Authors and Affiliations

Editor information

Editors and Affiliations

Rights and permissions

Copyright information

About this paper

Cite this paper

Download citation

Share this paper

Publish with us

Search

Navigation