References
- Atiya, A.F., Parlos, A.G. and Ingber, L. (2003). A reinforcement learning method based on adaptive simulated annealing,, pp. 121-124.
- Barto, A., Sutton, R. and Anderson, C. (1983). Neuronlike adaptive elements that can solve difficult learning problem,(5): 834-847.
- Cichosz, P. (1995). Truncating temporal differences: On the efficient implementation of() for reinforcement learning,: 287-318.
- Crook, P. and Hayes, G. (2003). Learning in a state of confusion: Perceptual aliasing in grid world navigation,, University of Edinburgh, Edinburgh.
- Ernst, D., Geurts, P. and Wehenkel, L. (2005). Tree-based batch mode reinforcement learning,: 503-556.
- Forbes, J. R. N. (2002)., Ph.D. thesis, University of California, Berkeley, CA.
- Gelly, S. and Silver, D. (2007). Combining online and offline knowledge in UCT,, pp. 273-280.
- Kaelbing, L.P., Litman, M.L. and Moore, A.W. (1996). Reinforcement learning: A survey,(1): 237-285.
- Krawiec, K., Jaśkowski, W.G. and Szubert, M.G. (2011). Evolving small-board Go players using coevolutionary temporal difference learning with archives,(4): 717-731, DOI: 10.2478/v10006-011-0057-3.
- Lagoudakis, M. and Parr, R. (2003). Least-squares policy iteration,: 1107-1149.
- Lanzi, P. (2000). Adaptive agents with reinforcement learning and internal memory,, pp. 333-342.
- Lin, L.-J. (1993)., Ph.D. thesis, Carnegie Mellon University, Pittsburgh, PA.
- Markowska-Kaczmar, U. and Kwaśnicka, H. (2005)., Wrocław University of Technology Press, Wrocław, (in Polish).
- Moore, A. and Atkeson, C. (1993). Prioritized sweeping: Reinforcement learning with less data and less time,(1): 103-130, DOI: 10.1007/BF00993104.
- Moriarty, D., Schultz, A. and Grefenstette, J. (1999). Evolutionary algorithms for reinforcement learning,: 241-276.
- Peng, J. and Williams, R. (1993). Efficient learning and planning within the Dyna framework,(4): 437-454.
- Reynolds, S. (2002). Experience stack reinforcement learning for off-policy control,, University of Birmingham, Birmingham, ftp://ftp.cs.bham.ac.uk/pub/tech-reports/2002/CSRP-02-01.ps.gz.
- Riedmiller, M. (2005). Neural reinforcement learning to swing-up and balance a real pole,, pp. 3191-3196.
- Rummery, G. and Niranjan, M. (1994). On-line q-learning using connectionist systems,, Cambridge University, Cambridge.
- Smart, W. and Kaelbing, L. (2002). Effective reinforcement learning for mobile robots,, pp. 3404-3410.
- Sutton, R. (1990). Integrated architectures for learning, planning, and reacting based on approximating dynamic programming,, pp. 216-224.
- Sutton, R. (1991). Planning by incremental dynamic programming,, pp. 353-357.
- Sutton, R. and Barto, A. (1998)., MIT Press, Cambridge, MA.
- Vanhulsel, M., Janssens, D. and Vanhoof, K. (2009). Simulation of sequential data: An enhanced reinforcement learning approach,(4): 8032-8039.
- Watkins, C. (1989)., Ph.D. thesis, Cambridge University, Cambridge.
- Whiteson, S. (2012). Evolutionary computation for reinforcement learning,M. Wiering and M. van Otterlo (Eds.),, Springer, Berlin, pp. 325-358.
- Whiteson, S. and Stone, P. (2006). Evolutionary function approximation for reinforcement learning,: 877-917.
- Ye, C., Young, N.H.C. and Wang, D. (2003). A fuzzy controller with supervised learning assisted reinforcement learning algorithm for obstacle avoidance,(1): 17-27.
- Zajdel, R. (2012). Fuzzy epoch-incremental reinforcement learning algorithm,L. Rutkowski, M. Korytkowski, R. Scherer, R. Tadeusiewicz, L.A. Zadeh and J.M. Zurada (Eds.),, Lecture Notes in Computer Science, Vol. 7267, Springer-Verlag, Berlin/Heidelberg, pp. 359-366.
Language: English
Page range: 623 - 635
Published on: Sep 30, 2013
Published by: University of Zielona Góra
In partnership with: Paradigm Publishing Services
Publication frequency: 4 issues per year
Related subjects:
© 2013 Roman Zajdel, published by University of Zielona Góra
This work is licensed under the Creative Commons License.