Dynamics-aware local trajectory control of a tracked vineyard robot via deep reinforcement learning
Abstract
Autonomous navigation between narrow vineyard rows requires respecting the platform’s dynamics, which trajectory optimizers such as CHOMP and STOMP resolve in batch before motion and again whenever the map changes. We instead cast local navigation as a continuous-control Markov decision process whose actions are the per-track torques of a skid-steer robot, so the platform’s dynamics enter the control law itself, and the optimization is paid once, during training. Soft Actor-Critic (SAC) and Twin Delayed Deep Deterministic policy gradient (TD3), both recurrent, and a feed-forward Proximal Policy Optimization (PPO) baseline are trained over ten seeds in simulation from a real vineyard passability map. On two held-out scenarios all three produce shorter routes than a conservative weighted A* reference at higher peak but comparable average impassability, and the safety ranking of the off-policy agents reverses between scenarios.
© 2026 Filip Zúbek, Oliver Halaš, Vendelín František Skokan, Ladislav Körösi, Ondrej Straka, Aleš Melichár, Martin Dekan, published by Slovak University of Technology in Bratislava
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License.