- Editors' Suggestion
- Access by Xinjiang University
Controlled gliding and perching through deep-reinforcement-learning
Phys. Rev. Fluids 4, 093902 – Published 6 September, 2019
DOI: https://doi.org/10.1103/PhysRevFluids.4.093902
Abstract
Controlled gliding is one of the most energetically efficient modes of transportation for natural and human powered fliers. Here we demonstrate that gliding and landing strategies with different optimality criteria can be identified through deep-reinforcement-learning without explicit knowledge of the underlying physics. We combine a two-dimensional model of a controlled elliptical body with deep-reinforcement-learning (D-RL) to achieve gliding with either minimum energy expenditure, or fastest time of arrival, at a predetermined location. In both cases the gliding trajectories are smooth, although energy/time optimal strategies are distinguished by small/high frequency actuations. We examine the effects of the ellipse's shape and weight on the optimal policies for controlled gliding. We find that the model-free reinforcement learning leads to more robust gliding than model-based optimal control strategies with a modest additional computational cost. We also demonstrate that the gliders with D-RL can generalize their strategies to reach the target location from previously unseen starting positions. The model-free character and robustness of D-RL suggests a promising framework for developing robotic devices capable of exploiting complex flow environments.
Physics Subject Headings (PhySH)
Article Text
References (58)
- C. D. Cone, Thermal soaring of birds, Am. Sci. 50, 180 (1962).
- A. Azuma and Y. Okuno, Flight of a samara, Alsomitra macrocarpa, J. Theor. Biol. 129, 263 (1987).
- R. Dudley, G. Byrnes, S. Yanoviak, B. Borrell, R. Brown, and J. A. McGuire, Gliding and the Functional Origins of Flight: Biomechanical Novelty or Necessity? Annu. Rev. Ecol. Evol. Syst. 38, 179 (2007).
- S. M. Jackson, Glide angle in the genus petaurus and a review of gliding in mammals, Mammal Rev. 30, 9 (2000).
- A. Mori, and T. Hikida, Field observations on the social behavior of the flying lizard, Draco volans sumatranus, in Borneo, in Copeia (American Society of Ichthyologists and Herpetologists, Lawrence, KS, 1994), pp. 124–130.
- M. G. McCay, Aerodynamic stability and maneuverability of the gliding frog Polypedates dennysi, J. Expt. Biol. 204, 2817 (2001).
- J. J. Socha, Kinematics: Gliding flight in the paradise tree snake, Nature 418, 603 (2002).
- S. P. Yanoviak, R. Dudley, and M. Kaspari, Directed aerial descent in canopy ants, Nature 433, 624 (2005).
- S. P. Yanoviak, M. Kaspari, and R. Dudley, Gliding hexapods and the origins of insect aerial behaviour, Biol. Lett. 5, 510 (2009).
- S. P. Yanoviak, and R. Dudley, The role of visual cues in directed aerial descent of Cephalotes atratus workers (hymenoptera: Formicidae), J. Exp. Biol. 209, 1777 (2006).
- J. M. V. Rayner, Bounding and undulating flight in birds, J. Theor. Biol. 117, 47 (1985).
- P. Paoletti and L. Mahadevan, Intermittent locomotion as an optimal control strategy, Proc. Roy. Soc. A: Math. Phys. Engg. Sci. 470, 20130535 (2014).
- D. Gurdan, J. Stumpf, M. Achtelik, K. M. Doth, G. Hirzinger, and D. Rus, Energy-efficient autonomous four-rotor flying robot controlled at 1 khz, in Proceedings of the IEEE International Conference on Robotics and Automation (IEEE, Roma, Italy, 2007), pp. 361–366.
- S. Lupashin, A. Schöllig, M. Sherback, and R. D'Andrea, A simple learning strategy for high-speed quadrocopter multi-flips, in Proceedings of the IEEE International Conference on Robotics and Automation (ICRA'10) (IEEE, Anchorage, Alaska, 2010), pp. 1642–1648.
- D. Mellinger, M. Shomin, N. Michael, and V. Kumar, Cooperative grasping and transport using multiple quadrotors, in Distributed Autonomous Robotic Systems (Springer, Berlin, 2013), pp. 545–558.
- M. Müller, S. Lupashin, and R. D'Andrea, Quadrocopter ball juggling, in Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS'11) (IEEE, San Francisco, California, 2011), pp. 5113–5120.
- J. Thomas, M. Pope, G. Loianno, E. W. Hawkes, M. A. Estrada, H. Jiang, M. R. Cutkosky, and V. Kumar, Aggressive flight with quadrotors for perching on inclined surfaces, J. Mech. Robot. 8, 051007 (2016).
- M. F. Bin Abas, A. S. B. M. Rafie, H. B. Yusoff, and K. A. B. Ahmad, Flapping wing micro-aerial-vehicle: Kinematics, membranes, and flapping mechanisms of ornithopter and insect flight, Chin. J. Aeronaut. 29, 1159 (2016).
- G. Reddy, A. Celani, T. J. Sejnowski, and M. Vergassola, Learning to soar in turbulent environments, Proc. Natl. Acad. Sci. USA 113, E4877 (2016).
- G. Reddy, J. Wong-Ng, A. Celani, T. J. Sejnowski, and M. Vergassola, Glider soaring via reinforcement learning in the field, Nature 562, 236 (2018).
- D. P. Bertsekas, D. P. Bertsekas, D. P. Bertsekas, and D. P. Bertsekas, Dynamic Programming and Optimal Control (Athena Scientific, Belmont, MA, 1995), Vol. 1.
- L. P. Kaelbling, M. L. Littman, and A. W. Moore, Reinforcement learning: A survey, J. Artif. Intell. Res 4, 237 (1996).
- R. S. Sutton, and A. G. Barto, Reinforcement Learning: An introduction (MIT Press, Cambridge, MA, 1998), Vol. 1.
- A. Andersen, U. Pesavento, and Z. J. Wang, Analysis of transitions between fluttering, tumbling and steady descent of falling cards, J. Fluid Mech. 541, 91 (2005).
- A. Andersen, U. Pesavento, and Z. J. Wang, Unsteady aerodynamics of fluttering and tumbling plates, J. Fluid Mech. 541, 65 (2005).
- Z. J. Wang, J. M. Birch, and M. H. Dickinson, Unsteady forces and flows in low Reynolds number hovering flight: Two-dimensional computations vs. robotic wing experiments, J. Exp. Biol. 207, 449 (2004).
- P. Paoletti, and L. Mahadevan, Planar controlled gliding, tumbling, and descent, J. Fluid Mech. 689, 489 (2011).
- S. Colabrese, K. Gustavsson, A. Celani, and L. Biferale, Flow Navigation by Smart Microswimmers Via Reinforcement Learning, Phys. Rev. Lett. 118, 158004 (2017).
- M. Gazzola, B. Hejazialhosseini, and P. Koumoutsakos, Reinforcement learning and wavelet adapted vortex methods for simulations of self-propelled swimmers, SIAM J. Sci. Comput. 36, B622 (2014).
- M. Gazzola, A. A. Tchieu, D. Alexeev, A. de Brauer, and P. Koumoutsakos, Learning to school in the presence of hydrodynamic interactions, J. Fluid Mech. 789, 726 (2016).
- G. Novati, S. Verma, D. Alexeev, D. Rossinelli, W. M. van Rees, and P. Koumoutsakos, Synchronisation through learning for two self-propelled swimmers, Bioinspir. Biomimet. 12, 036001 (2017).
- V. Mnih, K. Kavukcuoglu, D. Silver, A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al., Human-level control through deep-reinforcement-learning, Nature 518, 529 (2015).
- A. Waldock, C. Greatwood, F. Salama, and T. Richardson, Learning to perform a perched landing on the ground using deep reinforcement learning, J. Intell. Robot. Syst. 92, 685 (2018).
- K. Hang, X. Lyu, H. Song, J. A. Stork, A. M. Dollar, D. Kragic, and F. Zhang, Perching and resting—A paradigm for UAV maneuvering with modularized landing gears, Sci. Robot. 4, eaau6637 (2019).
- A. Belmonte, H. Eisenberg, and E. Moses, From Flutter to Tumble: Inertial Drag and Froude Similarity in Falling Paper, Phys. Rev. Lett. 81, 345 (1998).
- L. Mahadevan, W. S. Ryu, and A. D. T. Samuel, Tumbling cards, Phys. Fluids 11, 1 (1999).
- R. Mittal, V. Seshadri, and H. S. Udaykumar, Flutter, tumble and vortex induced autorotation, Theor. Comput. Fluid Dyn. 17, 165 (2004).
- U. Pesavento, and Z. J. Wang, Falling Paper: Navier-Stokes Solutions, Model of Fluid Forces, and Center of Mass Elevation, Phys. Rev. Lett. 93, 144501 (2004).
- H. Lamb, Hydrodynamics (Cambridge University Press, Cambridge, 1932).
- S. P. Yanoviak, Y. Munk, M. Kaspari, and R. Dudley, Aerial manoeuvrability in wingless gliding ants (Cephalotes atratus), Proc. Roy. Soc. London B: Biol. Sci. 277, 2199 (2010).
- S. Levine, C. Finn, T. Darrell, and P. Abbeel, End-to-end training of deep visuomotor policies, J. Mach. Learn. Res. 17, 1334 (2016).
- D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot et al., Mastering the game of go with deep neural networks and tree search, Nature 529, 484 (2016).
- R. Bellman, On the theory of dynamic programming, Proc. Natl. Acad. Sci. USA 38, 716 (1952).
- T. Duriez, S. L. Brunton, and B. R. Noack, Machine Learning Control-Taming Nonlinear Dynamics and Turbulence (Springer, Berlin, 2017).
- S. Verma, G. Novati, and P. Koumoutsakos, Efficient collective swimming by harnessing vortices through deep reinforcement learning, Proc. Natl. Acad. Sci. USA 115, 5849 (2018).
- R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour, Policy gradient methods for reinforcement learning with function approximation, in Advances in Neural Information Processing Systems (MIT Press, Cambridge, MA, 2000), pp. 1057–1063.
- G. Novati, and P. Koumoutsakos, Remember and forget for experience replay, in Proceedings of the 36th International Conference on Machine Learning Vol. 97 (PMLR, Long Beach, California, 2019), pp. 4851–4860.
- T. Degris, M. White, and R. S. Sutton, Off-policy actor-critic, in Proceedings of the 29th International Conference on Machine Learning (Omnipress, Edinburgh, Scotland, 2012), pp. 179–186.
- R. Munos, T. Stepleton, A. Harutyunyan, and M. Bellemare, Safe and efficient off-policy reinforcement learning, in Advances in Neural Information Processing Systems (MIT Press, Cambridge, MA, 2016), pp. 1054–1062.
- A. Y. Ng, D. Harada, and S. Russell, Policy invariance under reward transformations: Theory and application to reward shaping, in Proceedings of the 16th International Conference on Machine Learning (ICML'99) (Morgan Kaufmann Publishers Inc., San Francisco, California, 1999), Vol. 99, pp. 278–287.
- M. J. Lighthill, Introduction to the scaling of aerial locomotion, Scale effects in animal locomotion (Cambridge University Press, Cambridge, UK, 1977), pp. 365–404.
- J. M. V. Rayner, The intermittent flight of birds, Scale effects in animal locomotion (Academic Press, New York, 1977), pp. 437–443.
- M. Zhang, X. Geng, J. Bruce, K. Caluwaerts, M. Vespignani, S. V. Sun, P. Abbeel, and S. Levine, Deep reinforcement learning for tensegrity robot locomotion, in Proceedings of the IEEE International Conference on Robotics and Automation (ICRA) (IEEE, Singapore, 2017), pp. 634–641.
- F. A. Gers, J. Schmidhuber, and F. Cummins, Learning to forget: Continual prediction with LSTM, Neural Comput. 12, 2451 (2000).
- I. Sutskever, Training Recurrent Neural Networks (University of Toronto, Toronto, Ontario, Canada, 2013).
- S. Gu, T. Lillicrap, I. Sutskever, and S. Levine, Continuous deep q-learning with model-based acceleration, in Proceedings of the 33rd International Conference on Machine Learning, Vol. 48 (PMLR, New York, 2016), Vol. 48, pp. 2829–2838.
- J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, Proximal policy optimization algorithms, arXiv:1707.06347.
- J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel, High-dimensional continuous control using generalized advantage estimation, 4th International Conference on Learning Representations, (ICLR) (2016), arXiv:1506.02438.