Export citation

Export citation

Choose format for download:

Download Citation
  • Access by Xinjiang University

Deep reinforcement learning of airfoil pitch control in a highly disturbed environment using partial observations

Diederik Beckers*

Jeff D. Eldredge

  • *Contact author: beckers@caltech.edu

Phys. Rev. Fluids 9, 093902 – Published 12 September, 2024

DOI: https://doi.org/10.1103/PhysRevFluids.9.093902

Abstract

This study explores the application of deep reinforcement learning (RL) to design an airfoil pitch controller capable of minimizing lift variations in randomly disturbed flows. The controller, treated as an agent in a partially observable Markov decision process, receives non-Markovian observations from the environment, simulating practical constraints where flow information is limited to force and pressure sensors. Deep RL, particularly the TD3 algorithm, is used to approximate an optimal control policy under such conditions. Testing is conducted for a flat plate airfoil in two environments: a classical unsteady environment with vertical acceleration disturbances (i.e., a Wagner setup) and a viscous flow model with pulsed point force disturbances. In both cases, augmenting observations of the lift, pitch angle, and angular velocity with extra wake information (e.g., from pressure sensors) and retaining memory of past observations enhances RL control performance. Results demonstrate the capability of RL control to match or exceed standard linear controllers in minimizing lift variations. Special attention is given to the choice of training data and the generalization to unseen disturbances.

Physics Subject Headings (PhySH)

Article Text

References (43)

  1. X. An, D. R. Williams, J. Eldredge, and T. Colonius, Modeling dynamic lift response to actuation, in 54th AIAA Aerospace Sciences Meeting (AIAA, Reston, VA, 2016).
  2. A. R. Jones, Gust encounters of rigid wings: Taming the parameter space, Phys. Rev. Fluids 5, 110513 (2020).
  3. S. Skogestad and I. Postlethwaite, Multivariable Feedback Control: Analysis and Design (John Wiley, Hoboken, US-NJ, 2005).
  4. S. Ahuja and C. W. Rowley, Feedback control of unstable steady states of flow past a flat plate using reduced-order estimators, J. Fluid Mech. 645, 447 (2010).
  5. G. Sedky, A. Gementzopoulos, I. Andreu-Angulo, F. D. Lagor, and A. R. Jones, Physics of gust response mitigation in open-loop pitching manoeuvres, J. Fluid Mech. 944, A38 (2022).
  6. S. L. Brunton and C. W. Rowley, Empirical state-space representations for Theodorsen's lift model, J. Fluids Struct. 38, 174 (2013).
  7. S. L. Brunton, S. T. M. Dawson, and C. W. Rowley, State-space model identification and feedback control of unsteady aerodynamic forces, J. Fluids Struct. 50, 253 (2014).
  8. G. Sedky, F. D. Lagor, and A. R. Jones, Unsteady aerodynamics of lift regulation during a transverse gust encounter, Phys. Rev. Fluids 5, 074701 (2020).
  9. W. Kerstens, J. Pfeiffer, D. R. Williams, R. King, and T. Colonius, Closed-loop control of lift for longitudinal gust suppression at low Reynolds numbers, AIAA J. 49, 1721 (2011).
  10. K. Bieker, S. Peitz, S. L. Brunton, J. N. Kutz, and M. Dellnitz, Deep model predictive flow control with limited sensor data and online learning, Theor. Comput. Fluid Dyn. 34, 577 (2020).
  11. X. Xu, A. Gementzopoulos, G. Sedky, A. R. Jones, and F. D. Lagor, Design of optimal wing maneuvers in a transverse gust encounter through iterated simulation or experiment, Theor. Comput. Fluid Dyn. 37, 465 (2023).
  12. K. Arulkumaran, M. P. Deisenroth, M. Brundage, and A. A. Bharath, Deep reinforcement learning: A brief survey, IEEE Signal Process. Mag. 34, 26 (2017).
  13. R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction (MIT Press, Cambridge, MA, 2018).
  14. H. Kurniawati, Partially observable Markov decision processes and robotics, Annu. Rev. Control Robot. Auton. Syst. 5, 253 (2022).
  15. S. Verma, G. Novati, and P. Koumoutsakos, Efficient collective swimming by harnessing vortices through deep reinforcement learning, Proc. Natl. Acad. Sci. USA 115, 5849 (2018).
  16. G. Novati, L. Mahadevan, and P. Koumoutsakos, Controlled gliding and perching through deep-reinforcement-learning, Phys. Rev. Fluids 4, 093902 (2019).
  17. P. Gunnarson, I. Mandralis, G. Novati, P. Koumoutsakos, and J. O. Dabiri, Learning efficient navigation in vortical flow fields, Nat. Commun. 12, 7143 (2021).
  18. D. Fan, L. Yang, Z. Wang, M. S. Triantafyllou, and G. E. Karniadakis, Reinforcement learning for bluff body active flow control in experiments and simulations, Proc. Natl. Acad. Sci. USA 117, 26091 (2020).
  19. N. J. Nair and A. Goza, Bio-inspired variable-stiffness flaps for hybrid flow control, tuned via reinforcement learning, J. Fluid Mech. 956, R4 (2023).
  20. C. Xia, J. Zhang, E. C. Kerrigan, and G. Rigas, Active flow control for bluff body drag reduction using reinforcement learning with partial measurements, J. Fluid Mech. 981, A17 (2024).
  21. G. Novati, H. L. de Laroussilhe, and P. Koumoutsakos, Automating turbulence modelling by multi-agent reinforcement learning, Nat. Machine Intell. 3, 87 (2021).
  22. H. J. Bae and P. Koumoutsakos, Scientific multi-agent reinforcement learning for wall-models of turbulent flows, Nat. Commun. 13, 1443 (2022).
  23. P. I. Renn and M. Gharib, Machine learning for flow-informed aerodynamic control in turbulent wind conditions, Commun. Eng. 1, 45 (2022).
  24. L. P. Kaelbling, M. L. Littman, and A. R. Cassandra, Planning and acting in partially observable stochastic domains, Artif. Intell. 101, 99 (1998).
  25. L. P. Kaelbling, M. L. Littman, and A. W. Moore, Reinforcement learning: A survey, J. Artif. Intell. Res. 4, 237 (1996).
  26. L. Lin and T. M. Mitchell, Memory Approaches to Reinforcement Learning in Non-Markovian Domains (Carnegie Mellon University, Pittsburgh, PA, 1992).
  27. V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and A. Riedmiller, Playing Atari with deep reinforcement learning, arXiv:1312.5602.
  28. M. Hausknecht and P. Stone, Deep recurrent Q-learning for partially observable MDPs, in 2015 AAAI Fall Symposium Series (AAAI Press, Washington, DC, 2015).
  29. L. Meng, R. Gorbet, and D. Kulić, Memory-based deep reinforcement learning for POMDPs, in 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (IEEE, New York, 2021), pp. 5619–5626.
  30. S. Fujimoto, H. van Hoof, and D. Meger, Addressing function approximation error in actor-critic methods, in Proceedings of the 35th International Conference on Machine Learning, edited by J. Dy and A. Krause, Proceedings of Machine Learning Research (PMLR, 2018), Vol. 80, pp. 1587–1596.
  31. D. Beckers, Fast models and reinforcement learning control of unsteady aerodynamics, Ph.D. thesis, University of California, Los Angeles, 2023.
  32. T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor, in Proceedings of the 35th International Conference on Machine Learning, edited by J. Dy and A. Krause, Proceedings of Machine Learning Research (PMLR, 2018), Vol. 80, pp. 1861–1870.
  33. H. Wagner, Über die Entstehung des dynamischen Auftriebes von Tragflügeln, Z. Angew. Math. Mech. 5, 17 (1925).
  34. T. Theodorsen, General theory of aerodynamic instability and the mechanism of flutter, Tech. Rep. NACA-TR-496 (NASA, 1949), https://ntrs.nasa.gov/citations/19930090935.
  35. The MathWorks Inc., Control system toolbox version: 23.2 (r2023b) (2023).
  36. J. D. Eldredge, A method of immersed layers on Cartesian grids, with application to incompressible flows, J. Comput. Phys. 448, 110716 (2022).
  37. https://github.com/diederikb/AeroGym.
  38. https://github.com/diederikb/aero_gym_SB3.
  39. A. Raffin, A. Hill, A. Gleave, A. Kanervisto, M. Ernestus, and N. Dormann, Stable-Baselines3: Reliable reinforcement learning implementations, J. Mach. Learn. Res. 22, 1 (2021).
  40. T. von Kármán and W. R. Sears, Airfoil theory for non-uniform motion, J. Aeronaut. Sci. 5, 379 (1938).
  41. S. Neumark, Pressure distribution on an airfoil in nonuniform motion, J. Aeronaut. Sci. 19, 214 (1952).
  42. R. T. Jones, Operational treatment of the nonuniform-lift theory in airplane dynamics, Tech. Rep. NACA-TN-667, NACA, 1938 (NASA, 1938), https://ntrs.nasa.gov/citations/19930081472.
  43. I. E. Garrick, On some reciprocal relations in the theory of nonstationary flows, Tech. Rep. NACA-TN-629, NACA, 1938 (unpublished).

Outline

Information

Sign In to Your Journals Account

Filter

Filter

Article Lookup

Enter a citation