Profile Picture

Marlon Müller

I'm a MCML PhD student in the Cyber-Physical Systems Group led by Matthias Althoff. My research focuses on developing robust and predictable machine learning systems. In 2024, I visited Murat Arcak's group at UC Berkeley. In 2023, I spent a semester at UQ. I earned my bachelor's and master's degrees in Computer Science from TUM.

Publications

Falsification-Driven Reinforcement Learning for Maritime Motion Planning

Marlon Müller*, Florian Finkeldei*, Hanna Krasowski*, Murat Arcak, Matthias Althoff

Ocean Engineering

Compliance with maritime traffic rules is essential for the safe operation of autonomous vessels, yet training reinforcement learning (RL) agents to adhere to them is challenging. The behavior of RL agents is shaped by the training scenarios they encounter, but creating scenarios that capture the complexity of maritime navigation is non-trivial, and real-world data alone is insufficient. To address this, we propose a falsification-driven RL approach that generates adversarial training scenarios in which the vessel under test violates maritime traffic rules, which are expressed as signal temporal logic specifications. Our experiments on open-sea navigation with two vessels demonstrate that the proposed approach provides more relevant training scenarios and achieves more consistent rule compliance.

Provably Safe Reinforcement Learning: Conceptual Analysis, Survey, and Benchmarking

Hanna Krasowski*, Jakob Thumm*, Marlon Müller, Lukas Schäfer, Xiao Wang, Matthias Althoff

Transactions on Machine Learning Research

Ensuring the safety of reinforcement learning (RL) algorithms is crucial to unlock their potential for many real-world tasks. However, vanilla RL and most safe RL approaches do not guarantee safety. In recent years, several methods have been proposed to provide hard safety guarantees for RL, which is essential for applications where unsafe actions could have disastrous consequences. Nevertheless, there is no comprehensive comparison of these provably safe RL methods. Therefore, we introduce a categorization of existing provably safe RL methods, present the conceptual foundations for both continuous and discrete action spaces, and empirically benchmark existing methods. We categorize the methods based on how they adapt the action: action replacement, action projection, and action masking. Our experiments on an inverted pendulum and a quadrotor stabilization task indicate that action replacement is the best-performing approach for these applications despite its comparatively simple realization. Furthermore, adding a reward penalty, every time the safety verification is engaged, improved training performance in our experiments. Finally, we provide practical guidance on selecting provably safe RL approaches depending on the safety specification, RL algorithm, and type of action space.

EM27 Retrieval Pipeline: Automated EM27/SUN 2 Data Processing

Moritz Oliveira Makowski, Frank Hase, Friedrich Klappenbach, Andreas Luther, Lena Feld, Marlon Müller, Vyas Giridharan, Catherine Fait, Jia Chen

Preprint

Teaching

Fundamentals of Artificial Intelligence
Lecture · Teaching Assistant
Formal Methods for Cyber-Physical Systems
Lecture · Teaching Assistant
Safe Reinforcement Learning for Modular Robots
Practical Course · Supervisor
Cyber-Physical Systems
Seminar · Supervisor