SonarMapper
Elevator pitch
A PPO-based reinforcement learning agent for autonomous underwater navigation in a simulated environment under partial observability. Custom reward design, hyperparameter tuning, and training-behaviour analysis via reward curves and rollout visualisations.
The problem
Underwater navigation is a hard RL problem for three compounding reasons: visibility is limited (sonar returns, not clean visual data), positioning is noisy, and the reward signal is sparse (you learn from mapping progress, not from every step). It's a good testbed for perception-under-uncertainty methods that transfer to aerial and space contexts, where similar problems appear.
The approach
- Built the environment in Gymnasium with an observation space based on limited sonar-like returns
- Trained a PPO agent via Stable Baselines3, iterating on reward shape and observation encoding
- Analysed training behaviour through reward curves, entropy plots, and rollout visualisations
- Investigated the exploration-exploitation trade-off in a partial-observability setting
What I learned
- Reward shaping matters more than algorithm choice. PPO is fine; the win comes from teaching the agent what "good mapping" actually means. The reward function is the specification of the goal, and getting it right takes iteration
- Partial observability is not just "noisy" observations. It's structurally different. Framing the observation to give the agent a short memory of recent returns significantly improved sample efficiency