Skip to content
Aryan.
← All work

06 · Reinforcement learning, autonomous navigation

SonarMapper

Status
Completed (October 2025)
Role
Solo project
Started
2025

Stable Baselines3 / Gymnasium / PPO / Python

Elevator pitch

A PPO-based reinforcement learning agent for autonomous underwater navigation in a simulated environment under partial observability. Custom reward design, hyperparameter tuning, and training-behaviour analysis via reward curves and rollout visualisations.

The problem

Underwater navigation is a hard RL problem for three compounding reasons: visibility is limited (sonar returns, not clean visual data), positioning is noisy, and the reward signal is sparse (you learn from mapping progress, not from every step). It's a good testbed for perception-under-uncertainty methods that transfer to aerial and space contexts, where similar problems appear.

The approach

  • Built the environment in Gymnasium with an observation space based on limited sonar-like returns
  • Trained a PPO agent via Stable Baselines3, iterating on reward shape and observation encoding
  • Analysed training behaviour through reward curves, entropy plots, and rollout visualisations
  • Investigated the exploration-exploitation trade-off in a partial-observability setting

What I learned

  • Reward shaping matters more than algorithm choice. PPO is fine; the win comes from teaching the agent what "good mapping" actually means. The reward function is the specification of the goal, and getting it right takes iteration
  • Partial observability is not just "noisy" observations. It's structurally different. Framing the observation to give the agent a short memory of recent returns significantly improved sample efficiency

Next

Ask My Business →