startrek
Deep Q-Network written from scratch that learns to land in Gymnasium's LunarLander, with YAML configs, baselines and multi-seed experiments. Reaches an average return of about 240.
Team
2 people
Context
Epitech project, 2nd year
Status
Done
Context
LunarLander counts as solved at an average return of 200. The goal was to get there with a DQN written from scratch and to prove the result holds across random seeds.
What was built
- DQN implementation written from scratch in PyTorch
- YAML configs to run experiments reproducibly
- Baselines and multi-seed runs to compare results fairly
- Learning curves plotted with Matplotlib
Results
239.6mean eval return, tuned DQN over 5 seeds (solved at 200)
177.3same DQN before tuning
5seeds per experiment, 100 eval episodes each
