Atari benchmark
The Atari benchmark is the set of vintage Atari 2600 video games — Pong, Breakout, Space Invaders, Montezuma's Revenge and dozens more — packaged through the Arcade Learning Environment as a common testbed for deep RL. The appeal is that a single agent must learn all of them from the same raw pixels and score signal, with no game-specific engineering. It became the field's shared yardstick after DQN used it to show, game by game, that one algorithm could reach human level across very different challenges.
Performance is usually reported as a human-normalised score: zero is random play, a hundred is a human tester, and the median or mean across the full suite of fifty-odd games summarises an agent. It is a rich benchmark because the games span reflex twitch-play to long-horizon exploration — and Montezuma's Revenge in particular, with its sparse rewards, exposed how badly basic DQN explores. Critics note that deterministic Atari can be gamed by memorising trajectories, which is why sticky actions and other stochasticity were later added.