BreakRL#
Learn reinforcement learning through failure.
BreakRL is a bilingual, failure-first set of RL lecture notes. Each chapter explains a mechanism, then removes it so you can see the failure.
Start in three minutes#
Open the minimum demo and drag the coverage slider (a teaching plot; it does not train a model);
Open Failure Atlas #8 (healthy offline loss, collapsing returns);
Read the Offline RL chapter: text PDF and the saved experiment.
Reading the site needs no install. Pages render saved outputs and do not train. The catalog below is the rest of the book; start with the loop above, not Chapter 1.
How to read#
Read the chapter text for the problem, formulas, and mechanism;
Open the experiment notebook and watch the algorithm on a small task;
Compare the ablations: a finished run is not the same as a method that works.
When training does not converge, look up the symptom in the RL Failure Atlas (“symptom → mechanism → reproduction → fix”).
All chapters#
# |
Chapter |
Text |
Experiments |
Run |
|---|---|---|---|---|
1 |
Multi-Armed Bandits: Exploration vs Exploitation |
|||
2 |
Markov Decision Processes |
|||
3 |
Temporal-Difference Learning |
|||
4 |
DQN: Neural Value Learning |
|||
5 |
Policy Gradient / REINFORCE |
|||
6 |
Actor-Critic / A2C |
|||
7 |
PPO: Constrained Policy Updates |
|||
8 |
SAC: Maximum-Entropy Continuous Control |
|||
9 |
Offline RL: CQL and IQL |
|||
10 |
Model-Based RL: Dyna-Q |
|||
11 |
Decision Transformer: RL as Sequence Modeling |
|||
12 |
RLHF: From Preferences to Rewards |
|||
13 |
DPO: Preference Optimization without a Reward Model |
|||
14 |
GRPO and RLVR: Verifiable Rewards |
Chapters 12–14 use toy sequence tasks and two-digit addition. They show the mechanisms; they are not a substitute for production RLHF, DPO, or GRPO training.
The site displays saved notebook outputs and does not train models while you read. To run the experiments, use Colab in the table above or the local install in the README.