Minimum demo#
The loss goes down. The return collapses.
When offline data does not cover an action, the greedy Q-learning target can overestimate it. The loss looks healthy while the policy learns an action the data never contained.
Low coverage: unseen actions are easier for Q values to overestimate.
Failure mode: the loss falls while actual return collapses late in training.
Final return 0.08 · TD loss 0.11
TD loss (lower is better)
Actual return (higher is better)
Behavior-cloning baseline
This is a fixed teaching illustration: it does not train a model and does not report the exact numbers from one experiment. Same failure mode: this demo → Failure Atlas #8 → the Offline RL chapter.
Failure Atlas #8 Read the Offline RL chapter Open the chapter PDF