Minimum demo

Minimum demo#

The loss goes down. The return collapses.

When offline data does not cover an action, the greedy Q-learning target can overestimate it. The loss looks healthy while the policy learns an action the data never contained.

Low coverage: unseen actions are easier for Q values to overestimate.

Falling loss and actual return With low data coverage, the loss keeps falling while actual return collapses late in training.
Failure mode: the loss falls while actual return collapses late in training. Final return 0.08 · TD loss 0.11
TD loss (lower is better) Actual return (higher is better) Behavior-cloning baseline

This is a fixed teaching illustration: it does not train a model and does not report the exact numbers from one experiment. Same failure mode: this demo → Failure Atlas #8 → the Offline RL chapter.