Interactive demo of why training loss flattens: the same loss curve shown on normal axes where it hugs an irreducible floor, and on log-log axes where it becomes a straight line, plus a slider showing loss components draining at different speeds.

The same curve, two pairs of axes
Loss curve on two axis scales loss training compute irreducible floor
Looks finished: most visible progress happens early, then the curve hugs the floor.
Where the remaining loss lives
Start of training: everything still to learn. Total loss: 4.50
grammar, common words 1.60 everyday facts 0.80 rare knowledge 0.40 inherent uncertainty 1.70
Empty track = already learned. The gray block never drains.