Daily Paper Cast · Jingwen Liang, Gengyu Wang

Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability

April 10, 2026·25 min·2 clips
A 2025 paper claimed SFT memorizes while RL generalizes, but new research questions if those experiments were flawed.
SFT is not the villain here. Echo and another speaker start with the familiar post-training story: supervised fine-tuning improves in-domain behavior but tends to memorize, while reinforcement learning gets the credit for generalization. The paper pushes on that split. It asks whether SFT's limits are fixed, or whether optimization schedule, dataset exposure, and model capability change the result. The 640-step setup keeps the comparison grounded. With total gradient steps fixed, the paper separates repeated exposure on a smaller dataset from one-pass coverage of a larger one, then adds a larger batch, larger data condition. Repetition wins the first round. Setting 2, with 2,500 examples, batch size 32, and 8 epochs, beats the one-pass large data setting by a clear margin. Data still matters, though. Setting 1 combines 20,000 examples, batch size 256, and 8 epochs, and beats both alternatives. The hosts put the point plainly: more data helps, but one pass through more examples does not replace the signal from repeated exposure. Then response length turns into the useful tell. Early training produces long, inefficient answers as the model copies long-thinking traces, which drags performance down and sometimes causes format errors. They call it dip-and-recovery. As training continues, answers get shorter, performance improves, and length starts acting like a readout for optimization progress. Overfitting gets tested too. Four aggressive schedules probe failure modes, including one with the highest learning rate and no learning rate decay. That one overfits, performance drops, and response length climbs again. The close is conditional: model, data, algorithm, and schedule interact, so SFT generalization cannot be judged from one neat comparison.

As heard by us

A careful walk-through that tests SFT claims against optimization, data, and model capability.

Daily Papercast treats the claim with a steady hand: does reasoning SFT really just memorize, or does the answer change with optimization, data, and model capability?

Read the full review in PlayNext →

Why you'd press play

For a paper that tests whether repetition beats scale, press play here.

Read the full recommendation in PlayNext →
Listen to the show on