Skip to content
VibekollenBETAVibekollen
VideoAI Engineer

Scaling up Continual Learning — Ronak Malde, Trajectory

Ronak Malde from Trajectory presents self-distillation, a method for training AI models on long-horizon tasks with multiple steps.

The problem he solves is that when models are trained on long sequences using existing methods, they become uncertain and fill their responses with words like "but", "wait", and "maybe" — what he calls the "but wait problem". Self-distillation works by making the model its own teacher: you give one version of the model extra information ("hints") and train another version to match its decisions without these hints. Unlike other training methods that must choose between different benefits, Malde says self-distillation achieves all four desired properties: training on real data, ability to sample different paths, doing so efficiently without parallel runs, and providing reward for every individual token.

Open on YouTube →

Vibekollen prepared this summary with AI from the original publication. The content belongs to AI Engineer.

More to read