Skip to content
VibekollenBETAVibekollen
VideoAI Engineer

Compression at the Edge — Chris Alexiuk, NVIDIA

A panel discussion on how to shrink large AI models without making them less capable.

The key insight is that model layers have vastly different importance — the first and last layers are critical while many middle layers contribute almost nothing. GLM 5.2 was reduced from 1.5 terabytes to 250 GB (86 percent smaller) without corresponding performance degradation. NVIDIA advocates for NVFP4, a compact numerical format that shares precision information across groups of 16 values. The panel emphasizes that real-world testing in actual use cases is more important than abstract measurements.

Open on YouTube →

Vibekollen prepared this summary with AI from the original publication. The content belongs to AI Engineer.

More to read