LFM2.5-Encoders for Fast Long-Context Inference on CPU
Article image or reusable cover for Hugging Face
Liquid AI has launched LFM2.5-Encoders, small language models with 230 million and 350 million parameters optimized to run quickly on standard computer processors.
They can handle 8,192 tokens at once and are approximately 3.7 times faster than competing ModernBERT on long texts — an 8,192-token text takes about 28 seconds instead of 90 seconds. The models are trained to understand text bidirectionally and can be used for classification, routing, detecting personal information, and other tasks that run continuously. You can download them freely and fine-tune them for your own needs.
At 8,192 tokens, ModernBERT-base takes over a minute and a half per forward pass versus about 28s for LFM2.5-Encoder-230M.
Read the full story at Hugging Face →
Vibekollen prepared this summary with AI from the original publication. The content belongs to Hugging Face.