Hoppa till innehåll
VibekollenBETAVibekollen
VideoAI Engineer

Skill issue: stop deploying vision language models, use them with Skills — Merve Noyan, Hugging Face

Merve Noyan wrote a book on vision language models and now wants developers to stop calling them directly.

Put one in front of a camera and you will never get real time; a small detector trained for the task runs at forty frames per second on a toaster and beats the VLM anyway. Her other complaint is licensing: people deploy a popular detector without noticing its copyleft license. So she built a toolkit that hands her favorite Apache 2.0 models to a coding agent, which she calls a clueless computer vision engineer, plus what she calls vibe training. Give it a dataset with no labels and it labels images with a nine billion parameter open VLM, passes the overlaid bounding boxes to two smaller VLM judges, merges their verdicts on minimum agreement rather than consensus, and trains RF-DETR. The whole run costs three or four dollars on Hugging Face jobs and inference providers. On road signs the trained detector lands a good mean average precision against ground truth, and on document parsing it generalizes, catching a signature the labeling model itself missed.

Öppna på YouTube →

Texten är källans egen beskrivning av publiceringen. Innehållet tillhör AI Engineer.

Mer att läsa