Skill issue: stop deploying vision language models, use them with Skills — Merve Noyan, Hugging Face
Merve Noyan wrote a book on vision language models and now wants developers to stop calling them directly.
Put one in front of a camera and you will never get real time; a small detector trained for the task runs at forty frames per second on a toaster and beats the VLM anyway. Her other complaint is licensing: people deploy a popular detector without noticing its copyleft license. So she built a toolkit that hands her favorite Apache 2.0 models to a coding agent, which she calls a clueless computer vision engineer, plus what she calls vibe training. Give it a dataset with no labels and it labels images with a nine billion parameter open VLM, passes the overlaid bounding boxes to two smaller VLM judges, merges their verdicts on minimum agreement rather than consensus, and trains RF-DETR. The whole run costs three or four dollars on Hugging Face jobs and inference providers. On road signs the trained detector lands a good mean average precision against ground truth, and on document parsing it generalizes, catching a signature the labeling model itself missed.
Texten är källans egen beskrivning av publiceringen. Innehållet tillhör AI Engineer.
Mer att läsa
Server-Side Code Execution Tools for AI Agents, Compared
OpenRouter för 9 tim sedan
v0.40.0
Ollama för 9 tim sedan
Google froze its open source bug bounty program due to a ‘significant rise’ in AI submissions
TechCrunch AI för 13 tim sedan
Can ‘super intelligence’ and a non-binding safety pact solve AI’s image problem?
TechCrunch AI för 13 tim sedan