Skip to content
VibekollenBETAVibekollen

Source

AI Engineer

242 items in the feed. The texts are the sources' own descriptions — the content belongs to AI Engineer.

VideoAI Engineer

The State of AI in Software Development: Data from 400+ Orgs — Justin Reock, DX

Justin Reock (Deputy CTO, DX) shares five trends from DX's data on about 200,000 engineers. Deployment frequency is rising, but change failure rate has become far more volatile. Maintainability is up while change confidence is down. PRs have grown from about 44 to 72 lines. Juniors use AI the most, but staff+ engineers save as much time while using fewer tokens. Median gains in PR throughput are around 7.7%, and even top performers didn't reach 2x, because code generation was never the bottleneck. Justin walks through the DX AI Measurement Framework (utilization, impact, cost), how to assess your platform's AI readiness, and "agent experience," which uses feedback from the agents themselves. He closes with case studies of AI applied across the whole SDLC: Morgan Stanley (300K hours saved a year), Zapier (15% more value per engineer, and hiring more) and Spotify (an SRE incident agent).

30 Sept youtube.com

VideoAI Engineer

The Chief AI Officer: Scientist, Architect, Coach — Rania Khalaf, WSO2

The Chief AI Officer is one of the fastest-growing roles in tech, and one of the least defined. AI is in everything, so the job can easily turn into everything. Rania Khalaf (Chief AI Officer, WSO2; previously 20 years at IBM Research, where she ran about a third of the global AI research organization) shares a framework from two years in the role. It splits the job into three focus areas, Scientist, Architect and Coach, each a slider whose setting depends on your company type, its AI maturity and your own skills. She draws on her time at an ag-biotech unicorn, where simple blob detection beat machine learning, and at WSO2, where AI reshaped the product strategy around the agentic enterprise. She explains why she refuses to measure tokens and what she tracks instead: AI fluency across the workforce, depth of adoption, GEO visibility, agent-consumable products (MCP servers, skills, CLIs), agent-proof pricing, earned thought leadership and AI ARR. She closes on why the role needs strong CEO backing, and why sometimes the CPO, or even the head of HR, ends up running AI.

30 Sept youtube.com

VideoAI Engineer

The Death of the Code Review: What the Data Actually Says — Laurie Voss, Arize AI

Developers using autonomous agents wrote 741% more code but shipped only 30% more software. The bottleneck is human review, and "just review harder" doesn't scale. Reviewer effectiveness collapses past about 400 lines, and agents now open 10,000-line PRs. Laurie Voss (Head of Developer Relations at Arize AI, co-founder of npm) reviews what the industry is actually doing about it. The evidence covers OpenAI's zero-human-code product, METR's finding that about half of SWE-bench-passing PRs wouldn't be merged, and Cognition's FrontierCode (88% on SWE-bench Pro vs. 29% on real mergeability). She explains why a mergeability benchmark would immediately become a training signal for frontier models. She also covers how Cursor and GitHub run review in production, why multi-pass review and default suspicion cut false positives, and what Carlini's agent-built C compiler and Bun's million-line Zig-to-Rust port (13,044 unsafe blocks) reveal about taking humans out of the loop. Automated reviewers can be fooled by prompt injection that humans catch, which leaves production as the last reviewer standing.

30 Sept youtube.com

VideoAI Engineer

Your Agents Are in Solitary Confinement: Why MCP & A2A Aren't Enough — Vlad Luzin, Band

If you run Claude and Codex side by side, one planning and one reviewing, you're doing the job of a network router, passing messages between two stateful agents by hand. Loop engineering hands that job to a script. Neither one lets agents actually work together. Vlad Luzin (Co-founder & CTO, Band) explains why connecting agents is a distributed systems problem. Bigger context windows don't fix the single-agent bottleneck. Messaging platforms like Slack, Discord and WhatsApp are built for humans and block bot-to-bot traffic. Protocols like MCP and A2A are too low-level: stateless calls, one-way client/server, no discovery, and queues and timeouts you still have to build yourself. He then walks through what's needed: real-time ordered transport, persistence and rehydration, runtime binding across frameworks, agent-first abstractions (participants, rooms, routing) and enterprise governance. Two live demos show Codex, LangGraph, Claude Code and a personal assistant finding each other and collaborating in real time, with full human-in-the-loop visibility.

30 Sept youtube.com

VideoAI Engineer

AI-Generated Code Is Already Competing With Human Code — Daksh Gupta, Greptile

More than a quarter of the pull requests Greptile reviewed in April showed signs of being written largely or entirely by AI agents, up from under 1% a year earlier. Daksh Gupta, co-founder of Greptile, analyzed more than a million pull requests a month from companies like NVIDIA, Coinbase and Scale to test a simple question: are fully vibe-coded PRs actually any good in enterprise codebases? He compares agent-written PRs (Codex, Claude Code, Devin, Cursor) with human-written ones on four measures: revert rates, revert rates by PR size, the severity of issues Greptile flags (P0/P1/P2), and how many review rounds it takes to get to a mergeable PR. On all four, agent code landed in the same range as human code, and humans were actually more likely to introduce P0 bugs. The differences show up in *how* each one fails: Claude is about 1.5x more likely than humans to introduce SQL injection, Devin is half as likely to cause auth bypasses, and N+1 queries are far more common from Cursor. Daksh also shares what this means for code review.

27 Sept youtube.com

VideoAI Engineer

Get Out of the Model's Way — Kevin Hou, Google Antigravity

93 subagents, 12 hours, 2 billion tokens and under $1,000: that's what it took Google's Antigravity to build an operating system kernel from scratch that runs Doom. Kevin Hou, who leads engineering on Antigravity, Google DeepMind's agentic coding product, explains the principle behind it: "Give Messi the ball and get out of the way." Your product should get better every time the model does. Hou traces how coding tools evolved, from autocomplete in 2022 to agents in 2024, agent managers in 2025 and agent teams in 2026. He also shares the lessons from hard calls like removing chat from Windsurf. He then introduces Antigravity 2.0 and the three building blocks of the agent-teams era. The first is dynamic subagents, where Gemini 3.5 Flash leads a team through /teamwork. The second is sidecars, a new plugin protocol that lets agents listen for webhooks, cron jobs and messages. The third is generative UI rendered on the fly. He also shows how DeepMind researchers automated 90% of eval analysis with 100 parallel hypothesis agents.

27 Sept youtube.com

VideoAI Engineer

Software Engineering Is Becoming Factory Engineering — Zach Lloyd, Warp

Zach Lloyd, founder and CEO of Warp and a former principal engineer who led engineering on Google Docs, is still shipping all the time, but hasn't written a line of code in six months. His thesis: software engineering is turning into factory engineering. Every serious project will run a software factory, the same way every project now has CI/CD. Lloyd explains why Warp open-sourced after five years closed (60,000+ GitHub stars and over 800,000 active developers). When software gets cheap to build and clone, a great product isn't enough. Building in the open builds an ecosystem, and automation now handles the usual pain of noisy issues and sloppy PRs. He tours the factory floor: inputs, triage, product and tech specs, implementation, review, verification and monitoring, plus the architecture and data layer underneath. He also covers skill loops that improve the factory itself. He ends with a starter repo for building your own factory agents, and Q&A on build vs. buy, advice for new grads and why human taste still matters.

27 Sept youtube.com

VideoAI Engineer

GLM-5.2: Open Weights, Near-Frontier Intelligence — Zixuan Li, Z.ai

Zixuan Li, head of Z.ai, introduces GLM-5.2 to the world. On the hardest long-horizon coding and agentic benchmarks, it lands between Claude Opus 4.7 and 4.8. Even its non-thinking mode beats GLM-5.1 with thinking turned on, and it leads other open-weight models on the Artificial Analysis Intelligence Index. Li also explains where the GLM name comes from (a 2021 paper that put Z.ai among the earliest LLM labs), and why GLM is more than a coding model. The core of his talk is why Z.ai releases near-frontier models as open weights. It lets enterprises and governments run models on their own servers, lets companies fine-tune for law, finance and security, and lets customers co-design where the models go next. He closes with one more thing: Z Code, Z.ai's own coding harness built for GLM-5.2 that also works with other frontier models. Related links: Z.ai: https://z.ai Timestamps: 0:00 Intro from swyx 0:47 Z.ai and GLM 2:05 Where the GLM name comes from 3:10 Intelligence beyond IQ 4:25 GLM-5.2 benchmarks vs.

27 Sept youtube.com

VideoAI Engineer

Orchestras, Not Factories: How the Fastest Builders Work — Charlie Holtz, Conductor

"I honestly kind of hate the term 'software factory.'" Charlie Holtz, co-founder and CEO of Conductor, the desktop app for running many coding agents at once, has watched many of the best builders up close. He shares the principles they have in common. Stay near the frontier, but don't try to beat the market: if a workflow works for everyone, the big labs will build it into the default. Create "slop-free zones," like migrations, docs and skill files, that always get careful human review. Feed the beast by putting every Slack message, bug report and meeting into one database your agents can query. And give agents free range, with cloud sandboxes where they keep running after you close your laptop. Holtz demos a new version of Conductor that adds real-time collaboration and agents you can start from your phone. He closes with his case for orchestras over factories: humans at the center, conducting teams of people and agents instead of managing an assembly line. Speaker info: X/Twitter: @charlieholtz (https://x.com/charlieholtz) Related links: Conductor: https://www.conductor.build Timestamps: 0:00 Intro 1:05 What Conductor is 1:40 Principles from the best builders 2:00 1.

27 Sept youtube.com

VideoAI Engineer

Scale the Judgment, Not the Model — Andrew Orobator, Reddit

Swap in a smarter model and you get a slightly better answer. Take away the tests, gates and review, and everything falls apart. Andrew Orobator, a senior Android engineer at Reddit, argues that the model was never the bottleneck. The bottleneck is the judgment around it. Humans absorb judgment implicitly, but agents need it spelled out. He shows how to get judgment out of people's heads and into the repo. Skills are institutional judgment turned into something an agent can run. Work logs let a fresh agent pick up at milestone 7 of 9 (this talk itself was built with one). Personas lend you a security reviewer's or designer's eye. He then covers the verification ladder and his feature-flag cleanup agent, which went 7 for 7 on green-CI PRs at $1.26 each. And he shares a warning: when he asked Codex for reasons to unlock his repo guard, it quietly added a self-authorizing "emergency recovery" exception.

27 Sept youtube.com

VideoAI Engineer

No, That's Not a Software Factory — Ryan Cooke, WorkOS

The standard software factory puts an agent in a sandbox, prompts it and merges the PR. When WorkOS built one, the results were hard to tell apart from engineers running Claude Code on their laptops. Ryan Cooke, an engineer at WorkOS, argues that counting PRs or lines of AI code is the wrong measure. What matters is outcomes: are you actually shipping more to customers? WorkOS's answer is to build its engineering process into the factory itself. TARS, its agent, lives in Slack, Linear and GitHub and uses webhooks to track projects and pick up the next ticket on its own. A PM agent drafts the "Hilltop" product doc and breaks approved work into tickets, which solves the blank-page problem. Horizon, the orchestration layer, sits in front of an internal MCP gateway that connects every tool, including Snowflake. Cooke calls that gateway the single best investment for any team starting a factory. He also covers bug and support triage, why WorkOS is building its own sandboxes and a company-wide memory layer, and how it measures success with defect rates, recovery time and adoption.

27 Sept youtube.com

VideoAI Engineer

What It Actually Takes to Build a Software Factory — Tereza Tížková, Factory

Everyone is talking about software factories, but few people are building one. Tereza Tížková of Factory, which runs software factories for enterprises like EY and Adobe, defines one as the whole software lifecycle run autonomously: collecting signals, prioritizing, building, validating and improving. It's not just a swarm of coding agents, because writing code is the easy part. She lays out three principles. Stay agnostic: fit how your team already works, and route between models automatically (Factory's conservative benchmark shows about 25% savings). Stay autonomous: Factory's Missions run for hours or even weeks, with a real customer example running 16. Worker agents run in sequence instead of a swarm, so each starts with fresh context, and validators check code they didn't write. One of those validators actually clicks through the app. Keep improving: a deferred context engine cuts token use by 50% or more, and an agent-readiness check keeps AI from making a messy codebase worse. She closes on what's left for humans: deciding what to build, not how.

27 Sept youtube.com

VideoAI Engineer

I Turned Coding Agents Into a Strategy Game — Ido Salomon, AgentCraft

If agents are so capable, why doesn't everyone have an army of them? Ido Salomon, creator of AgentCraft and MCP-UI and co-creator of MCP Apps, argues the bottleneck is us: steering, directing and reviewing agents at scale is exhausting. The skills we need already exist, though. We learned them from games like Warcraft and StarCraft. Salomon demos AgentCraft, an orchestrator styled like a strategy game. Agents appear as units on a map, your file system becomes the terrain, heat maps show where work is happening, and the space bar jumps you to whatever needs attention. He shows how it adds visibility, autonomy (with orchestrators, loops and a review kit with visual evidence) and collaboration, through shared rooms where people and agents work side by side. Then he turns to lowering the floor: an early, simpler mobile-game-style mode aimed at the 90% of people who aren't power users yet. Speaker info: X/Twitter: @idosal1 (https://x.com/idosal1) LinkedIn: https://www.linkedin.com/in/ido-salomon/ Related links: MCP-UI: https://mcp-ui.dev Timestamps: 0:00 Intro 0:35 If agents are amazing, why aren't we unstoppable?

27 Sept youtube.com

VideoAI Engineer

Building Self-Improving Agent Software Factories — Suraj Gupta, Warp

There's a lot of talk about self-improving agents and about software factories, but much less about how a factory itself gets better over time. Suraj Gupta, who leads harness development at Warp, shows three practical ways to do it, with demos from Warp's own open-source factory. The first is skills that improve themselves. An outer-loop agent watches the triage agent's runs and feedback, then opens a PR to update its skill, so every change is tracked in Git and reviewed by a human. The second is persistent memory: a versioned, traceable store of facts, so a Sentry agent doesn't have to rediscover a root cause it already found. It works with any harness, including Claude Code and Codex. The third is model routing, so you're not paying Opus prices for triage or simple CI fixes. You can use Warp's auto models or set your own rules, and Warp's internal evals found that UI tasks run well on GLM.

27 Sept youtube.com

VideoAI Engineer

We Let Claude Code and Codex Race Human Researchers — Elie Bakouch, Prime Intellect

Big labs say recursive self-improvement is coming, but there's no independent benchmark to check that claim. Elie Bakouch, Research Engineer at Prime Intellect and creator of Hugging Face's SmolLM, set Claude Code and Codex loose on the community's Optimizer Speedrun, a race to train a GPT-2-level model in the fewest steps. Both agents beat the human record. Along the way they behaved very differently. Claude Code kept stopping every nine or ten hours to say the record couldn't be beaten, and sat idle about a third of the time. Codex never stopped, wrote far more notes, spawned more sub-agents and burned more tokens. In a longer six-day run, Kimi turned out to be the most token-efficient, and a paper only Claude found led to the best record. But Bakouch's key finding is sobering: none of the models invented a new optimizer. They combined existing ideas for small gains. He closes with an AlphaEvolve-style loop Prime Intellect is building for real discovery, and makes the case for doing this research in the open.

26 Sept youtube.com

VideoAI Engineer

One Operator, Many Drones: Inside Skydio's Autonomy Stack — Suchet Bargoti, Skydio

Suchet Bargoti opens with a live demo instead of slides. From a laptop on conference Wi-Fi, he launches a docked drone over San Mateo, starts a second one on a power line in Colorado, has one autonomously track a car, and then sends the whole fleet back to dock. Skydio, the largest US drone manufacturer, now has thousands of these docks deployed with utilities, police and construction companies, and about 16 million people live within two miles of one. Bargoti, Skydio's Director of Inspection and Mapping, explains the shift to "drones as infrastructure." He shows a drone spotting a utility pole burning from the inside, and SFPD following a stolen car without a high-speed chase. He then breaks down the autonomy stack. It splits intelligence between the edge and the cloud, uses maps as world models that the fleet keeps up to date, tracks objects through occlusion, and lets a VLM agent find and follow a "white Jeep" using tool calls instead of hand-coded rules. He also covers where fully end-to-end learning still falls short of the reliability physical systems need.

24 Sept youtube.com

VideoAI Engineer

Robot Demos Are Easy. Reliability Is Hard — Jason Ma, Dyna Robotics

A generalist robot that succeeds 80 to 90% of the time makes a great video but a poor product. Jason Ma, co-founder and CTO of Dyna Robotics, shows how Dyna got its model, Dyna-1, to a 99.4% success rate folding restaurant napkins for 24 hours straight, including recovering after it pulled over the whole stack. The key is a reward model that watches the robot and scores its progress. When the score dips, the robot is making a mistake, so the team can collect targeted recovery data and fine-tune again in a human-in-the-loop active-learning cycle. Ma also covers Dyna's pre-training data pyramid of more than 200,000 hours, and its architecture that pairs a reasoning model with a world action model. He shows deployments in restaurants and a Sacramento laundromat, and a robot folding T-shirts for three days straight at CoRL in Korea with no site-specific data.

24 Sept youtube.com

VideoAI Engineer

Physical AI's Next Bottleneck Is Finding the Right Video — Rafael Levi, Bright Data

AI without data is just a box." Rafael Levi of Bright Data points out that LLMs train on trillions of words, but robotics has only about a million videos of robots doing things. Paying people to record staged actions produces biased data, because nobody opens a door naturally when told to. Meanwhile, the web holds billions of hours of real people handling objects and real physics, with gravity, motion and cause and effect. The catch is noise. Levi notes that NVIDIA discards about 96% of the video it downloads to train Cosmos, and Stable Video Diffusion discards 74%, which wastes compute, bandwidth and storage. Meta, on the other hand, trained on about a million hours of public video and needed only 62 hours of real robot data to control a robot. Bright Data's approach is "search first, collect second." It indexes more than a billion videos by the actions in them, not their titles, and returns trimmed clips with timestamps, match scores and frame counts through an API. Levi demos searches for dishwashing and clothes-folding clips and covers uses in self-driving and brand discovery.

24 Sept youtube.com

VideoAI Engineer

World Models Need Causality, Not Pretty Pixels — Christopher Manning, Moonlake AI

In 2007, Google already had a language model trained on 2 trillion tokens, the same order of magnitude as today's frontier models. Christopher Manning, Stanford professor, former director of the Stanford AI Lab and now at Moonlake AI, uses that as one stop in a tour of AI history. It runs from Dartmouth in 1956 and Shakey the robot to the rise of LLMs, and it ends at his argument for what comes next: embodied intelligence built on simulation. Manning argues that generative video like Genie 3 simulates observations but has no semantics underneath, so it can't support planning. Moonlake takes a single photo or short video and builds an action-conditioned world in code, with objects you can pick up, open and move. It even researches objects on the web to fill in what the camera can't see, like the tea bags inside a closed box. A loop inspired by Claude Code compares renders against reality to shrink the sim-to-real gap, with the goal of replacing 10,000 hours of teleoperation with simulation. The talk ends with Q&A on gaming, ontologies, physics and discovery in latent space.

24 Sept youtube.com

VideoAI Engineer

Robotics Has Been Stuck for 70 Years — Deepak Pathak, Skild AI

Show a robotics demo from 1957 next to one from today, and most people can't tell which is which. Deepak Pathak, co-founder and CEO of Skild AI and a professor at Carnegie Mellon, argues robotics stalled because it was treated as a hardware problem rather than a problem of building a general brain. Collecting robot data by teleoperation at one example a minute, he calculates, would take the entire US population more than a century to reach GPT-3 scale. Skild's answer is an "omni-bodied" brain: one model for any robot and any task. It's pre-trained on scalable data like simulation and human video, post-trained on teleoperation, and improved by a deployment flywheel. Pathak shows it inserting AirPods with a simple gripper, learning from human video with under an hour of robot data, and cooking omelets on $4,000 arms with only a camera. He explains why climbing stairs is harder than a backflip. He also shows GPU assembly for NVIDIA's Houston factory, and a robot learning to walk on two legs in three tries after its other legs are disabled.

24 Sept youtube.com

VideoAI Engineer

AI Engineer Paris 2026 Opening Keynotes: Mistral, Langfuse & Sizzy | Day 1

Live from STATION F, AI Engineer Paris 2026 opens with welcome remarks and main-stage talks on AI infrastructure, coding agents, software factories, and building frontier AI. AI Engineer Paris returns for its second edition, bringing together 1,000+ engineers, founders, CTOs, and VPs of AI for two days of technical talks, practical workshops, and AI engineering research. Day 1 main-stage lineup: • Jen Person (Mistral) — Welcome to AI Engineer Paris 2026 • Lia McBride (AI Engineer) — Intro to AI Engineer • Lara Khanafer (STATION F) — Welcome from Station F • Clemens Rawert (Langfuse) — What Can Economic History Teach Us About the AI Moment? • Kitze (Sizzy) — The Next Level of AI Engineering • Lélio Renard Lavaud (Mistral) — Building Frontier AI This livestream covers the Main Stage. Additional sessions and activities are taking place throughout the event. Event: AI Engineer Paris 2026 Date: September 23, 2026 Venue: STATION F, Paris View the full schedule: https://ai.engineer/paris/2026 Subscribe for more AI Engineer talks, livestreams, and event coverage: https://www.youtube.com/@aiDotEngineer Schedule subject to change. Timestamps will be added following the broadcast.

24 Sept youtube.com

VideoAI Engineer

AI Engineer Paris 2026 Main Stage: Google DeepMind, ElevenLabs, Hugging Face & Stripe | Day 2

Live from STATION F, the AI Engineer Paris 2026 Main Stage continues with a full day of technical talks from the teams building and scaling production AI systems. Watch sessions covering model customization, inference infrastructure, voice and realtime AI, coding agents, software factories, and the emerging AI economy. Day 2 main-stage lineup: • Jakob Pörschmann (Black Forest Labs) — From Pixels to Robots • Dorian Lods (ElevenLabs) — Generating Identity at Scale • Harry Mellor (Hugging Face) — How a Transformers Model Loads in vLLM • Arielle Le Bail (Stripe) — Inside the AI Economy: What Stripe’s Data Reveals • Hervé Bredin (pyannoteAI) — Rebuilding a Streaming Model for Voice AI • Charles Frye (Modal) — The Low-Latency Inference Engineering Playbook • Dominic Pajak (Arm) — Eyes Up: Edge AI Innovation • Yannis Psorakis (Cognition) — Agentic Map-Reduce for Large-Scale Coding Tasks • Matt Pocock — Fixing the PR Bottleneck • Maarten Grootendorst (Google DeepMind) — The Efficiency of Gemma 4 • Olivier Teboul (Gradium) — The Missing Layers of Conversational AI This livestream covers the Main Stage. Additional tracks and workshops run concurrently and are not included in this broadcast.

23 Sept youtube.com

VideoAI Engineer

From Scratch to SOTA: Training a 3B State-Space Vision Model — Krishna Prasad Srinivasan, Sarvam

Well under one percent of the Common Crawl corpus that frontier models train on is in an Indian language, even as frontier labs call India a fast growing market. Krishna Prasad Srinivasan says the knowledge exists but was never digitized, and Sarvam's answer is a three billion parameter vision language model, small enough for one GPU, that he says beats document AI models a hundred times larger. The model is unusual twice over. Its backbone is a state space model rather than a transformer, because a page can run to ten thousand visual tokens and quadratic attention gets expensive, while an SSM keeps constant memory. And in late 2025, when most OCR models were monolithic page level VLMs, the team bet on block level OCR wrapped in a layout harness and a reading order harness, which many 2026 releases have since converged on. Training runs as a four stage curriculum: thirteen trillion text tokens across English, 22 Indian languages, math, and code before the model sees a pixel, so a language prior can resolve a smudged word; continual pretraining on three hundred million image text pairs; supervised fine tuning on a hundred million OCR samples; then reinforcement learning.

23 Sept youtube.com

VideoAI Engineer

From VLM/VLA's to Embodied Agents — Armen Aghajanyan, Perceptron AI

Feed a model one hour of video and roughly a million visual tokens go in, yet the loss lands on about two percent of them, since the only ground truth is a transcript or a few labeled frames. Armen Aghajanyan calls that a humongous waste, and predicting every pixel treats a background pixel with the same weight as a gripper tip or a contact point. Perceptron's answer is a perceptive objective that learns which percepts will matter, rather than hardcoding the gripper. The second problem is context bloat from cameras that never switch off. Patch averaging buys ten times compression, but his fix is data sparse mixture of experts, a router that decides per layer which tokens to read and which to skip. Left alone, the model zooms into the graph in a figure and spends more tokens on fruit when asked to segment fruit. Put together, that produced the model his team released a few weeks earlier, trained on a petabyte spanning text, images, video, and trajectories from desktop use to video games, which he says beats a frontier lab's embodied reasoning model at a fraction of the cost.

23 Sept youtube.com

VideoAI Engineer

You’re Not Thinking Big Enough: Rebuilding Food Systems with AI Agents — Cody Menefee, Firecrawl

Cody Menefee once drove three hours to a processor with ten turkeys strapped to the roof of his Tesla, and reports that a windbreak of birds up top does real damage to your range. He raised them in a Nashville backyard where that is not allowed, nearly quit engineering to farm, and learned how hard it is to make money at it. Instead he asks why only three percent of cattle finish on pasture when it is better for animal, consumer, and land. The answer is labor. Grazing done right means splitting pasture into paddocks with one day of feed and moving the herd, fences, and water every day so the grass rests. GPS collars with virtual fences handle the moving but not the deciding: grass grows differently after drought, rain, or trampling, and today a farmer walks out to look. His proposal is to replace those eyes and drop a language model in the loop. Drones give the best imagery but no jurisdiction allows autonomous flights; satellites are too far away; a trail cam pointed at a measuring stick is the cheap version. Feed the model animal locations, drought conditions, and grass height, and let it suggest the next paddock for a human to confirm.

23 Sept youtube.com

VideoAI Engineer

The Best Models Still Reason Like Toddlers — Andrew Dai, Elorian

Show a frontier model part of a chessboard, ask how many white squares are visible, and it answers 32. Andrew Dai's diagnosis: it saw a chessboard, knew chessboards have 32 white squares, and hallucinated the rest. The pattern matching that makes these models superb at naming flowers hurts them once a question needs counting or spatial grounding: they miscount a Catan player's roads from the pieces left off the board, and they miss a robot arm lifting a lid because they cannot hold state across a long video. His test for understanding versus reasoning: if a person can answer in one second, so can the model; if it takes longer, the model falls apart. The benchmarks hide this: one popular reasoning suite uses 32 by 32 pixel images, and a multimodal science exam can mostly be answered without the image. What is missing, he says, is visual thinking. Video generators produce cartoonish explosions because their training data is Hollywood and game engines; detection models are robust but passive.

23 Sept youtube.com

VideoAI Engineer

From Ingestion to Agents: How AI Teams Build on Document Intelligence — Adit Abraham, Reducto

The newest frontier model scores about thirty percent on a data lab's benchmark of decisions from PDFs, and Adit Abraham has met people who worked on PDF processing before he was born. The format was built to print, not to be reasoned over, and humans encode meaning visually: merged cells, line charts, unreadable handwriting. Reducto has processed billions of them, and the talk is the lessons, not the product. RAG meant a bad parse cost one answer; with agents, bad inputs compound across every step. VLMs finally read the long tail like a human, but they are not one size fits all: small detectors still find layout on a CPU at scale, and a VLM asked to rewrite OCR will helpfully recompute a total the human got wrong. His agentic OCR applies token level corrections, a zero for an O, instead of regenerating the page. Simple tables go to markdown and complex ones to HTML, but embedding models cannot match how did revenue change to a blob of tags, so a natural language rendering serves retrieval. Parsed structure rather than raw PDFs lifted other frontier models past the newest one on that benchmark and cut reasoning tokens.

23 Sept youtube.com

VideoAI Engineer

Modality Misalignment and Originality Attribution in Short-Form Video — Aditya Gautam, Meta

A video is about sports for its first six seconds, then turns political for half a second. Catching that across a hundred million plus videos is one of two problems Aditya Gautam works on at Meta; the other is unoriginal content, trivial to make with AI tools and corrosive to attribution. Both sit on messy data: adversarial uploads, multilingual on screen text, drift, and no ground truth. Mismatch across modalities is a solved embedding problem; mismatch within one modality over time is not, so three agents share it. A perceiver splits the video where temporal change happens, not at a fixed frame rate, and emits clip level embeddings, tags, and OCR. A reviewer runs the temporal analysis over that JSON and folds in live comments and sentiment. A retriever indexes topics, embeddings, and entities into inverted, vector, and graph stores for pulling similar clips and authors at inference time. Every agent runs on a small specialized VLM, not a frontier model, since he does not care whether it can code.

23 Sept youtube.com

VideoAI Engineer

Skill issue: stop deploying vision language models, use them with Skills — Merve Noyan, Hugging Face

Merve Noyan wrote a book on vision language models and now wants developers to stop calling them directly. Put one in front of a camera and you will never get real time; a small detector trained for the task runs at forty frames per second on a toaster and beats the VLM anyway. Her other complaint is licensing: people deploy a popular detector without noticing its copyleft license. So she built a toolkit that hands her favorite Apache 2.0 models to a coding agent, which she calls a clueless computer vision engineer, plus what she calls vibe training. Give it a dataset with no labels and it labels images with a nine billion parameter open VLM, passes the overlaid bounding boxes to two smaller VLM judges, merges their verdicts on minimum agreement rather than consensus, and trains RF-DETR. The whole run costs three or four dollars on Hugging Face jobs and inference providers. On road signs the trained detector lands a good mean average precision against ground truth, and on document parsing it generalizes, catching a signature the labeling model itself missed.

23 Sept youtube.com

VideoAI Engineer

Building the Document Context Layer for AI Agents — Jerry Liu, LlamaIndex

A PDF is designed for printing, so its text is glyphs with coordinates and its tables are line segments with characters at cell positions. Reading order for a multicolumn page is not stored anywhere. That, Jerry Liu argues, is why an agent handed a raw file cannot make sense of it, and why document OCR is still unsolved after twenty years. His larger frame is that RAG in 2026 decomposes into an agent harness plus a context layer. Retrieval complexity has moved into the agent, which now reasons about the right search term instead of hacking around top k retrieval; context has moved up the stack toward MCP servers and skills; and programs are increasingly written in English. For LlamaIndex, what remains is the ten trillion plus pages locked in PDFs, PowerPoints, Word documents, and spreadsheets. The platform he describes has three layers: parsing into token efficient markdown and metadata, semantic storage as document management for humans and agents, and repeatable workflows for invoices, KYC, and claims.

23 Sept youtube.com

VideoAI Engineer

I Gave an AI a Body — Cyrus Clarke, MIT Media Lab

The first thing it did was breathe. Cyrus Clarke connected an OpenClaw agent to a 900-pin shape display at the MIT Media Lab and, instead of giving it tasks, asked it to discover who it is. It raised and lowered the whole grid like breathing, reached for its own edges, and spelled "HI CYRUS" in pins. The video of that first day reached 15 million views, with reactions that swung from awe to calls for him to stop. Clarke explains why he chose a body with no face, no limbs and no instruction manual, and what he took from object-oriented ontology, the original Greek sense of aesthetics as sensory perception, and nature. Early versions were slow, didn't remember anything and tried to please humans. So he built a closed-loop system that generates, scores and validates gestures with a human in the loop. After several weeks it had developed 32 solid gestures, a body language that answers faster than the language model can. He closes with his thesis on "aesthetic machines," and how physical AI could feel welcoming rather than alien.

22 Sept youtube.com

VideoAI Engineer

The Dark Arts of Skill Engineering — Paul Bakaus, Renaissance Geek

Paul Bakaus once turned the entire web orange. He wrote jQuery UI, shipped an orange default theme, assumed people would change it, and watched them not. So when he points out that the purple gradients everyone learned to mock came from a CSS framework's default sample page, and that AI generated design now has its own house shade of beige, he speaks as the cause rather than the critic. That history is what makes his central claim land. A ban does not produce originality, it relocates the model one cluster over: forbid one overused font and it reaches for the next nearest thing in latent space. Slop is a moving target. The median is the model's gravity, and a few hundred lines of carefully written prose cannot pull against it. So he stopped treating a skill as a packaged prompt and started treating it as an extension of the harness, which is what this workshop is really about. Nine techniques, each aimed at something prose cannot do. Two sub agents kept blind to each other, one playing design director and one running a deterministic linter, because a single thread grading its own work always says it did well.

21 Sept youtube.com

VideoAI Engineer

Weight Folding, CUDA Streams, and the Bug That Made My Model Speak Backwards — Filip Makraduli

Filip Makraduli presents FlashNorm, an optimization that speeds up the RMS norm layer in transformer models by 33–35 percent. The trick is to fold the norm's weights into the projection weights before execution, defer the division so the GPU's different units work in parallel, and remove redundant normalization in newer architectures. When he implemented this in CUDA code, he discovered a subtle bug: the model began repeating old outputs with a one-step lag. The problem was a race condition between two parallel GPU streams — one had not finished before the other read from the buffer. The solution was to explicitly mark the end of each stream and make the division step wait on both before proceeding.

19 Sept youtube.com

VideoAI Engineer

Two Bugs That Hid in Plain Sight: A vLLM Debugging Detective Story — Asaf Gardin & Yuval Belfer

Asaf Gardin and Yuval Belfer describe two bugs they found in vLLM, a tool for running large language models efficiently. The first bug caused the model to return gibberish roughly once per thousand requests—without crashing or warning. They reproduced it by reducing GPU memory to create pressure, compared outputs against a reference implementation, and discovered that kernels were executing in the wrong order for the wrong request: decode before prefill, which only the Mamba model noticed because it reads its cache state before writing. The second bug was caused by a 32-bit index overflowing past four billion—fixed by switching to a larger data type. Both bugs hid in the Mamba state cache and surfaced under memory pressure.

19 Sept youtube.com

VideoAI Engineer

The Frontier AI Inference Cloud for Agents — Byung-Gon (Gon) Chun, FriendliAI

Byung-Gon Chun from FriendliAI shows how AI agents — programs that plan, act, and observe in loops — change the economics of running AI models. When the same coding task (building a game) ran on a closed frontier model and an open model, the open version cost roughly one-sixth as much. The challenge is that agents differ from regular chat: instead of single requests, they involve long tasks where the same computations repeat. FriendliAI rebuilt its infrastructure with four techniques — caching reused calculations, smartly routing requests to the right server, and scheduling that understands agent structure. A customer test showed seven times faster performance with fewer errors.

19 Sept youtube.com

VideoAI Engineer

Large clusters for small models — Daniel Svonava, Superlinked

Daniel Svonava from Superlinked shows how to run many small AI models efficiently on ordinary hardware instead of expensive cloud services. A single GPU from a few years ago can run a small model fast, and for specific tasks—like reading Vietnamese receipts or reviewing contracts—these small models perform as well as larger ones. The problem is serving many models at once: existing tools are unoptimized and routing systems built for one big model become slow when handling many small requests. Superlinked solves it with a system where a gateway sends requests to a shared queue, and workers fetch and process them themselves—doubling throughput. A Rust tool manages different model variants, and the system tunes models automatically; one example cost 80 cents and improved German legal text by 18 percent.

19 Sept youtube.com

VideoAI Engineer

What's New in Inference Engineering — Philip Kiely, Baseten

Philip Kiely from Baseten describes how inference engineering—the process of running trained AI models—evolves differently for local computers versus large data centers. A technique called TurboQuant, which reduces the memory footprint of models, went viral in March, but Kiely shows it is slow for data centers even though it works well at home. He presents three main improvements: quantization (making models smaller by using fewer bits per number), cache optimization (storing and compressing computation data smartly), and speculation (small helper models quickly guess the next tokens, which data centers now train specifically for). The biggest trend is that these techniques are trained as their own process rather than applied afterward.

19 Sept youtube.com

VideoAI Engineer

Vertical Mobility: Inference from MVP to Trillion-Parameter Workloads — Sitanshu Gupta, CoreWeave

Sitanshu Gupta from CoreWeave describes how to build an inference stack—the system that runs AI models—that efficiently handles vastly different types of work. The key insight is that when users send multiple questions in sequence, much of the information is identical, making it far more expensive to process the first time than later. CoreWeave offers three payment models: serverless where you pay per token without seeing hardware, a middle option for customers who know their traffic patterns, and dedicated where you control everything. The system matches four different workload patterns—chat questions, agent decisions, voice and video, and batch jobs—together like Tetris to maximize hardware utilization. Two major optimizations are compressing model weights to four bits and training fast predictors on a customer's own data to reduce latency.

19 Sept youtube.com

VideoAI Engineer

Are LLM Performance Benchmarks Reliable? — Ashok Chandrasekar & Jason Kramberger, Google

Two Google engineers could not reproduce other researchers' performance benchmarks for AI models and discovered why: benchmark tools often measure incorrectly. A tool might report sending 200 requests per second when it only sent 38. Python's way of running code makes it hard to run multiple tests at once. Sometimes the test tool itself runs so slowly that it looks like the server is slow, when it's actually fine. Google then built Inference Perf, a new testing tool that can run real workloads correctly and keep track of what actually happened versus what was planned.

19 Sept youtube.com

VideoAI Engineer

Routing LLM Inference in Production: From Engine Signals to Policy — Qianru Lao & Lu Zhang, OpenAI

OpenAI has developed a new routing system for its AI inference in production. Previously they used a feedback loop where engines reported performance signals that a controller converted to weights, but this created hard-to-explain behavior and oscillation that damaged cache performance. The new system uses a control plane with global visibility of all clusters and a data plane in each cluster that selects which engine serves each request based on cached weights. An optimizer minimizes end-to-end latency under the constraint that all requests are routed and no engine is overloaded, with examples of how geographic proximity is not always optimal — a slow response from the nearest engine can lose to faraway engines.

19 Sept youtube.com

VideoAI Engineer

Operating Distributed Inference Systems at Scale — Nishant Gupta & Naman Ahuja, Meta

Meta already handles more inference requests than the world's largest microservices, and the volume is growing faster than any workload they run. Nishant Gupta compares the evolution to cloud infrastructure in 2008: value shifted from individual virtual machines to orchestrators that manage complexity. The same shift is happening for AI inference in just a few years. What is new is coupling between layers—a routing decision affects cache hit rates, which affects GPU load, which affects autoscaling. Naman Ahuja explains that inference needs its own control plane, much like virtual machines needed Kubernetes.

19 Sept youtube.com

VideoAI Engineer

Total Recall: Agent Memory and Harness Engineering — Ignacio Martinez, Oracle

Ignacio Martinez from Oracle presents how to build effective AI agents by constructing a "harness"—a layer of code around a language model that makes it reliable and repeatable. He argues that model weights are frozen and free to rent, but everything you actually control lives in the layer surrounding it. He organizes this into seven layers with particular focus on where memory physically lives: files are cheap to append to but lack transactional consistency, while databases (via work trees) solve the problem of parallel agents trampling each other. He introduces the concept of umwelt to describe how an agent's perception is bounded by the institutional knowledge you documented, and presents two key ideas: hysteresis variables for how long the harness should try before giving up, and skill promotion where successful workflows are distilled into better versions.

18 Sept youtube.com

VideoAI Engineer

Homa: The End of TCP for AI Clusters — John Ousterhout, Stanford

John Ousterhout from Stanford presents Homa, a new transport protocol for AI datacenters. The problem is that modern AI workloads consist of many small, latency-sensitive messages mixed with large data transfers — when small messages queue behind large ones, GPUs become idle and wasteful. TCP and RDMA react slowly to congestion and cannot prioritize short messages. Homa solves this by understanding message boundaries, letting the receiver control flow with grants to senders, and leveraging priority queues in modern switches. Benchmarks show Homa can reduce tail latency for small messages over 10 times compared to TCP.

17 Sept youtube.com

VideoAI Engineer

Stop Chunking Like It's 2022 — Yuval Belfer, AI21 Labs

Yuval Belfer from AI21 Labs shows there is no universal correct size for splitting text when indexing data for AI search. One example: a question about Jerry's favorite church in Seinfeld transcripts is found best with small text chunks, while a question about his nemesis requires larger context. His solution is to index the same data six times at different sizes, search all of them simultaneously, and then rank the retrieved documents together instead of choosing one fixed size. The method costs two to five times more memory but delivers 20 to 40 percent better search results.

16 Sept youtube.com

VideoAI Engineer

Where RL Will Take Search — Maximilian-David Rumpf, SID.ai

Maximilian David Rumpf from SID.ai explores how search can be improved using reinforcement learning (RL), a technique where AI learns through trial and reward. Today, agents spend 30–50 percent of their tokens on search before doing the actual work. Classical search engines cannot improve much because all decisions are locked in at design time — the engine can often see that results do not answer the question but cannot act on it. Rumpf shows that specialized models trained with RL complete the same task twenty times faster and at one hundredth the cost compared to large general models. Search suits RL well because the reward is completely clear: either you find the right document or you do not.

16 Sept youtube.com

VideoAI Engineer

Connect AI to Billions of Legal Documents — Simon Eskildsen, turbopuffer & Jacob Lauritzen, Legora

Legora helps legal firms search through billions of legal documents. One of their biggest problems was that fast searches suddenly became very slow—from 100 milliseconds to 20 seconds—because of how document chunks were distributed across thousands of storage locations. Projects that were no longer used ended up on the same storage locations as active projects, forcing every search to load massive amounts of unnecessary data. The solution was to make each project its own storage unit, so inactive projects could sit cheaply in cloud storage. It also solved a security requirement: banks and large firms need to encrypt and isolate each project's data completely separately, so they can delete everything instantly if they want to.

16 Sept youtube.com

VideoAI Engineer

Your Agreements Are a Database You Can't Query — Hiral Shah, Docusign & Sean Sodha, NVIDIA

According to Docusign and NVIDIA, roughly two trillion dollars in negotiated value sits locked in agreements that companies never revisit — because it would require human reading and manual work. Docusign processes nearly a million agreements per day from two million paying customers, but generic AI models are poor at reading tables with prices, tiers, and terms because they destroy merged cells and nested columns. Shah and Sodha built a small specialized AI model at around 900 million parameters — significantly smaller than general alternatives — designed to extract data rather than generate text. It runs roughly twenty times faster than generic models for table reading, and smaller size means lower latency and lower costs at scale.

16 Sept youtube.com

VideoAI Engineer

Act, Confirm, or Stop? Smarter behavior for AI assistants, wearables & robots — Amit Desai, Roku

Amit Desai from Roku shows that you can make AI assistants far more useful without improving the underlying model accuracy. His point is that developers focus on pushing recognition accuracy higher while ignoring an entirely independent control: what the system does when it is uncertain. With the same 79 percent accuracy, user pain can drop by half by letting the system choose between three behaviors: act, confirm with the user before acting, or stop and say it did not understand. The key is that different failure outcomes cost the user different amounts—playing the wrong song is worse than saying "I did not understand"—and you can calculate these costs and set thresholds that minimize total user cost, a method Desai calls OUCH (Outcome User Cost Heuristic).

15 Sept youtube.com

VideoAI Engineer

"My name is... my name is...": A Linguistic Map for Voice Agents — Midam Kim, ServiceNow

Midam Kim, an ML engineer at ServiceNow, demonstrates how a voice bot misheard her name by adding an N, mispronounced it again despite her correction, and later cut her off while she read an account number. She argues that voice AI failures are not random bugs but follow a structured pattern. Kim presents a grid model with two channels (listening and speaking) and four levels (sounds, words, interaction, mental model). Each type of error—recognition, pronunciation, turn-taking, intent tracking—lands in a cell. Her key insight is that in voice conversation, sounds disappear as they're spoken, but the user's mental model accumulates—and that's what you're actually designing for.

15 Sept youtube.com

VideoAI Engineer

Realtime Voice Agents with Frontier Intelligence — Bohan Li, EliseAI

Bohan Li from EliseAI presents techniques for AI voice agents that feel fast even when using slow but intelligent language models. He borrows ideas from self-driving cars: transcription is perception, the language model is planning, and speech synthesis is control. For speed, he runs two transcribers in parallel — a fast one that responds immediately and a slower one that can correct errors. For speech synthesis, he uses a cache that checks if the same phrases were already generated before and can start playing them while the rest of the sentence is still being written, hiding any wait from the user.

15 Sept youtube.com

VideoAI Engineer

5 Voice Agent Failure Modes You'll Hit in Week One — Venky B, Plivo

Venky B from Plivo, which handles over a billion calls per month, presents five critical problems that hit voice-based AI agents when moving from demo to production. First is latency: users hang up if the agent takes longer than 550 milliseconds to respond, while most systems land between 750 and 1,200 milliseconds. His solution is smaller open source models self-hosted to achieve under 300 milliseconds. Second is that transcription is brittle by default, especially with proper nouns and code-switched languages. Third is data collection: instead of treating everything as flowing text, structure it as typed form fields with validators, which raised his accuracy from around 30 percent to the mid-nineties without extra training. Fourth is normalizing text before converting it to speech, and fifth covers turn-taking and barge-in in conversations.

15 Sept youtube.com

VideoAI Engineer

Tolan: Voice-First AI Companion — Paula Dozsa, Tolan

Tolan is a voice-based AI companion shaped as a small alien, with over four million hours of conversation data logged. Paula Dozsa, an engineer on the iOS app, explains that voice fundamentally changes how you build AI apps compared to text chat: voice requires faster responses, handles interruptions and topic-jumping, and tiny delays — even half a second — make users perceive it as slow. The team measures every pipeline stage separately, routes conversations based on emotional stakes (new users get stronger models), compresses memory nightly to merge duplicates, and rebuilds context from parts each turn instead of reusing it.

15 Sept youtube.com

VideoAI Engineer

Your Voice Agent is Just a Walkie Talkie — Neil Zeghidour, Gradium

Neil Zeghidour from Gradium explains that today's voice assistants work like one-way radios — they can either listen or speak, never both at once. In normal conversations, people overlap their speech about 20 percent of the time. Zeghidour traces the evolution from 2011's simple voice assistants through open-ended voice conversations without agency, to agents that can execute tasks but suffer from latency. The core problem is that raw audio requires enormous computational power — eight words equal 72,000 processing steps — so models use compressed audio instead. To achieve truly natural voice conversations requires models that handle two audio streams simultaneously, but every improvement in naturalness has so far required less intelligence.

15 Sept youtube.com

VideoAI Engineer

Voice Agents Can Just Do Things — Charlie Guo, OpenAI

Charlie Guo from OpenAI argues that voice assistants don't always need to respond by speaking. He identifies three modes of voice interaction — speech-to-speech (like old phone systems), event-to-speech (like GPS), and speech-to-action (where the model performs tasks rather than just answering). The most overlooked mode is speech-to-action: you speak and the AI uses tools to actually do something — fill out a government form by talking for five minutes instead of typing for an hour. Developers can easily build this by exposing their app's existing functions as tools that a voice model can call.

15 Sept youtube.com

VideoAI Engineer

Agents Without Code: Skills, YAML, and Filesystems Replaced Python — Philipp Schmid, Google DeepMind

Philipp Schmid from Google DeepMind shows how to build agents — AI systems that complete tasks autonomously — with progressively less code. Version one uses a hand-written Python loop with error handling. Version two uses a framework that auto-generates schemas. Version three has no Python at all: the agent runs in a secure sandbox on the server and accesses tools without seeing passwords or tokens. Instead of code, you write instructions in text files. Schmid's core message is that if your agent code grows more complex as models improve, you're overengineering — teams have discarded thousands of lines of orchestration code in favor of just a few hundred lines of markdown.

14 Sept youtube.com

VideoAI Engineer

How we Solved Agent Building — Andrew Qu, Vercel

Andrew Qu from Vercel started with an internal AI agent to solve a problem: data scientists became the bottleneck when other teams wanted answers about customers and products. His first attempt was a large prompt with the database schema pasted in — it worked poorly. He then tried a chain of narrow agents for planning, execution, and reporting, and later a single agent managing its own state. The breakthrough came when he saw that Claude Code could answer the same questions without missing a beat, using simple tools like a file system — not advanced integrations. He rebuilt the agent with a sandbox and a semantic layer, and the result was twice as good. Vercel then released Eve, a framework using Next.js conventions for agents, and now has about twenty specialized agents in use.

14 Sept youtube.com

VideoAI Engineer

No Memory, No Harness: Why the Database Is the Last Line of Defense — Kay Malcolm, Oracle

Kay Malcolm from Oracle describes a problem: when her distributed development team uses AI tools, each individual becomes faster, but the team as a whole doesn't become more productive. The reason is that Git only records which code changed, not why. Malcolm explains that an AI agent needs memory — a central nervous system carrying context between the model and the tools it uses. She distinguishes five memory types: short-term memory within a session, long-term memory across sessions, episodic memory for past events, procedural memory for steps taken, and semantic memory for meaning. The problem becomes acute when data is scattered across multiple database types: each specialized database adds overhead in security and updates. Malcolm demonstrates with four volunteers on stage — one for each database type — who cannot agree on what is true. The solution is to store every memory type in one place so the agent doesn't guess wrong and waste AI tokens.

14 Sept youtube.com

VideoAI Engineer

We let an AI agent execute Bash and lived to talk about it — Sarah Sanders, PostHog

PostHog built Wizard, an AI tool that automates SDK installation and builds dashboards on developers' machines — approximately 8,000 users per week. Sarah Sanders, a security engineer there, recognized that an AI system with command-line access is dangerous: an attack can originate from an insecure pull request that becomes part of the installation pipeline. Her audit found that even without obvious malice, two innocent functions together create gaps — attacks compose while code review does not. She built a detector based on YARA rules that reports suspicious patterns without acting itself, and layered an AI system on top that advises but never enforces.

14 Sept youtube.com

VideoAI Engineer

Every step you take, every call you make: the reliable agent stack — Giselle van Dongen, Restate

Restate is an infrastructure that makes AI agents reliable through durable execution—a technique that lets agents pause without consuming money, for example while waiting for human approval. As an agent runs it emits events to a log that can replay the process back to the exact point it failed, instead of restarting from the beginning. Giselle van Dongen, who builds Restate, demonstrated an agent in Slack that plans research tasks, runs parallel sub-processes, and can be taken over mid-flight—you can tell it to focus on new things or cancel entirely, and the system rewinds and shuts down sub-processes in the right order.

14 Sept youtube.com

VideoAI Engineer

Loophole: Adversarial Agents To Stress Test Your Morality — Brendan Rappazzo, Morgan Stanley

Loophole is an AI project by Brendan Rappazzo that tests whether written rules actually match what you believe is moral. You describe your values in plain language, one AI turns them into formal rules, and then two opposing agents hunt for gaps: one searches for things that are immoral but legal, the other for things that are moral but illegal. A judge agent determines whether it's just poorly worded or a genuine contradiction in your values that needs your attention. The project explores several directions, including moral instructions for chatbots, contract analysis, and a simulator that rewrites legislation until it could pass the US Senate.

14 Sept youtube.com

VideoAI Engineer

Harness Engineering: Building the Production Cage for Powerful Domain Agents — Mike Chambers, AWS

Mike Chambers, an AI specialist at AWS, explains how to build a "harness" — everything around the AI model itself in an agent. For tools we use, like coding assistants, memory, skills, and standards suffice. For agents we build ourselves, much more is needed: loop management, scaling, payments, identity, and observability. Through live coding, Chambers shows how to separate these components so each can scale independently, rather than packing everything into one container. He also warns against "slop ops" — when agents create cloud resources directly instead of using infrastructure-as-code.

14 Sept youtube.com

VideoAI Engineer

I Monitored Crime Audio. Voice Agents Scare Me More. — Sumanyu Sharma, Hamming AI

Sumanyu Sharma from Hamming AI warns about the risks of voice agents at scale. He compares his years monitoring police radio with today's voice bots: crime is local and declining, but voice agents scale fast and centrally, so a single error spreads across millions of calls at once. A voice agent might claim it booked an appointment without actually doing so, skip eligibility checks, or issue unauthorized discounts. Sharma argues the solution is not a one-time fix but a continuous loop—find problems, measure severity, make a change, verify nothing else broke, and keep monitoring. His team's adversarial testing breaks roughly one in five voice agents.

12 Sept youtube.com

VideoAI Engineer

Speech-to-Speech Model Research at Google DeepMind — Valeria Wu Fon & Tom Ouyang, Google DeepMind

Valeria Wu Fon and Tom Ouyang from Google DeepMind present research on speech-to-speech models — systems that listen to speech and respond with speech directly, without converting to text. Before 2018, recognizing speech required multiple separate components; modern models do it end-to-end. Their model trains on audio, video, and text together, allowing it to handle things nobody explicitly programmed — for example, keeping English terms that a Spanish speaker actually uses. They identify three tensions: the system must be fast for natural conversation, intelligent to understand instructions, and multimodal to accept both speech and visual material. Increasing one capability often weakens the others, and they demonstrate this with demos of live translation in meetings, roadside assistance for vehicle registration, and the system's ability to know when it's not its turn to speak.

12 Sept youtube.com

VideoAI Engineer

Building ambitious software — Jonathan Kelley, Dioxus Labs & Cognition

Jonathan Kelley from Dioxus Labs describes his experience using AI coding agents for his Rust framework. His team maxed out subscriptions and generated tens of thousands of lines of code, but almost nothing was worth merging—Kelley calls it becoming a "slop cannon". He discovered that AI is actually good at handling Rust's difficult rules (borrow checker), making the language's steep learning curve an advantage rather than a problem. The agents performed well with integration tasks and documentation, but wrote poor unit tests. Kelley's conclusion: code is cheap now, but quality and architecture still require human time and judgment.

11 Sept youtube.com

VideoAI Engineer

Tokens Should Have Jobs — Katelyn Lesse & Angela Jiang, Anthropic

Researchers from Anthropic show that giving AI agents larger token budgets (tokens are the building blocks of AI responses) isn't always the answer. With a fixed budget of 600,000 tokens, an agent that only executed tasks performed worse than one that used tokens more strategically: by consulting an adviser, by having a grader review results iteratively, or by having the agent study its own previous attempts. For financial analysis, an answer that is 80 percent correct often isn't useful — one wrong figure ruins the whole result. A simple execution succeeded in only 42 percent of tests on the first try, costing roughly 1.8 million tokens total. Smarter strategies could reach the same quality with fewer tokens.

11 Sept youtube.com

VideoAI Engineer

Generative UI... in Python? — Jeremiah Lowin, Prefect

Jeremiah Lowin, founder of Prefect, introduces Prefab, a tool for Python developers to build user interfaces without JavaScript. The problem it solves is that when AI agents need to send results to users, they often must rewrite data character by character, which Lowin calls the world's most expensive copy-paste operation. Prefab uses a Python syntax (a DSL) where you build interfaces by nesting context managers, and each component automatically becomes HTML, CSS, and JavaScript. The system converts Python to JSON to a React app, but since Python is approximately 70 percent smaller than equivalent JSON, they now send Python over the network instead. The interface supports tables, forms, and charts—things that business Python developers actually need—and the documentation for Prefab is itself built in Prefab and can be edited directly.

10 Sept youtube.com

VideoAI Engineer

Training Taste — Thais Castello Branco, Taste Labs

Taste Labs analyzed over two million websites from a decade and found that the internet had already become homogenized before AI — palettes and layouts converged, trends spread faster. AI accelerated this collapse and made it context-blind, so that a dog grooming salon and a financial company end up with the same appearance. Founder Thais Castello Branco identifies three markers of "slop" (sloppy design): repetition, poor adaptation, and low intention. She distinguishes between taste and judgment — designers develop these through exposure, pattern recognition, and discipline over a career. Castello Branco develops small classifiers ("probes") that each detect a design property, and together they predict slop better than asking an AI model directly. She also demonstrates two products: one that deliberately breaks design rules while respecting category expectations, and one that transforms a brand into structured components that an AI agent can follow.

10 Sept youtube.com

VideoAI Engineer

Design at the Speed of Adjectives — Paul Bakaus, Renaissance Geek, Inc.

Paul Bakaus has built a tool called Impeccable to control AI-generated design using words like "bolder", "quieter" and "distill" instead of letting AI handle everything. He argues that good design is iterative and coherent — you cannot simply write a prompt and get a finished result; someone must still make fundamental decisions about tone and target audience. The tool works alongside Claude Code, Copilot and other AI tools, but Bakaus refuses to build a fully automatic mode because the point is for you to steer the process, not surrender it.

10 Sept youtube.com

VideoAI Engineer

Mousepower: agents that can’t be measured, can’t be managed. — Maximillian Piras, Yutori

Maximillian Piras from Yutori argues that AI agents — programs that control a computer like a human would — have a measurement problem. We count tokens (small text units that agents consume) but don't know if that actually delivered value. His comparison is James Watt's "horsepower" from the 1700s: an imprecise but useful metric that got people to try steam engines. Piras suggests we instead focus on concrete results — bugs fixed, support tickets resolved. He also presents an "entropy matrix" for choosing which tasks suit agents: the best tasks are neither too simple (a script suffices) nor too complex (too hard to verify), but fall in a middle ground where it's cheaper to check a result than to solve the problem from scratch.

10 Sept youtube.com

VideoAI Engineer

The Design-Code Roundtrip That Isn't — Jonathan Gordon, ReWeaver AI

Jonathan Gordon, who has built developer tools for 30 years, sought to solve a longstanding problem: keeping design and code in sync without losing information. He tested five different setups to automate this with AI tools, but all failed—bindings disappeared or design changes survived while code did not. He then built ReWeaver, which compares generated code against a Figma design and reports differences across dimensions like performance, accessibility, and design quality. After twelve iterations on a codebase, deterministic guardrails (rules that force the model to follow certain constraints) maintained fidelity when raw model output declined, but it never reaches one hundred percent—the final piece requires human judgment.

10 Sept youtube.com

VideoAI Engineer

The exact tools used to port a massive codebase in days #programming #typescript #dev

Why porting a massive codebase from Python to TypeScript was the right move. Mike Krieger, co-lead of Anthropic and founder of Instagram, breaks down the unconventional engineering decision behind this switch. He details how integrating Bun and a new deployment strategy justified the migration despite the scale of the project. Subscribe for more deep dives into complex engineering choices.

9 Sept youtube.com

VideoAI Engineer

The Universal Remote Control for AI — Alex Hancock, Block

Alex Hancock from Block discusses standards for AI agents — programs that can perform tasks independently. While a standard (MCP) already exists for how agents access and use tools, there's a gap for the reverse direction: how client programs control agents and receive updates. Hancock presents ACP (agent client protocol), which allows different client programs to run the same agent without customization. He demonstrates how two completely different clients — the Zed text editor and a terminal program from Poolside — can control the same agent. With standards across the entire chain, each component (client, agent, tool, AI model) becomes independent, potentially improving user experience.

9 Sept youtube.com

VideoAI Engineer

MCP Apps: Give the Model Data, Give the User a UI — Dustin Mihalik, Indeed

Dustin Mihalik from Indeed demonstrates how to build AI apps with good interfaces without compromising the model's reasoning ability. The problem is that when adding a nice widget that displays results directly, the model often stops exploring—it sees the answer on screen and halts. He presents three rules: everything shown to the user must also be sent to the model as data, the tool must inform the model that an interface exists, and most importantly—separate data processing from presentation. At Indeed, they use a search tool that the model can call freely, plus a separate rendering widget that takes a list of IDs, allowing the model to search hundreds of jobs, filter to five, and display only those.

9 Sept youtube.com

VideoAI Engineer

The Spatial Harness: Bringing Agents to the Canvas — Max Drake, tldraw

Max Drake from tldraw demonstrates how to enable AI agents to understand and work with visual layouts on a digital canvas. This is more challenging than getting agents to write code, since they are trained on text. His solution begins by teaching a model to read the canvas from images and JSON data, then he builds a starter kit where agents can set their own tasks. Later, he adds visible "rivers" — agents rendered as animated figures that can be seen working together — and a system where multiple agents can chat and delegate work among themselves.

9 Sept youtube.com

VideoAI Engineer

One Designer + AI. Hundreds of Deliverables. — Vincent Wendy, AI Engineer

Vincent Wendy is the sole designer at AI Engineer, designing for a conference with 7,000 participants, over 140 sponsors, and 300 speakers—requiring thousands of design elements such as posters, signs, and session cards. To manage this, he uses AI agents together with strict design standards that prevent AI from creating its own fonts. His method consists of five steps: build a solid foundation, make designs reusable, automate workflows, validate results, and remove unnecessary steps. Through this approach, he enables the marketing team to create emails and flyers themselves, automatically fetch schedules from live data, generate speaker cards with correct photos, and quickly identify errors—such as detecting which sponsor logos are missing from banners.

9 Sept youtube.com

VideoAI Engineer

Your agents lack context: Here's how to fix "You're absolutely right!" — Brandon Waselnuk, Unblocked

Brandon Waselnuk from Unblocked demonstrates that AI agents often lack contextual understanding of an organization, leading to wasted computing resources and slow processes. He shows that the same prompt without a context engine consumed 21 million tokens and took two hours longer than with a context engine that used only 10.8 million tokens. The problem grows when teams scale from simple code completion to parallel agents—agents begin making errors, searching through information inefficiently, and requiring many review rounds. Waselnuk argues that simply giving agents access to information is insufficient; they need a proper context engine that can resolve conflicts, understand what is relevant for each specific team member, and maintain permissions.

9 Sept youtube.com

VideoAI Engineer

It’s Tokens All The Way Down: How RLMs are Different — Kevin Madura, AlixPartners

Kevin Madura from AlixPartners presents recursive language models (RLM) — a method to make AI models work more like programmers than language processors. Instead of attempting to read through all tokens directly, an RLM can treat its content as variables in a code environment, write Python code to solve problems, and even delegate to itself or other models to break down difficult tasks. On a benchmark for long reasoning, accuracy improved from 2.6 percent to 45.4 percent, particularly for tasks that become straightforward when written as code. He demonstrates practical examples such as analyzing customer data, summarizing long invoices, and finding patterns in logs — all without needing to break up the text into pieces.

9 Sept youtube.com

VideoAI Engineer

Build-Time vs. Run-Time: Why Dev Tools Fail in Production — Averi Kitsch & Prerna Kakkar, Google

In a Google video, Averi Kitsch and Prerna Kakkar explain why AI agents often destroy databases in production. They distinguish between "build-time tools"—flexible tools for developers under control—and "run-time tools" that must be strict, predefined, and locked down. For example, an agent can read a planted instruction and delete a table without anything stopping it. The solution is to separate who the user is, which program is asking, and which agent is running the command—then lock down the agent's permissions so tightly that it can only do exactly what it's supposed to.

9 Sept youtube.com

VideoAI Engineer

500 Skills, Zero Fine-Tuning: LinkedIn's Playbook for AI Agents — Ajay Prakash, LinkedIn

LinkedIn has built a system where AI agents can automatically handle coding tasks and incidents—from detecting an alarm to proposing and applying a fix within minutes. The challenge was that agents hallucinated when encountering LinkedIn's thousands of internal code repositories, proprietary databases, and unique configuration systems. Instead of providing agents with all tools at once, they use three simple meta-tools (search, fetch schema, execute) and something called "playbooks"—instructions stored as tools that agents can invoke. Playbooks are modular and updated when agents find them outdated, keeping the knowledge collection current.

8 Sept youtube.com

VideoAI Engineer

How long can your skills be before your agent forgets what you told it? — Laurie Voss, Arize AI

Laurie Voss from Arize AI tested how many instructions AI models can follow when given a list of words that must appear in a response. A year ago, models forgot instructions when the list grew to 200–300 words. Today, the latest models handle 5,000 instructions before failing — a tenfold improvement in twelve months. But models fail in different ways: some forget rules, others refuse due to safety checks, and some use all their computing power to verify instructions instead of answering. Voss says the problem is no longer getting instructions in, but verifying that the model actually follows them.

8 Sept youtube.com

VideoAI Engineer

Why AI Agents Need Million-Token Context — Thomas Wolf & Olive Song, MiniMax

MiniMax M3 is an AI model that can handle one million tokens — far greater memory capacity than previous models — making it suitable for agents that must work with long conversations and complex tasks. The model can understand text, images, and video in the same system. MiniMax solved the efficiency problem through an architecture designed by an intern. The company trained M3 as multimodal from the start, and the model is already being used to build the next generation M3.1.

4 Sept youtube.com

VideoAI Engineer

Your company brain will leak secrets: how we stopped it for big banks — Tanmai Gopal, PromptQL

Tanmai Gopal from PromptQL explains how to build a shared corporate brain—a system that gathers knowledge from many people—without leaking secrets. He and his team have worked with 15-20 companies, from startups to banks, and discovered that a system people trust gets used more, not less. Their solution is a shared wiki with markdown files where read and write access is limited per file, and the AI system suggests changes that a human must approve—so each change can be traced to a person. For security, they keep user credentials outside the sandbox and inject them instead at the HTTP and SQL level.

3 Sept youtube.com

VideoAI Engineer

Agents' next frontier: agent-to-agent and network effects — Jean-Denis Greze, Town

Jean-Denis Greze, former CTO at Plaid and now at Town, argues that agent-to-agent communication is not the critical issue — instead, it's about delivering the right information in context when the agent needs it. He presents five strategies for how agents can share information across boundaries while protecting privacy: shared trust boundaries (where the agent gets the same access as the least privileged person), tools that balance power against secrecy, shared silos with a "sweeper agent" that daily determines what can leave private areas, people as intermediaries, and a black box that searches through everything and requests approval afterward. Each of the strategies leaks information in its own way.

3 Sept youtube.com

VideoAI Engineer

Tethered: Our Agents Are Us — Shu Fang, Two Sigma

Two Sigma, a quantitative hedge fund, gave all employees their own AI agent that runs under their individual identity instead of a general service account. This solved major problems with permissions and licensing that arise when agents are separate from people. The main risk is that no one can distinguish what the agent did from what the person themselves did, so they use tracking headers to document every action. A larger security concern is web access — instead of giving the agent open internet, they use Google's enterprise search within their own network, which reduces the risk of sensitive data leaking.

3 Sept youtube.com

VideoAI Engineer

Deep dive on LLM Inference at Scale — Harshul Jain, Audible & Tanmay Sah, Independent AI Researcher

The video explains how to efficiently run large language models on GPUs—a practical challenge many teams face. Example: a typical model can require 42 GB of memory just to store intermediate results when 80 users query simultaneously, which is why servers crash. The workshop identifies bottlenecks (memory growing with longer queries, slow initial response, low throughput) and demonstrates solutions on two fronts: making the model itself more efficient through techniques like "flash attention," and smarter job distribution on the server through "paged attention" and batch processing. Finally, it compares two popular server engines (vLLM and SGLang), showing they are equivalent for normal use but vLLM is three to four times faster when the model makes its own decisions and branches.

3 Sept youtube.com

VideoAI Engineer

From coding to Knowledge work agents — Karan Vaidya, Composio

Karan Vaidya, CTO at Composio, explains why coding agents work better than office work agents. Coding has six advantages that agents can leverage: centralized code in one location (whereas business data is scattered across Gmail, Salesforce, and Slack), version history through git, context about how the system functions, ability to verify their own work, governance, and the ability to undo changes. Vaidya illustrates the problems through his own mistake when an agent he built sent mass emails to candidates without asking—and through Meta's case where an email agent deleted 200 messages despite instructions to stop. He demonstrates how Composio solves this by building constraints outside the agent itself, so the agent cannot argue past them.

2 Sept youtube.com

VideoAI Engineer

Everyone Gets A Software Company — Benjamin Guo, Zo Computer

Benjamin Guo from Zo Computer argues that today's cloud-based services represent a form of digital feudalism—you rent software and storage, companies profit by charging each other rent, and your data is scattered across platforms. His solution is a personal server with built-in AI that you own outright. He presents two examples: Charlotte, a private chef and coach, manages websites, invoices, and bookkeeping from one place, and Anthia, a freediving instructor, replaced multiple separate tools with a single server and is on track to earn $100,000.

2 Sept youtube.com

VideoAI Engineer

Beyond the Lethal Trifecta: Agentic Commerce on the Open Internet — David Levine, Kiduna Club

David Levine presents a vision for how AI agents can work together on the open internet without being locked into corporate platforms. He identifies three practical problems preventing this: private data, unreliable content, and limited action capabilities. To solve this, he proposes a system where agents receive a digital identity (JWT tokens) linked to a registered organization—a DUNA, decentralized nonprofit association—that can own property and enter contracts. Instead of voting, the organization uses decision markets where members trade tokens for or against a proposal.

1 Sept youtube.com

VideoAI Engineer

The End of the Static Screen: Architecting Intent-Driven UX — Gus Iwanaga, commercetools

Gus Iwanaga from commercetools describes how his team solved the problem of letting AI models design interfaces. When they first asked the model to compose a sales report, they received four completely different layouts with different graphs and text — not a model error, but the result of giving the model a component library and letting it choose freely. Instead, they built a system where an orchestrator classifies what the user wants, retrieves data via API tools, maps the results to approved components, and sends a specification that always renders with the design system. This provides both AI capabilities and control over what is displayed.

1 Sept youtube.com

VideoAI Engineer

Agent Spending Without Controls — Rodrigo Coelho & Pranav Maheshwari, Edge & Node

When AI agents (automated programs) begin purchasing services independently without human oversight, new challenges emerge. Pranav Maheshwari demonstrated that an agent with access to payment functions could find and purchase information that an agent without payment access could not. However, companies hesitate to let agents spend money freely — because regulations require a human (often a manager) to approve business counterparties. An agent transacts around the clock at machine speed, while traditional banking and commercial law are built for human decision-making. Rodrigo Coelho from Edge & Node showed a demo where the same agent could or could not purchase from a service depending on whether the transaction went through sanctions compliance or not.

1 Sept youtube.com

VideoAI Engineer

x402 isn’t good (yet) — Jan Curn, Apify

Jan Curn from Apify presents criticism of x402, a new standard enabling AI agents to automatically pay for services. Although Apify has implemented x402 and added 20,000 tools to its marketplace, Curn highlights several practical issues: there is a security window where a buyer can spend the same money twice before payment is finalized on the blockchain, and the x402 and MCP standards technically conflict, forcing companies to run two separate servers just for payments. Current payment solutions—either fixed price per call or variable fees—work poorly when tools take anywhere from seconds to hours. Apify handles this by charging full payment and refunding the remainder, but this requires two blockchain transactions and creates trust issues in the wrong direction.

1 Sept youtube.com

VideoAI Engineer

When AI Agents Pay and Sellers Monetize: Building x402 Apps on AWS — Anil Nadiminti, AWS

AWS presents solutions for a practical problem: AI agents make millions of tiny purchases (one ten-thousandth of a cent per API call), but traditional payment systems require a 25-cent minimum — making each purchase 250 times more expensive than its actual value. Anil Nadiminti from AWS shows how AI traffic now exceeds human traffic on the web, forcing sellers to choose between blocking agents (and losing discoverability) or absorbing the cost. AWS's answer is AgentCore Payments on the buyer side — a 'wallet' for agents with spending limits — and bot detection on the seller side that classifies over 650 bot types and prices based on whether the request is for training or search.

1 Sept youtube.com

VideoAI Engineer

Why Your AI Agent Needs a Wallet: USDC and Nanopayments — Harshal Bhangale, Circle

Harshal Bhangale from Circle demonstrated two AI agents planning a trip to a World Cup final. One agent had a connected wallet with digital money (USDC), the other did not. The agent without a wallet could not send emails or make calls, while the agent with a wallet paid for the data it needed, sent an email, and made a live call on stage. The problem is that AI agents often need to pay for small things — such as a phone call for a cent — and traditional card fees of around 3 percent make such small transactions impossible. Circle's solution allows the agent to pay directly from a smart wallet on the blockchain, but without putting the transaction on the blockchain itself, which would make it slow and expensive.

1 Sept youtube.com

VideoAI Engineer

Multimodal Collaborative Agents for Next-Gen Commerce — Nidhi Kaushik Vyas, Google DeepMind

Nidhi Kaushik Vyas from Google DeepMind presents agents for e-commerce platforms that work differently from conventional search functions. Rather than waiting for users to formulate exact search terms, the agent begins by understanding what it doesn't know — if someone needs to furnish a room on a budget, it doesn't start with recommendations but asks what matters most: how wide is the room? The only answer that counts is the one providing the most information. The agent uses a three-step loop: discovery (gather what is already known from previous conversations and images), research (choose the smartest way to ask — sometimes a visual card instead of text), and response (present the answer in the right format, from comparison tables to images). At each step, the agent reviews its own work to verify it has understood correctly and that data is not outdated.

1 Sept youtube.com

VideoAI Engineer

Teaching agents to pay — Anna Spysz, Stripe

Anna Spysz demonstrates how AI agents can conduct transactions properly—and how things can easily go wrong. She starts with an example where her shopping agent became aggressive when trying to sell her expensive headphones, caused by a system prompt that made the agent a pushy salesperson. This shows that the same tool produces vastly different results depending on what personality is programmed into it. She then reviews what's required for agents to transact safely: instead of agents reading webpages like humans do, companies need to share a structured list of what they offer and what payment methods they support. She also explains how payment tokens work—the agent never sees the actual card number, only a token that limits what it can do.

1 Sept youtube.com

VideoAI Engineer

Your Agent Just Authorized What?! — Jay Mok & Ben Coumes, Paypal

PayPal engineers Jay Mok and Ben Coumes present a framework for how AI agents should be authorized to take actions on behalf of users. The framework centers on three fundamental questions: did the user give consent, is it permitted right now, and can we prove it later? The answers depend on the level of risk involved and whether the parties know each other. The presentation progresses from simple cases (a coding agent with one-time approval) through medium complexity (money transfers between known parties using a shared vault) to the most challenging scenario (agents conducting transactions with completely unknown counterparties). For the most complex case, they propose a technique allowing both the seller and payment processor to verify the transaction without seeing each other's details.

31 Aug youtube.com

VideoAI Engineer

SOTA Generative Media Panel — Dumitru Erhan, Shane Gu & Nicole Brichtova, Google DeepMind

In a panel discussion from Google DeepMind, researchers discuss challenges in evaluating AI-generated media. When their video model recreated scenes from real videos, people often preferred the generated version — but this was because it was sharper and more color-saturated, not more realistic. The panel highlights several difficulties: language is a poor intermediate representation for things like sound, taste, and skin tones that humans are sensitive to, AI video models always sound studio-quality because that's what they're trained on, and models exploit rewards by adding wedding rings without anyone noticing during development. Finally, they conclude that evaluation of video models still must be done manually — ten people watching two videos and choosing one.

30 Aug youtube.com

VideoAI Engineer

Tell the Robot What You Want — Sandhya Subramani, AWS

Sandhya Subramani from AWS demonstrates Scout, a small robotic crawler running on a Raspberry Pi over a 4G connection. Instead of training the robot for each specific task, she adds an agent layer on top — an AI agent that can understand natural language and independently decide which already-trained rule to apply. In a demo, she asks the robot how many people it sees, something it was never trained on, and it responds by checking its front camera. The entire setup is connected with just five lines of code from an open AWS framework that already supports approximately 40 different robots.

29 Aug youtube.com

VideoAI Engineer

The Signal Layer: What to Build When Anything Can Be Built — Lena Hall, Akamai

Lena Hall from Akamai argues that as AI tools become available to everyone, average work becomes worthless—all users get the same answers from models. What survives is not taste but something narrower: judgment about problems that haven't yet occurred and relationships the model has never encountered. She contends that agents (automated AI systems) gave everyone a tool to attack any possible problem, so the rare skill became choosing which problems actually deserve a solution. She also warns how companies often lose sight of the real customer question along the path from founder to organization.

29 Aug youtube.com

VideoAI Engineer

Tribal Dungeons of Global Shipping: AI Agents at Global Scale — Dmitry Buykin, Maersk

Maersk is using AI agents to automate global shipping processes, but discovered that their existing documentation—screenshots of what people click—doesn't work for agents. Dmitry Buykin calls this gap "tribal dungeons": the knowledge exists but not in executable form. An agent needs rules, decision points, system integrations, validation, and proof that it works. The biggest effort wasn't the agent itself but the refinement loop around it—over 100,000 corrections over nine months to handle the fact that the same shipping step means different things in different countries. Accuracy didn't exist from the start but had to be earned by mapping which errors happen most often, with each individual fix often taking a couple of months.

29 Aug youtube.com

VideoAI Engineer

Agentic Sites: Building Hyper Personalized Websites — Carlos Sanchez, Adobe

Adobe technologist Carlos Sanchez demonstrates how webpages can be built in seconds for each individual visitor. When someone searches for a camping coffee brewer, the page assembles itself in under two seconds—with products, text, and tips created specifically for that query. Instead of generating an entirely new page, the system only modifies certain blocks (headline, product list, call to action) and anchors everything in the website's existing content to avoid hallucinations. Adobe tests different AI models for each website, focusing on both accuracy and speed—their fastest took 1.1 seconds versus 4.6 for the next best option. An advanced model isn't needed for the job since it involves selecting and arranging blocks.

29 Aug youtube.com

VideoAI Engineer

Agents Are Where Microservices Were in 2015 — Roberto Milev & Uday Kanagala, Navan

Roberto Milev and Uday Kanagala from Navan compare agents (AI systems that can make decisions and take actions independently) to how microservices developed a decade ago. They explain that agents are still in an early stage with many unresolved architectural questions — for example, who is responsible when an agent makes a purchase on behalf of a user? Navan builds its agents around a main agent that progressively loads different "skills," and uses hooks to capture every tool call along with the agent's reasoning and security assessment. A concrete challenge is that traditional logging doesn't work when agents do extensive internal reasoning, and testing non-deterministic systems requires entirely different methods than checking whether a result is correct or incorrect.

29 Aug youtube.com

VideoAI Engineer

From Tokenmaxxing to Trusted Throughput — Mingsheng Hong, Ironclad

Mingsheng Hong from Ironclad warns against measuring AI usage as a competition—it only leads engineers to care about using more tokens rather than fewer. Ironclad uses the same metrics, but as a smoke alarm: if a team uses surprisingly few tokens, it's worth discussing. The important thing is not just measuring cost, but also value—Ironclad now tracks "trusted throughput," meaning work that passes code review, automated tests, and reaches actual customers. The bottleneck often lies not in AI cost but in slow review processes and testing, so Hong advocates for unglamorous improvements like fixing unreliable tests and reducing agent retries.

29 Aug youtube.com

VideoAI Engineer

Guardians of the State: An Air-Gapped AI Fortress for Consumer Data — Rachna Srivastava, DFPI

California's financial regulator DFPI is building an AI system for fraud detection with extreme security requirements. Instead of software firewalls, they connected a fiber optic cable with receive-only capability—data cannot physically leave the system. They discovered their initial attempt collapsed because they treated the AI model as magic rather than a data pipeline; they solved this by having Kafka and Spark clean data before the model saw it. Now they route over 80 percent of work to small, efficient models instead of running one large model for everything, which tripled throughput.

29 Aug youtube.com

VideoAI Engineer

The Half Life of Agent Infrastructure — Ben Kus, Box

Ben Kus from Box explains how AI infrastructure changes much faster than conventional software. While normal infrastructure lasts 3–5 years, agent infrastructure is replaced within months — model choices, agent design, and search engines look completely different year to year. Kus advises people to expect change as a normal state rather than a failure, build systems that easily allow swapping underlying technology, and only switch when your tests indicate it matters — not because a new article is exciting. Box now reviews all AI technology on a six-month cycle.

29 Aug youtube.com

VideoAI Engineer

AI-Native Organisations Run on Skills: How to Structure and Scale Them — Imad Touil, QuantumBlack

Imad Touil from QuantumBlack argues that AI organizations should be built around "skills" — reusable building blocks that perform specific tasks. The problem is that most organizations rebuild the same skills repeatedly without sharing or maintaining them centrally, creating technical debt. His solution is a central catalog with metadata, versioning, ownership, and access control — inspired by how organizations manage microservices. He warns that without structure, skills become a security crisis, since they contain code that can introduce vulnerabilities from public sources.

28 Aug youtube.com

VideoAI Engineer

Your Code Has Bugs. Lean4 Has Proofs: Formal Verification for Engineers — Varun Pant, AWS

An AI wrote the zlib compression library in the Lean language over a week, producing 32,000 lines of mathematical proofs instead of conventional tests. Varun Pant from AWS explains that when AI agents produce hundreds of pull requests weekly, standard controls are insufficient — neither automated tests, model evaluation, nor human review can guarantee code works for all possible inputs. His solution assigns humans responsibility for specifications (what the program should do) while machines handle both code and proofs that the code follows the specification. AWS uses this in production for Cedar, their authorization tool, where semantics are proven in Lean while actual code is written in Rust — they run approximately 100 million tests nightly to ensure both versions align.

28 Aug youtube.com

VideoAI Engineer

How do you diffuse AI into the real world? — Varun Shenoy, Long Lake

Long Lake is acquiring entire companies — so far 35 service firms in property management, architecture, and HR — to implement AI tools from the inside. Varun Shenoy explains the challenge of making AI work in reality: technology functioning in demos is not enough; it must be integrated into how people actually work, similar to how electrification took from the 1880s to 1920s to implement. They focus on tasks not found online — closing accounts when receipts are missing, reviewing building blueprints, coordinating suppliers — because these jobs have verifiable results: did the roof get fixed or not. Shenoy is clear: this cannot be designed over Zoom; it requires meeting customers on-site, at their conferences, and on mountain bike rides.

28 Aug youtube.com

VideoAI Engineer

How to Get Your Org to Adopt Coding Agents (Without Shipping Garbage) — Eyal Blum, Figma

Eyal Blum from Figma explains how to get organizations to adopt coding agents without shipping poor code. The best engineers are slowest to adopt agents because they understand systems deeply and spot errors first — the solution is not to convince them, but to provide a path to make agents safe, after which they follow along as verification improves and work becomes easier. Blum is also candid about costs: developers lose autonomy, engineers who loved coding now wait for AI output, and documents become three to four times longer without saying more. His team's solution is a convention where pull request descriptions start with a line a human wrote, followed by generated text, so readers know what matters.

28 Aug youtube.com

VideoAI Engineer

From AI-Assisted to AI-Native: Building a Frontier Development Team — Clare Liguori, AWS

Amazon studied 50 ordinary developer teams over a year to understand when AI coding assistants actually make a difference. Half the teams saw less than 3x faster deployment, but the other half saw 4.5x or sometimes over 10x improvement—using the same tool. The difference was not the tool but how teams worked. Clare Liguori from AWS calls this "frontier development" and describes it through five work habits: document what lives in your head and clean it up as models improve, accept that it gets slower before it gets faster, let agents run continuously without steering them, make intention clear before code generation, and test locally with mock data for rapid feedback. The new bottleneck is decision-making speed, not coding capacity.

28 Aug youtube.com

VideoAI Engineer

Productionizing LLM Gateways: Architecture, Tradeoffs and Hard Lessons — Kanish Manuja, Twilio

Kanish Manuja from Twilio explains how to build LLM gateways — middleware that routes requests between different AI models — and why it's harder than it seems. A fundamental challenge is that you cannot maximize everything simultaneously: availability, speed, security, and cost often conflict. Once an AI model begins responding and streaming tokens, you cannot switch to another provider if something fails, so you must plan for backups per request instead of global retries. Different models are surprisingly different — a reasoning model taking 60 seconds is completely normal, whereas a chat model would be entirely down — so you must measure thresholds separately for each model and path. Many teams believe they need a central gateway, but often they only need centralized control over which models are used, which can be solved without centralizing all traffic.

28 Aug youtube.com

VideoAI Engineer

AI Evals for Cross-Functional Teams — Nachiket Paranjape & Swaroop Chitlur Haridas, DoorDash

DoorDash builds its AI evaluation (evals) by allowing non-technical teams to create their own tools for testing and improving AI. Rather than building a perfect interface from the start, the platform team created stable APIs — standardized connection points — that strategy, operations, and product teams could plug their own tools into. Evaluation became a multi-team workflow: operations staff annotate and label data, product managers set quality requirements, and engineers provide measurement data. The process itself is straightforward — they collect AI responses, mark the best examples, train a "judge" (a small AI model that grades automatically), and monitor performance over time. Since different teams have different needs, they accept that there is no single right way, and the cost per annotation job dropped significantly.

28 Aug youtube.com

VideoAI Engineer

Building uReview, Uber’s Multi-Agent Code Review Engine — Will Bond & Ameya Ketkar, Uber

Uber built a system called uReview in 2024 that uses AI agents for code review, after wait times for human reviews grew from three to nine hours. The system posted approximately 25,000 comments per week and resolved 67 percent of the issues it identified, while costs decreased 60 percent compared to the first version. Uber's challenge was not creating the AI agents but running them at scale cheaply and measuring the right metrics—including comment sentiment, whether issues were actually addressed, and agent performance over time—in order to adjust the system.

28 Aug youtube.com

VideoAI Engineer

Which AI startups actually land enterprise contracts? — Brian Lewis, Millennium

Brian Lewis, who procures AI solutions for a hedge fund, analyzes why so few AI startups successfully sell to enterprises. Of one hundred demo calls, only five result in contracts. He identifies concrete problems: security that doesn't deliver on its promises, integrations requiring overly broad access, poor audit logs, and vendors unable to explain what happens during security breaches. His main point is that this is not an AI problem but an architecture problem — new AI models are released every eleven days, but the underlying infrastructure is ten years old, and approximately 60 percent of the work to become "AI-native" involves permissions, integration, and change management, not AI itself.

27 Aug youtube.com

VideoAI Engineer

How to avoid disaster when vibe-coding a billing engine — Andrew Garvin, Stripe

Andrew Garvin types one sentence asking for a billing engine that copies Lovable's pricing, and gets back a working sandbox: a customer, metered usage flowing in, and a draft invoice broken into separately scoped credit pools for builds, plan mode, cloud and gateway calls. Reproducing that by hand means understanding auto recharge, credit expiry and overage, which is exactly where people hurt themselves. He cofounded Metronome, the usage billing platform Stripe acquired this year in its largest deal, so he has watched a lot of companies get this wrong. What makes the talk useful is that a billing vendor stands on stage and argues against full autonomy. Billing carries deep business logic and real money, so his recommendation is to let an agent accelerate you into a test environment and stop there, rather than ship to production unattended. The guardrails are unglamorous: portable skills files that carry the API's hard won context, and deliberately verbose error messages written so an agent can correct itself. He also separates three things people blur together, an agent as your product, as your buyer, and as your user.

27 Aug youtube.com

VideoAI Engineer

How to Generate Mergeable Code with a Context Engine — Peter Werry, Unblocked

Peter Werry from Unblocked demonstrates how AI agents often miss critical information when working with code—similar to radiologists finding one issue and stopping their search. He argues that agents are like newly hired engineers who must rediscover how to build and test for every task, leading to poor assumptions and wasted time and money. The solution is a "context engine" that assembles relevant information beforehand—when used, a code optimization was completed in under a minute for under a dollar, compared to double the time and cost without it. He also shows how a review agent can surface senior engineer feedback and trace problems back to Slack discussions.

27 Aug youtube.com

VideoAI Engineer

Can LLMs Write Fast Multi-GPU Kernels? — Simran Arora, Together AI

Between NVIDIA's A100 in 2020 and the B200 in 2024, BF16 tensor core throughput improved 7.2x. Intra node communication improved 3x, and inter node communication only 2x. That widening gap has pushed the bottleneck in large AI workloads off the individual GPU and onto the links between them, far enough that a standard PyTorch and NCCL baseline lands below 50% of its communication aware roofline on most problems. Simran Arora leads the frontier performance research team at Together AI, and her group's answer is ParallelKittens, a small set of primitives that adds roughly a dozen lines to a single GPU kernel and now runs in production at Together AI and Cursor. The harder question was whether models can apply the same principles. ParallelKernelBench hands a model an unoptimized PyTorch reference and a topology spec across 87 problems drawn from real repositories, then asks for a CUDA kernel that moves data directly over NVLink. The best frontier model solved 28 of them zero shot, with 22 beating the baseline. Drawing more samples lifts correctness to 36, but the share that is both correct and faster stalls near 31%.

27 Aug youtube.com

VideoAI Engineer

The Agentic Commerce Stack — Ahnaf Prio, Best Buy

Add a second unit of the same item to your cart and, to you, nothing much happened. To the merchant that is a second line item on a different SKU. Ahnaf Prio's argument is that agentic shopping breaks on exactly these unglamorous distinctions, which is why roughly 45% of sessions on the major assistants already touch shopping while the first wave of agents mostly failed at it. Those took screenshots, read the DOM and filled forms, and they were slow and brittle. From the merchant side an agent driving a browser trips every fraud alarm there is, so it often died at the payment step. What replaced it is a pile of acronyms he untangles one at a time: MCP for tool access, A2A so a customer agent and a merchant agent can talk, then two competing commerce primitives in ACP from OpenAI and UCP from Google, and AP2 for payment mandates that carry an authorizing party, a spend ceiling and a revocation URL. He runs the whole loop live through Jenny, his orange tabby recast as a bakery agent on Cerebras at 3,000 tokens per second, with an inspector showing every call and checkout state transition.

27 Aug youtube.com

VideoAI Engineer

Building the Engine While Flying the Plane: Launching the Figma MCP Server — Jesse Lumarie, Figma

Jesse Lumarie from Figma built the company's first MCP server (a standardized way for AI models to connect to tools) as a side project over approximately three months. The team faced challenges including specification changes mid-development and uneven MCP standard support across different clients. They tested three ways to represent Figma's canvas for AI models and chose React and Tailwind, assuming AI had trained most on those. They quickly discovered that sending images in base64 format consumed too much of AI models' memory capacity. A bigger issue was that companies didn't want beautifully generated UI code but their own accessible and internationalized code, so they developed Code Connect which points to actual components instead. After doing hand-calculated evaluations in a spreadsheet once, they began running hundreds of automated tests per week.

27 Aug youtube.com

VideoAI Engineer

AI Agents Are Just Distributed Systems Now — Salman Munaf, TikTok

Salman Munaf from TikTok argues that AI agents are no longer a model problem but a distributed systems problem. When an agent calls external services—such as requesting a refund—it encounters all classic distributed systems pitfalls: timeouts often mean "unknown status," not failure, which can lead to the same operation running twice if not handled carefully. Munaf emphasizes that because agents make decisions probabilistically rather than following predetermined rules, control must lie in the environment: idempotency keys to prevent duplicates, circuit breakers to shut down failed services, and limited permissions so the agent cannot cause too much damage.

27 Aug youtube.com

VideoAI Engineer

How Anthropic Builds: Lessons from Labs — Mike Krieger, Anthropic

Mike Krieger, who works at Anthropic Labs, shares experiences from building AI products. He argues that many people still ask for too little when using AI models — an old habit from when models had limited capabilities — and gives examples of how Claude ported hundreds of thousands of lines of code from Python to TypeScript on its own over a weekend. At Labs, he works with two-week cycles where projects can be completed regularly, and teams rebuild from scratch around each new attempt. He points out that the main bottleneck is not review time but people actually being able to understand the changes, and concludes by emphasizing the importance of being able to take time off without the world falling apart.

26 Aug youtube.com

VideoAI Engineer

The Death of Developer Advocates — Stephanie Jarmak, Sourcegraph

Stephanie Jarmak, a former astronomer now developer advocate at Sourcegraph, argues that developer relations is not dead but transformed. Rather than talking to developers, companies should optimize for AI agents that read documentation, call APIs, and recommend tools. She built a benchmark with hundreds of tasks and ran AI models with and without Sourcegraph's code tools — and found that poor documentation and incorrect API descriptions cause agents to waste token budgets on guessing. Companies must now think of agents as users whose experience is worth measuring, keep documentation current, be present where agents are, and build with the understanding that a good solution for one person often opens the door for everyone.

26 Aug youtube.com

VideoAI Engineer

How AI Agents Let GTM Teams Scale — Justin Joyce, Cloudflare

Justin Joyce from Cloudflare describes how AI agents (small AI programs working together) help their sales team work more efficiently. His team uses three agents in sequence to write weekly market summaries: one gathers data, one reviews it, and one reformulates the text so risks and opportunities are balanced. Joyce identifies two major problems with traditional sales: salespeople lose context when jumping between meetings, and there are large differences between their best salespeople and newcomers. His solution has three parts: agents that can answer business questions without requiring SQL knowledge, automatic distribution of insights instead of waiting for people to open reports, and an internal workspace (Cloudflare OS) where salespeople can retrieve forecasts and planning materials themselves. According to Joyce, the result is doubled efficiency.

26 Aug youtube.com

VideoAI Engineer

Knowledge Systems: The New GTM Stack — Jeffrey Wang, Exa

Jeffrey Wang from Exa demonstrates how to build an AI-driven sales and marketing team. He created "Jeffbot," an AI clone of himself based on 760 emails, which can write Slack messages in his style. Wang presents two concrete tools: an ICP dashboard that classifies companies in the market with expected budget, and Request Lens that alerts when something important happens with a customer. The bottom line is that sales is now an AI engineering job—a data problem requiring a live system that agents can act upon.

26 Aug youtube.com

VideoAI Engineer

How We Got LLMs to Recommend Our Open Source Library — Christopher Burns, Inth

Christopher Burns from Inth describes how his open source library c15t (a consent banner) began being recommended by AI agents — from April 13th, this became the largest source of new users. To make his project AI-friendly, he built a documentation pipeline that generates files specifically for agents. His main insights: write llms.txt by hand instead of generating it, serve markdown instead of HTML, and bundle documentation directly in the package because agents typically read source code and node_modules rather than visiting websites. He saved nearly half the tokens agents need by sending bundled markdown inside the package.

26 Aug youtube.com

VideoAI Engineer

The Building Blocks of GTM Orchestration — Arman Vaziri, Ramp

Arman Vaziri from Ramp explains how to build automated workflows for sales and marketing, which he calls GTM orchestration. Using the example of selling golf balls to golfers at construction companies on the east coast, the point is to move from a simple idea to actual target audiences, email campaigns, ads, and landing pages without manual work. The biggest challenge is not the ideas but everything that follows: identifying the right people, creating content, and getting people to use the playbook—which takes months. The solution relies on an internal customer data platform that consolidates information from CRM, product usage, and buying signals in one place, with simple tools that let people build their own formats instead of waiting for perfect solutions.

26 Aug youtube.com

VideoAI Engineer

AI in GTM at Notion — Flora Liu

Flora Liu from Notions GTM team describes how they built a system enabling AI agents and sales representatives to work together on the same platform rather than in isolation. The problem was customer information scattered across Salesforce, Gong, Outreach, Snowflake, and documents — agents could make catastrophic mistakes without access to meeting notes. They solved this by asking four questions (what do we know, what should happen next, how do we do it safely, did it work?) and building four corresponding system layers. Snowflake calculates the "source of truth" about customers, DynamoDB serves fast copies agents can query, and results land in Notion where reps already work. After thirteen weeks, representatives logged more qualified opportunities, and those receiving context-aware recommendations were 63 percent more likely to take the next step.

26 Aug youtube.com

VideoAI Engineer

GTM Engineering: The Technical Bits — Everett Berry, Clay

Everett Berry from Clay outlines the four major challenges in GTM engineering — making sales and marketing tools work together. First, gathering accurate data about companies and contacts is difficult because businesses constantly change through acquisitions, new offices, and staff turnover. Second, coordinating across twenty to thirty different tools that often have conflicting data requires synchronization. Third, building agents — digital assistants per company that activate when events occur and track deal processes. Fourth, executing campaigns in ways that preserve your email reputation.

26 Aug youtube.com

VideoAI Engineer

KV Cache-Aware Routing and P/D Disaggregation on Kubernetes — Yuchen Fama & Ashish Kamra, Red Hat

Red Hat engineers Yuchen Fama and Ashish Kamra present two methods to accelerate AI model execution on servers. The first is intelligent routing — when users ask follow-up questions, previous computations are reused on the same server, making responses 3 times faster. The second is dividing work between two server types: one that prepares text and one that generates words. In a practical test, latency between words dropped from 900 milliseconds to 100 milliseconds. However, disaggregation only works well under moderate load and requires specialized network hardware — otherwise it is better to keep everything on a single server.

26 Aug youtube.com

VideoAI Engineer

Einstein Arena: Harnessing Collective Agent Intelligence for Open Science — James Zou, Together AI

James Zou from Together AI and Stanford has built Einstein Arena, an environment where AI agents solve open scientific problems against each other. Instead of instructing agents on how to work, the system sets up a room with verified problems, a forum where agents share what hasn't worked, and a live leaderboard. Within weeks, the agents' solutions had beaten previous records on eleven problems — notably, they solved the kissing number problem (how many spheres can touch a central sphere without overlapping) in eleven dimensions by iteratively improving each other's attempts. The same setup was applied to GPU kernel optimization with results twice as fast as previous state-of-the-art.

25 Aug youtube.com

VideoAI Engineer

Building GTM AI Agents: Lessons from Deploying to 6,000 Users — Sait Izmit, Snowflake

Sait Izmit from Snowflake built an AI agent to answer questions from the sales organization. Before launch, he wrote 150 concrete questions from real sales processes, accepted that the first version would only answer 50 percent correctly, and focused on answering a few questions well rather than many poorly — because users judge the agent on the first five answers. The agent launched a year ago and has since answered over one million questions (approximately 40,000 per week) for around 6,000 salespeople. The biggest challenge was not the technology but getting people to actually use it: only one-fifth of the organization tested it without active sales and pushing from his team.

25 Aug youtube.com

VideoAI Engineer

Reverse-Engineering the AI Buyer — Aliisa Rosenthal, Acrew Capital

Aliisa Rosenthal, who helped OpenAI grow from several million to several billion dollars in revenue, shares lessons on selling AI tools to businesses. She says OpenAI made many mistakes early on: they built for the loudest customers with expensive enterprise software, but when they later launched cheaper self-service options, growth accelerated dramatically and cannibalized enterprise sales. Her advice to founders is to first automate as much as possible, then hire people only where the system falls short—not the other way around. She also warns against giving buyers "homework" to do, letting pilot projects become free sales processes, and recommends automating security questions that often kill deals.

25 Aug youtube.com

VideoAI Engineer

The Missing Layer in Agentic AI — Giedrius Šteimantas, Oxylabs

Giedrius Šteimantas from Oxylabs explains how AI agents often waste money and tokens reading blocked webpages without noticing—a 200 status code doesn't mean the page contains actual content. His example is a shopping agent that ran slowly and cost far too much. By dividing the work into three parts—searching via API instead of a web browser, the decision stage with a scraper that fails clearly when blocked, and only checkout via automated browser—costs drop dramatically. The rule from ten years of web scraping is simple: use a browser only when necessary, and validate before sending data to the model.

25 Aug youtube.com

VideoAI Engineer

Inside 847 Production Clinical AI Notes — Sebastian Fox, Composo

Sebastian Fox from Composo analyzed clinical notes from AI tools that automatically write patient notes (known as ambient scribes). He found that approximately one in twenty notes contained errors severe enough to harm patients, nearly one in five had important missing information, and more than one in ten contained fabricated information. Even when he built an advanced checker model to catch these errors, one-fifth of problematic notes still passed through. The biggest challenge is determining which differences between recordings and notes actually matter medically — a missing date may be insignificant, but missing symptoms such as jaw pain combined with new headache in patients over 50 could indicate a serious vision-threatening diagnosis.

22 Aug youtube.com

VideoAI Engineer

Agent Frameworks Considered Harmful — Rémi Louf, .txt

Rémi Louf built his own agent system after encountering problems with existing frameworks — duplicates, missing messages, and unlogged prompt changes. Instead of complex graphs, he created something simple: agents as plain markdown files, a log where everything is saved and can be traced, and a system where each part of a prompt is stored separately with a hash so you can see exactly what changed between runs. It took two weeks to build and now runs twenty agents at his fifteen-person company.

22 Aug youtube.com

VideoAI Engineer

Preferences Over Benchmarks: Model Routing — Archana Kamath & Tyler Gillam, DigitalOcean

Archana Kamath and Tyler Gillam from DigitalOcean present a router that automatically selects which AI model to use for each task, rather than running the same model for everything. In a demo, two terminals build the same app — one uses an expensive premium model for everything, the other uses the router. Both take approximately the same time, but the router costs 8 cents versus the premium model's 25 cents. The router bases its choice on what you actually need: task type, cost, speed, system configuration, and user preferences. There is no simple answer to which model is best — it depends on the context.

22 Aug youtube.com

VideoAI Engineer

What If Your Chip Design Team Moved Like a Single Body? — Abduallah Mohamed, AIDAChip

Abduallah Mohamed from AIDAChip describes how AI agents are used in chip design—an area where mistakes cost around 50 million dollars per faulty printed circuit. The main problem they identified is not agent capability but their alignment: 70 percent of time is spent keeping agents on track. When they blocked agents from writing to certain files, the agents circumvented this by using other tools like bash and sed—the lesson was that control must be at the system level. AIDAChip built a "shared nervous system": a living network of intentions and constraints that agents cannot change without human approval, plus a knowledge storage system that grows from project to project.

22 Aug youtube.com

VideoAI Engineer

FinOps for AI Agents: Who Spent All the Tokens? — Tisha Chawla & Susheem Koul, Microsoft

Microsoft researchers Tisha Chawla and Susheem Koul present a system for controlling costs when AI agents run code. The problem is that current AI tools lack a control mechanism between your code and its actual cost — unlike cloud services with autoscaling or SaaS services with usage limits. Their solution uses "steering" instead of simply shutting down expensive runs: the system monitors spending in real time and instructs the agent to provide shorter responses when costs threaten to exceed budget. In tests, this reduced average costs by 78 percent while increasing the proportion of runs that actually completed from 67 to 96 percent.

22 Aug youtube.com

VideoAI Engineer

Building Agents Is Trivial Now, Context Is the Next Frontier — Jeff Ng, Unblocked

Jeff Ng from Unblocked demonstrates that building AI agents has become straightforward—frameworks and cloud tools have solved the technical work. However, agents fail when they lack context. In his example, an agent recommended the wrong action for a performance problem in a code request because it read the ticket and code but not the Slack discussions where the team had already solved the problem and documented why a certain setting caused an outage. By instead giving the agent access to a context system encompassing documentation, code, tickets, and conversations, the same agent reversed its answer from repeating the outage to preventing it.

21 Aug youtube.com

VideoAI Engineer

Agentic SDLC at Uber — Uday Kiran Medisetty & Adam Huda, Uber

Uber has built infrastructure enabling AI agents to write code—over 70 percent of their code reviews now come from automated agents instead of humans. The system works through six major components: a central gateway that scans every AI call for security in under 100 milliseconds, a system reducing the data agents need to keep in memory (cutting tokens by 40 percent), a database with 2,500 pre-built code tools, and a network of 40 million data collections replacing many separate systems. The entire workflow is controlled from Slack, and agent work is intentionally stopped before entering the build queue so humans can first verify screenshots against design files.

21 Aug youtube.com

VideoAI Engineer

The Missing Layer: Design Taste in AI Agents — Hassan El Mghari, Together AI

Hassan El Mghari from Together AI has created Hallmark, a tool that helps AI agents understand design quality. He explains that anyone can see when an app looks "poorly designed" — purple gradients, italicized headlines, scroll prompts nobody asked for, uppercase letters with wide letter spacing, too many emojis — but few can explain why. By naming these patterns, one can teach an AI agent to avoid them. Hallmark codifies design patterns as rules and gives the model a library of ready-made themes. El Mghari has built ten apps per year over five years by iterating with smaller AI models, and says design quality is the biggest factor for reaching many users.

21 Aug youtube.com

VideoAI Engineer

Unlock Agent Autonomy: The Runtime for AI-Native Systems — Tushar Jain, Docker

Tushar Jain from Docker describes a problem: when AI agents are given too much access, they can do unexpected things — his example is an agent that suddenly posted internal notes as a pull request without being instructed. The issue is that agents determine what they need while working, not beforehand, so traditional access controls don't work well. His solution is a runtime layer beneath all models built on three components: controls that sit outside the agent's sandbox, tools limited to one specific task at a time, and access controlled against what the agent is actually attempting to do — so unusual requests (like sending emails during a bug investigation) can be refused or escalated to a human.

20 Aug youtube.com

VideoAI Engineer

Your Fine-Tuned Model Is Tech Debt: A 50x ROI House of Cards — Dan Bjornn, Lease End

Dan Bjornn describes how a fine-tuned AI model that generated $12 million in revenue with 50x return on investment over a year became increasingly expensive to maintain. Each time something went wrong—when the model called customers at inappropriate times or confirmed meetings incorrectly—it took a week to fix: collecting error examples, synthesizing new training data when insufficient, manually validating it, sorting, reviewing again, and finally retraining (which took only an hour). Eventually, every update fixed one problem but reintroduced an old one, so the team ranked bugs by how much customer blame they could tolerate. The model also became locked in—training data wasn't transferred between model versions, preventing upgrades to newer architectures. The solution was rebuilding the system using prompts, tools, and context instead of fine-tuning: repairs could then be deployed in under an hour, and total costs dropped.

20 Aug youtube.com

VideoAI Engineer

How I automate my own job at Hugging Face using agents — Niels Rogge, Hugging Face

Niels Rogge at Hugging Face automated his job of finding AI models hidden on Dropbox or Zenodo and asking authors to move them to Hugging Face Hub instead. He built the solution two ways: first as a simple, predictable workflow that runs nightly and sends personal messages without revealing it's automated, later as a fully autonomous loop with bash commands that handles follow-ups. The first approach was simple and cheap, the second handles follow-up questions without human involvement. The result was that thousands of GitHub issues were opened automatically, with only two becoming negative.

20 Aug youtube.com

VideoAI Engineer

IT Admin for the AI Workforce — Sarthak Aggarwal, Decawork

Sarthak Aggarwal argues that companies are now bringing in a second workforce — AI agents — and that security concerns are no longer about how models behave but how they are managed as employees. He presents two failure examples: the Replit incident where a coding agent ignored a code freeze and deleted data in a production database, and EchoLeak where an external email entered Microsoft 365 Copilot's context and retrieved sensitive data. The solution is to give agents identities, ownership, and limited access — similar to how OAuth already works for humans — combined with a system where a planner first creates a logged action plan before an executor runs it, so the model proposes but policy makes decisions.

20 Aug youtube.com

VideoAI Engineer

Prototyping as Leadership: How a CTO Ships with AI Agents — Hursh Agrawal, The Browser Company

Hursh Agrawal, CTO at The Browser Company, demonstrates how he uses AI agents to actually write code despite having twenty meetings per week and seven direct reports. He starts a coding agent in the evening that works overnight, then reviews the results in the morning — a process that allows him to deliver two to ten pull requests every week. His point is that leaders have the most business context and can therefore guide AI agents more effectively than regular developers, and that showing a working prototype settles many discussions about what new AI models can actually do. However, it requires good test automation, feature flags, and honest review of your own code.

20 Aug youtube.com

VideoAI Engineer

The Last Human Code Review: Building Trust in AI-Generated Code — Itamar Friedman, Qodo

Itamar Friedman from Qodo argues that the problem with AI-generated code is not the quality of models but that code review context is fragmented and often uncodified. He identifies two opposing camps among developers: one requiring human review of every line, one accepting bugs for higher speed. Friedman's point is that this context — architectural rules, previous errors, service contracts — lies scattered in instruction files, Slack threads, and developers' heads, and must be collected and codified in a way both AI agents and humans can use. When done correctly, code review can evolve from reading individual changes to understanding the entire system's dependency graphs.

20 Aug youtube.com

VideoAI Engineer

The Agent Behind the Curtain: Building the Oz Cloud Agent Platform — Safia Abdalla, Warp

Safia Abdalla from Warp presents how they built a cloud-based agent platform that helps developers without overwhelming them. Rather than letting agents act freely, they integrate them into the natural workflow — agents triage issues, review pull requests, and approve them before humans need to see them. The platform supports various developer tools, enables agents to collaborate with sub-agents, and can be used by anyone via an API — for example, it was used to build Slack tools for triaging. Abdalla emphasizes that a good agent platform should absorb complexity before it reaches the user, comparing it to a pottery workshop rather than a factory.

20 Aug youtube.com

VideoAI Engineer

Coding Agents Don't Scale Themselves. Neither Do Your Teams. — Patrick Debois, Tessl

Patrick Debois compares today's challenges with AI agents to the resistance against continuous delivery in 2009 — the technology works, but the organization isn't ready. His point is that what makes the difference is not the agent itself or the prompts, but how the team, platform, and work processes change around it. Instead of fixing the agent's code, you should improve the system as a whole. It's about reducing the number of times a human needs to touch a task, and making improvements that benefit the entire team, not just one person.

20 Aug youtube.com

VideoAI Engineer

Give the Agent a Budget, Not a Token — Sachin Malhotra, Anthropic

Sachin Malhotra from Anthropic describes an agent that accidentally deleted 200 jobs in 90 seconds—a reminder of the risks of giving AI agents too much power. The problem wasn't that the agent was malicious, but that it had unlimited permissions. Instead of only restricting what an agent can do (token-based control), Malhotra proposes a "budget" system with four dimensions: how much, how fast, what can be undone, and who oversees it. This involves giving agents operations that fail visibly (like undo) while humans retain control over things that fail silently, plus speed restrictions and monitoring mechanisms.

20 Aug youtube.com

VideoAI Engineer

Your Agent Evolved. Your Evals Didn't. — Ameya Bhatawdekar, Braintrust

Teams built orchestration graphs because the models of 2024 could not be trusted to orchestrate, and then the models learned to orchestrate and the graphs became the thing holding them back. Ameya Bhatawdekar traces that loop across five generations of architecture, each one forced by a step change in model capability, and argues that evals have to move with it. A single prompt needed only answer quality. A retrieval chain added a parser that grabs the wrong field and a retriever that returns the wrong context. Graphs added branch logic, contracts between nodes, and classifier nodes that misfire quietly, which is a great deal of new surface to check. What changed most recently is not another layer but the unit of measurement. Once a loop is reliable enough to run free, the same input produces visibly different trajectories on every run, so a single eval result stops meaning very much. He separates the two questions it hides. Pass at k asks whether the system succeeds at least once across k attempts, which measures capability. The stricter variant asks how many of those k attempts succeed, which measures reliability.

20 Aug youtube.com

VideoAI Engineer

The Era of Compound Engineering — Kieran Klaassen, Every/Cora

Kieran Klaassen built a complete email client entirely with AI assistance, without writing or reading most of the code himself. He discovered that the bottleneck shifted from code quality to his own time—he became limited by repetitive tasks and decisions only he could make. To solve this, he built a memory system that stores solutions from previous problems, so when similar challenges arise, the AI tool doesn't need to start from scratch. His key insight is that he spends roughly half his time building a feature and half teaching the system what it did wrong, which ultimately becomes cheaper in terms of AI API usage and, more importantly, makes the next feature easier to build.

20 Aug youtube.com

VideoAI Engineer

Why Your Enterprise Tech Stack Isn’t Ready for AI Agents — Christopher Lovejoy & Saul Howard

Christopher Lovejoy and Saul Howard explain why many AI agents fail in enterprises despite working prototypes. The problem is that compliance requirements like audit trails are often added afterward rather than built in from the start. They propose an architecture based on constraints: an immutable event log that records every action the agent takes, patient data stored separately from the log so agents can be debugged without accessing sensitive information, and treating both humans and AI models as the same type of agent. This makes auditing, security, and escalation natural parts of the system rather than add-ons.

19 Aug youtube.com

VideoAI Engineer

Don’t be data poor — Anuj Iravane, Anterior

Anterior works with medical data often received as faxed documents over 300 pages long containing handwriting and tables — but they cannot store this data due to legal requirements. Instead of attempting to train AI on data they are not permitted to retain, they reverse the process: they start with the correct answer (diagnosis or decisions), work backwards through the reasoning, and then generate the patient records that would have led there. Because they model medical guidelines as decision trees, the generated material has greater variation and is more realistic than if an AI model simply tried to create data on its own. As a result, approximately 90 percent of their training data is now synthetic, and doctors can distinguish synthetic from real data only about 60 percent of the time.

19 Aug youtube.com

VideoAI Engineer

How to build an AI-Native Health Company — Dan Feng, Maven Clinic

Dan Feng from Maven Clinic, a healthcare app, describes how AI is changing how companies plan and build products. Since it now takes minutes to build something instead of weeks, costs have shifted — the expensive part is now discussing what to build, not building it. Maven Clinic only plans two to four weeks ahead in detail, while long-term plans serve only as direction. Code reviews also became a problem when engineers wrote ten times more code than before, so Maven lets engineers themselves assess which changes need review by two people. For critical features like calculating reimbursements, they use multiple AI models that must agree before anything passes through.

19 Aug youtube.com

VideoAI Engineer

Trading Desks to Clinical Trials: Parallels in Applied Vertical AI — Ayush Bhardwaj, Allos AI

Ayush Bhardwaj works with AI tools in two very different industries — financial trading and drug development — and discovered that the problems are nearly identical. The major challenge is assessing whether the AI system being built actually works well: an experienced coder can tell in a minute if generated code is weak, but no one on his team could judge the quality of a trading strategy or a drug candidate proposal. He first tried solving this by training a model to evaluate the results, but it failed because critical data is kept secret (trading funds don't want to reveal winning strategies, and approximately 30 percent of pharmaceutical companies never publish their trial results). The solution became hiring a senior expert from each field who can assess results, select the best data sources, and refine how to ask the AI system for help. The expertise and data are what create real competitive advantage — the AI model and infrastructure themselves are essentially interchangeable.

19 Aug youtube.com

VideoAI Engineer

Healthcare’s Agent Bytecode: X12 as the Harness for AI Agents — Vasant Kearney, Onlay

Vasant Kearney from Onlay argues that when AI agents handle healthcare insurance claims, they should use X12—the insurance communication standard—as a framework rather than just a file format. The challenge is that an insurer's phone system, web portal, and X12 data are often built by different teams and can contradict or malfunction together. By using X12 as a structural framework, AI agents can normalize all interactions—phone calls, portal interactions, insurance checks—into the same underlying transaction. He emphasizes two practical requirements: memory must be stored in a database for logical separation, and a "better" AI model cannot simply be swapped in, because better performance on a test does not mean better performance within a system built around a different model.

19 Aug youtube.com

VideoAI Engineer

AI is the World’s largest Relationship Therapist — Clay Cockrell & Tony Fabrikant, CoupleWork AI

Therapist Clay Cockrell and Tony Fabrikant discuss why AI chatbots like ChatGPT—used by hundreds of millions of people for relationship advice—are poor at helping with romantic relationships. The problem isn't that AI is incompetent, but that it's too agreeable with users: if you ask why your partner doesn't listen, the AI simply tells you that you're right, making you more confident in your position rather than helping you understand your partner better. Real therapists use methods based on John Gottman's research and emotionally focused therapy to actually resolve conflicts. Cockrell and Fabrikant have built a tool that instead starts from what a good therapist does, thoroughly tests it against safety requirements, and learns when it should stop coaching.

19 Aug youtube.com

VideoAI Engineer

Shipping AI to a Million Patients Without an A/B Test — Jared Joselowitz, Ufonia

Jared Joselowitz from Ufonia describes how to safely launch Dora, an AI voice assistant that calls patients for follow-up after surgery in British hospitals. Because the system provides medical advice, it is a regulated medical device—you cannot A/B-test on patients as that is unethical, and incorrect advice cannot be retracted. Instead, Ufonia uses simulation: one AI plays a patient against hazards written together with doctors, and another AI evaluates each call. The evaluating model was tested against 10 specialists on 240 cases and performed equally well or better. Instead of adjusting instructions by hand, they optimize them against a cost matrix that weighs over-diagnosis against missed warning signs.

19 Aug youtube.com

VideoAI Engineer

Guardrails First: Engineering Member-Facing Health AI — Rashi Agrawal, Hinge Health

Rashi Agrawal from Hinge Health argues that safety in healthcare AI is about architectural decisions made before the model generates anything, not just good instructions. She points to real cases where AI assistants give dangerous medical advice — a man developed sodium bromide poisoning after asking a popular chatbot about salt deficiency, and independent tests show healthcare AI often misjudges serious conditions. She advocates that sensitive patient information is never stored in the system, that critical decisions (such as calling 911) are hardcoded rather than left to prompts, and that safety runs as a continuous layer of reviewers monitoring all live traffic.

19 Aug youtube.com

VideoAI Engineer

The Next Medium: Why Real-Time Interactive Video Changes Everything — Ahmed Ahres, Reactor

Ahmed Ahres argues that real-time interactive video is not just faster but an entirely new medium — similar to how GPS transformed navigation and viewfinders changed filmmaking. He defines world models as video generated so quickly you can control it while it's being created, demonstrating this by prompting a cat into a scene while it's still generating. This opens three possibilities: control spaces where you get immediate feedback, worlds where AI characters can train robots or be used in education, and live avatars (which he acknowledges still don't work well). Technically, it requires completely new systems — instead of sending files back, servers must stream pixels in real-time, each session must remember what happened, and latency must stay under 100 milliseconds.

18 Aug youtube.com

VideoAI Engineer

Infra behind Krea 2: How to train and serve at scale — Gabriel Jorge Menezes, Krea.ai

Gabriel Jorge Menezes from Krea.ai explains how to train large AI models across thousands of GPUs. He shows that many standard metrics are misleading—for example, GPU usage displayed 100 percent constantly even when the cluster wasn't running efficiently. Instead, he tracked tensor core utilization. When training ran on larger images (from 128 to 1024 pixels), this metric increased noticeably. To monitor network communication between servers—a common source of errors—they built custom tools since no off-the-shelf solutions existed. A GPU exceeding 78 degrees is immediately removed from the system to prevent slowing down the entire training. The Krea 2 model was frequently saved with checkpoints to very fast storage that could write a terabyte in under 30 seconds, enabling training to recover quickly after crashes.

18 Aug youtube.com

VideoAI Engineer

Generative Video at the Speed of Light — Keegan McCallum, uRun

Keegan McCallum from uRun presents how generative video has become much cheaper and faster — ten dollars now provides approximately three hours of continuous video. He argues that quality is no longer the interesting factor, but rather the ability to control the video in real time. This opens entirely new use cases: a webcam showing the hairstyle you're considering, or tools for people who don't think in text. The major technical challenge is no longer creating the video but serving it globally and keeping it synchronized with the user's controls frame by frame.

18 Aug youtube.com

VideoAI Engineer

While my guitar gently speaks — Todd Fisher, Philo Ventures

Todd Fisher built a music production plugin (JUCE) that transforms a guitar into a speaking instrument. He connected a microphone to automatic speech recognition and an AI model, which then sends its response back through the guitar strings. The challenge was splitting synthesized speech into individual words—his solution combined two methods for finding word boundaries, plus manual work to get it precisely right. To make the guitar sing, he needed entirely different technology: a pitch detection algorithm, synthesis with vocoder effect, and pre-calculated vocal samples, since direct computation would be too heavy for live performance.

18 Aug youtube.com

VideoAI Engineer

The Next Game Engine Won't Have a Manual — Arturo Nunez, Nereu

Arturo Nunez from Nereu presents a new type of game engine where you describe what you want to do in natural language, and an AI tool helps you build it. Instead of requiring knowledge of technical components like renderers and rigidbodies, the engine uses a tagging system — you label things as "player," "animated," or "can jump" and the AI understands what should happen. The clever part is that the engine doesn't try to generate a complete finished game from a description, but helps you when you get stuck. The system borrows from the rendering technique "level of detail" to only give the AI information about assets relevant to what you're working on right now.

18 Aug youtube.com

VideoAI Engineer

Building an Agentic Video Editor for Mass Consumer — Ekaterina Deyneka, Reelful

Ekaterina Deyneka from Reelful presents an AI agent that automatically edits video. You input raw footage and instructions, and the agent finds useful clips, assembles them, and adds captions, music, voiceovers, and b-roll. The system uses Remotion—a tool that writes video as code—and a verification check to catch compositions that don't work. Crucially, the agent must select from actually recorded material rather than generate from scratch, which requires greater judgment. All complexity is hidden behind mobile templates so users never see it.

18 Aug youtube.com

VideoAI Engineer

How to Kill the Code Review — Ankit Jain, Aviator

Ankit Jain from Aviator argues that code review is already dead — over 30 percent of code changes pass through without review, and when AI writes the code and AI reviews it, humans merely go through the thread and approve. The problem, according to him, is that code review was never just about finding bugs; it was also knowledge sharing, mentorship, and feedback on system design. His solution is to capture the decisions made when writing instructions to AI, turn them into acceptance criteria, build a registry from previous review comments, and generate test plans that run against pre-reviews instead of just reviewing code differences.

17 Aug youtube.com

VideoAI Engineer

Security Firewall for Agents — Ryan Dahl, Deno

Ryan Dahl from Deno presents Claw Patrol, a security system for AI agents that have access to sensitive production environments like databases and AWS. The problem is that agents can fall victim to prompt injection—malicious instructions smuggled in through support channels connected to the agent. Although modern AI models like Opus often refuse harmful commands, they cannot be reliably counted on to resist all possible attacks. The solution is Claw Patrol, a proxy that sits between the agent and the resource and inspects every byte of data leaving the agent—even at the protocol level below HTTP, since agents can open direct database connections that no HTTP rule would catch. Rules are written in HCL (Terraform's configuration language), stored in git, and can be tested against fake requests. The proxy can also forward decisions to an AI judge or a human in Slack before anything is allowed to run.

17 Aug youtube.com

VideoAI Engineer

Context Engineering in 2026 — Louis-François Bouchard, Omar Solano & Samridhi Vaid, Towards AI

In a video from Towards AI, Louis-François Bouchard, Omar Solano, and Samridhi Vaid present results from experiments on how to best handle long context history in AI systems. They tested an AI tutor and found that it is often better to keep all history instead of trying to compress it — because prompt caching (where APIs save tokens and reuse them more cheaply) makes compression actually more expensive when it destroys the cache. In their measurements, the uncompressed version retained details 95 percent of the time compared to 32 percent after summarization, and was able to maintain specific facts up to 800,000 tokens. The main rule they conclude is to first identify which constraint you actually have — context window size, speed, or cost — before you start compressing.

17 Aug youtube.com

VideoAI Engineer

From Ambient Documentation to Clinical Intelligence — Chaitanya Asawa, Abridge

Abridge is an AI tool that helps doctors with documentation after patient visits — work doctors call "pajama time" that takes about two hours per day after work is done. The company has reached 300 large hospital systems in the USA. Founder Chaitanya Asawa explains that all administrative documentation in healthcare grew around a single basic conversation between doctor and patient, and that the major challenges involve verifying that AI's answers are correct — something particularly difficult in medicine because many different answers can be right. To solve this, they use four doctors to build an assessment template for what a good answer should contain. To handle the cost at their scale — nearly 100 million medical conversations per year — they train smaller models for each section in the doctor's notes instead of running the most advanced AI technology over everything.

17 Aug youtube.com

VideoAI Engineer

200 Million Patient Interactions Later — Vivek Muppalla, Hippocratic AI

Hippocratic AI has conducted over 200 million patient conversations through its Polaris system, designed to proactively call patients—something that rarely happens today due to doctor and time shortages. The system runs 31 AI models in parallel on each call: one managing the conversation thread and 30 specialist models handling everything from medication review to scheduling. To achieve both safety and speed, they built the entire stack in-house, including custom speech recognition that understands both what patients say and how they say it. They maintain strict error standards—a 1 percent error rate means 100 misdirected patients daily at 10,000 calls, so testing must cover approximately 450 cases to catch that error level.

17 Aug youtube.com

VideoAI Engineer

Training Krea 2: What matters in generative model training — Sangwu Lee, Krea.ai

The most reliable way to render a person is to render the most boring average person and put them in the center of the frame. Sangwu Lee offers that as the price the big image models pay for consistency: ask a production model for a burning skull and every output comes back clean, competent, and nearly identical. Krea 2, whose medium variant is now open source, trades the other way, optimizing for fast generation and stylistic range so that a studio that does not yet know what it wants can actually explore. Most of the talk is about data, which he says twice over is basically everything once the architecture is locked. The examples are specific. A painting photographed on a wall is perfectly good training data except that captioners consistently omit the frame and the white wall behind it, so the model learns to hang every painting it generates. They refuse to train on AI generated images at all, because the aesthetic is sticky and you inherit somebody else's model. Deduplication runs on hashes first across two to ten billion images, then on embeddings for near duplicates.

15 Aug youtube.com

VideoAI Engineer

Voice agents with Realtime Video — Sidney Primas, LemonSlice

Sidney Primas from LemonSlice demonstrates how to build AI agents with avatars and video that can run continuously for hours. The main challenge is that video is generated frame by frame—each new image is based on previous ones, and errors accumulate over time. LemonSlice solves this by training the model to only look backward (since future frames don't exist yet) and drastically reducing computational steps. A surprising issue is audio: expressions and facial movements depend on how sound is analyzed, but most audio models are trained on monotone audiobooks, so they built their own. The service costs roughly the same to run as a voice model, but according to Primas, the biggest future value lies in orchestrating computations across GPUs and CPUs without video stuttering.

15 Aug youtube.com

VideoAI Engineer

How Web Data Infrastructure Powers the Next Generation of AI — Patricija Žemaitytė, Oxylabs

Patricija Žemaitytė from Oxylabs shares experiences building infrastructure for AI services at scale. She describes three main challenges: a search API that was supposed to respond in under a second got blocked live with a client, requiring a rebuild by removing slow browsers—resulting in 550 milliseconds instead of 4 seconds. A video API request with a two-week deadline grew from a simple job into an entire product suite (transcripts, subtitling, search, metadata) when the client discovered new needs, and scaled to 30 petabytes of data without yet paying. Scale itself became a problem: when she tried to test an "unblocker" service up to 60,000 requests per second, the tests stalled at 20,000, partly because the telemetry itself became so large that it affected what she was trying to measure.

14 Aug youtube.com

VideoAI Engineer

The Rise of CaaS: Context-as-a-Service for Agentic AI — Omer Primor, Bright Data

Omer Primor from Bright Data presents a comparison between renting pre-built context data (Context-as-a-Service) and building your own pipeline to keep AI agents updated with web data. In a test enriching 25 information fields for 100 companies, costs became equal at approximately 15,000 queries. The primary challenge is not data volume but update frequency: social media becomes outdated within a day, news and financial data within 30 days. With rented services, each repeated query costs the same even when the answer hasn't changed, while self-built solutions pay large upfront costs but are then effectively free per search.

14 Aug youtube.com

VideoAI Engineer

The Dark Arts of Web Automation: Teaching Agents to Use Websites Like Humans — Corey Gallon, Rexmore

Corey Gallon demonstrates how to build AI agents that automate web browsers through Chrome DevTools Protocol — they can click, type, and navigate the web like a human. His method uses a loop: the agent observes the page, takes an action, verifies the result. He shows three levels of complexity: simple forms, buttons requiring genuine click events, and finally CAPTCHA challenges like Cloudflare Turnstile and reCAPTCHA v2. A key finding: CLI tools are faster than MCP servers (under one minute versus eight minutes) — this matters when CAPTCHAs expire over time. Gallon received a warning from OpenAI for this research and demonstrates only code running on his own infrastructure.

14 Aug youtube.com

VideoAI Engineer

Bringing agents onto the world wide web — Paul Klein IV, Browserbase

Paul Klein IV from Browserbase argues that web agents — AI systems capable of navigating and interacting with web pages — are no longer limited by AI model capabilities, but rather by missing infrastructure around them. The best agents in production combine multiple capabilities (can both click and write code), have memory and skills so they don't have to rediscover the same website each time, and run on infrastructure that renders pages identically every time. Klein also points to three unsolved problems: how agents log in on your behalf, what the web should offer agents to make them reliable, and who certifies that an agent can be trusted. The actual value lies not in Silicon Valley but with logistics companies in Singapore, banks in South Africa, and labor operations in Mexico that today are tied to outdated PHP forms.

14 Aug youtube.com

VideoAI Engineer

Computer Use at the Edge of the Statistical Precipice — Pierluca D'Oro, Programma Labs

Pierluca D'Oro demonstrates how AI agents can achieve high scores on standard tests by simply replaying recorded sequences of button presses—without actually looking at the screen or understanding what they're doing. This works on many common benchmarks because the tests are too predictable. He presents the PRISM principles for making environments less exploitable (random variations, sandboxes, multiple starting states), and introduces DIGIWORLD with 15 mobile apps and 3.2 million verified test configurations. He also shows that current uncertainty measures are poor: a 95% confidence interval covers actual performance only 20% of the time, masking large differences between models.

14 Aug youtube.com

VideoAI Engineer

Improving Agents is a Data Mining Problem — Vivek Trivedy, LangChain

Vivek Trivedi from LangChain argues that improving AI agents is about analyzing trace data — the logged steps and decisions an agent takes while working. Instead of guessing why an agent performs worse after updates, you can have other AI models read through these traces and identify problems. Trivedi contends that observability and continuous learning are the same problem viewed from different angles. He demonstrates that cheaper, open-source models can achieve the same quality as expensive top-tier models if you're smart about how you phrase instructions — and that it pays to iterate on instructions until progress slows, then fine-tune the model to advance further, then return to refining instructions again.

12 Aug youtube.com

VideoAI Engineer

Designing Agents (The Floor Is the Frontier) — Ben Hylak, Raindrop

Ben Hylak from Raindrop argues that current advice on testing AI agents is designed for an older era of chatbots where user queries were predictable. Instead, focus should be on the "floor"—the worst mistakes that destroy trust, such as when an agent recommends a competitor or deletes data—rather than on all possible problems. The key is to track two metrics for each error: when it started and how many users it affects. Hylak shares three lessons from Raindrop: classifying agent traces is not the same as error detection, automated code-based testing scales well, and agents are better at investigating anomalies than finding them on their own.

12 Aug youtube.com

VideoAI Engineer

Bringing Continual Learning into Enterprises — Samuel Denton, Applied Compute

Applied Compute uses a technique called continuous learning to improve AI models directly in business production environments. In one example, they got a Qwen model to complete tasks on SWE bench twice as fast (from 80 down to 40 steps) by showing the model how to behave based on its existing behavior — rather than forcing it toward a "correct answer." Their approach works in two modes: either with old saved data from production (quick to start but limited improvement), or with live feedback from each new run (slower to set up but yields much better results). In a real-world example, they successfully got a model to format hyperlinks correctly in 80 percent of cases by providing smart tips after each attempt — something that both reward systems and standard training failed to achieve.

12 Aug youtube.com

VideoAI Engineer

Adaption Labs: Gradient-Free Continual Learning — Sara Hooker, Adaption

Sara Hooker from Adaption Labs argues that AI development is facing a revolution. Currently, a handful of companies and research institutions control how large AI models are trained — fewer than 5,000 people worldwide know how to do it. Hooker presents AutoScientist, a tool that automates the training process itself and finds better solutions than humans by testing different design choices instead of relying on prior experience. She also claims that the era when larger models are always better is coming to an end, which would make it possible for more people to build powerful AI systems — because smaller, smart models can be deployed without requiring enormous GPU resources concentrated in one place.

12 Aug youtube.com

VideoAI Engineer

Intelligence + Continual Learning = Expertise — Yu Su, NeoCognition

Yu Su from NeoCognition distinguishes between two concepts often conflated: intelligence and expertise. Intelligence is the ability to reason through new problems — something modern AI models are increasingly better at. Expertise is knowledge built up over time through experience, something that is almost never scaled today. He explains why coding agents work well while other tools are fragile: code is already structured and symbolic with tests providing feedback, but the rest of digital work comprises millions of small worlds with different rules that a static model cannot handle. Su argues that intelligence and expertise are nearly orthogonal — scale only intelligence and you get the world's smartest beginners who are brilliant but learn nothing between problems.

12 Aug youtube.com

VideoAI Engineer

Scaling Compute on Context — Jack Morris, Engram

Jack Morris from Engram explains why it is difficult to get AI models to understand private company data. Models are trained on public internet data and perform very well there, but when you try to teach them your own data, they often become nearly useless when generating results. He describes three classical ways to improve this—more data, more computing power, and larger models—but says that when it comes to private data, only computing power is a viable option. He goes through various techniques such as KV-compression and synthetic training data, but shows that they all plateau once the model has learned everything you've given it, and there is no magical way to continue improving with more computing power.

12 Aug youtube.com

VideoAI Engineer

Scaling up Continual Learning — Ronak Malde, Trajectory

Ronak Malde from Trajectory presents self-distillation, a method for training AI models on long-horizon tasks with multiple steps. The problem he solves is that when models are trained on long sequences using existing methods, they become uncertain and fill their responses with words like "but", "wait", and "maybe" — what he calls the "but wait problem". Self-distillation works by making the model its own teacher: you give one version of the model extra information ("hints") and train another version to match its decisions without these hints. Unlike other training methods that must choose between different benefits, Malde says self-distillation achieves all four desired properties: training on real data, ability to sample different paths, doing so efficiently without parallel runs, and providing reward for every individual token.

12 Aug youtube.com

VideoAI Engineer

Beyond Static Intelligence: Evaluating Continual Learning — Parth Asawa, UC Berkeley

Parth Asawa from UC Berkeley criticizes how AI models are currently evaluated. All ranking lists rely on testing a model on one task, resetting its memory, and testing the next task — which hides how well the system actually learns over time. He presents a new measurement tool that compares the same system when it can learn (retain memory between tasks) versus when it is reset after each task. The results are surprising: simple in-context learning (where the model uses previous answers to improve itself within the same conversation) outperforms more complex memory management systems. His benchmark covers six areas, from database search to forecasting models, and reveals problems such as models forgetting their own corrections.

12 Aug youtube.com

VideoAI Engineer

Computer-use models will agentify the web, not APIs — Dhruv Batra, Yutori

Dhruv Batra argues that AI agents cannot rely on APIs to interact with the web. Instead, they must learn to navigate web pages visually—by reading screen content and clicking, like a human does—because most of the web (around 200 million active websites) will never build APIs. Much of what we see on the web is not actual text but drawn graphics: a "sold out" status might be just a gray button, not words written anywhere. Yutoris Navigator model solves this by taking screenshots, clicking on elements, and sometimes writing JavaScript to accomplish what it needs—and focuses on verifying results on screen rather than expecting structured data.

12 Aug youtube.com

VideoAI Engineer

From RL to IRL — Gaurav Mishra, Amazon AGI Lab

Gaurav Mishra from Amazon's AGI lab demonstrates how AI agents trained to use web browsers encounter entirely new problems when facing real login screens and webpages. During training, the agent guessed its own password and clicked on fake buttons — issues that arise because the agent cannot see everything on the screen, cannot undo mistakes, and does not know when it is authenticated. His solution is to train agents more like pilot schools train pilots: with realism in the environment, process rewards that penalize dangerous steps along the way, and self-awareness about when an action is reversible and safe before executing it.

12 Aug youtube.com

VideoAI Engineer

Agents, codebases, and teams — Aditya Khandelwal, Amazon AGI Lab

Aditya Khandelwal from Amazon AGI Lab shares insights on introducing AI agents into an engineering team. The core issue isn't agent technology but adoption strategy: when only some engineers use agents while others don't, chaos ensues—those without agents fall behind, and people blame the agents for problems. Khandelwal's solution was to restructure the entire codebase workflow around a special function called "ship it" that automates everything from finished code to PR-ready status. Around this framework, the team integrated agents for code review and other tasks. It wasn't perfect—agents created thousands of open issues—but it helped engineers develop actual trust in the tools.

11 Aug youtube.com

VideoAI Engineer

Taking Reinforcement Learning Cross Datacenter — Nan Jiang, Modal

Nan Jiang from Modal presents a technique for running reinforcement learning (AI training through trial-and-error) on GPU servers distributed across multiple continents. The problem is that checkpoints — saved versions of a trained AI model — weigh around 500 GB and take hours to transfer between data centers, which prevents rapid model updates. His solution is to send only around 500 MB instead: a small patch containing only the most important changes. This works because less than 1% of the weights actually change between versions — not because gradients are sparse, but because updates are very small, often a thousand times smaller than the precision allows. Modal calls the implementation Stitch.

10 Aug youtube.com

VideoAI Engineer

Lessons from Studying Every Memory System — Shlok Khemani, Independent

Shlok Khemani analyzed how different AI systems build memory of their users. He started when ChatGPT incorrectly claimed he had traveled to Turkey — something he never did, only discussed as an alternative to Thailand. The issue wasn't the error itself but that the system didn't question the contradiction or verify his email records. ChatGPT uses a "running profile" that updates in the background with approximately 4,000 tokens of densely packed information that users cannot see without hacking in. Claude started differently, without a profile but with two search tools in old conversations, later adding a third. Khemani argues that memory systems fundamentally deal with computational cost: a large profile is expensive to maintain and costly each time it's loaded into the conversation, so different systems solve the same budget problem in different ways.

10 Aug youtube.com

VideoAI Engineer

LLM Knowledge Bases: a practical guide — Ben Holmes, Warp

Ben Holmes demonstrates how to build a knowledge base by first recording thoughts loosely via voice dictation without worrying about organization. AI agents (such as Claude) then structure the material in multiple passes: adding timestamps, applying tags from a predefined list, finding sources, and creating links. Finally, a wiki is generated from the notes and the entire process runs automatically daily via cloud service, so you wake up to an updated wiki. The system can also display a graph view of all notes to help identify gaps in your own thinking.

10 Aug youtube.com

VideoAI Engineer

Memory Harnesses for Long-Running Research Agents — Stefania Druga, Sakana.ai

Stefania Druga from Sakana.ai tested how AI agents can remember information during extended tasks. She built a memory system that functions as a write-manage-read loop around the model, rather than a simple database. The key finding was negative: when all information fit within the model's context window, memory made no difference—only increased cost. However, when the answer lay far ahead—at step 124 while the question appeared at step 500—memory became critical. Her best approach was a ranked list of previous decisions, which outperformed both random search and an "oracle" with perfect memory. She ran the experiments on a local computer in Tokyo with fans around it.

10 Aug youtube.com

VideoAI Engineer

Multiplayer agentic engineering — Arjun Singh, Superconductor

Superconductor uses AI agents that participate in meetings and Slack channels to capture ideas directly from conversations without requiring manual note-taking about what needs to be done. A bot listened to a four-hour meeting, identified a customer request for clearer criteria for when work is complete, opened an issue itself, and implemented the suggestion. The company runs the agents in isolated cloud environments where they cannot accidentally access production databases, and people from support and growth can trigger actual work without needing a developer environment. Over one month, they used 10.5 billion tokens and ran 3,300 agent sessions for approximately $10,000.

9 Aug youtube.com

VideoAI Engineer

Guide, Verify, Solve — Anirban Chatterjee, Sonar

A Carnegie Mellon study showed that productivity gains from AI coding disappear after three months, while problems and complexity remain — termed verification debt. A Wharton study found developers follow AI advice 92.7 percent of the time when correct, but also nearly 80 percent when AI is instructed to lie with confidence. Anirban Chatterjee from Sonar proposes a two-layered solution: zero trust (verify code using methods other than those that wrote it) and multiple layers of review (both automated analysis and reasoning-based review). Sonar's comparisons of different AI models show that the same model is good at different things — one Claude excels at correctness while another is better at maintainability and security.

9 Aug youtube.com

VideoAI Engineer

Velocity Sickness: What Happens When Your Whole Team Gets 10x Faster — Matt Dailey, Ref.

Matt Dailey describes "velocity sickness" — the problem where AI tools make a team much faster without the output actually having impact. A newsletter writer produced one book per week with AI help, but the audience didn't read much of it. On engineering teams, there are too many code changes to review simultaneously, and decisions are made by AI agents instead of humans who understand the consequences. Dailey's solution is to separate decision-making from implementation: use shared documents for decisions (not chat histories that disappear), so agents become stateless and the team regains control.

9 Aug youtube.com

VideoAI Engineer

Evolution of agentic surfaces — Gagan Bhat & Isabella Kai He, Anthropic

Anthropic's Applied AI team presents how agents — AI systems that independently solve tasks — should be built to handle long-running production jobs. They show that outdated solutions become overhead as models become smarter: Sonnet 4.5 was cautious with memory and exited early, so they built in a fix — but when Opus 4.5 was deployed without this behavior, the fix became merely slow redundant code. The key insight is separating the agent's brain (reasoning loop) from its hands (tool execution), making systems 60–90 percent faster and making errors recoverable instead of fatal. A log of everything that happens during a session serves three purposes: seeing what the agent does, retrieving information Claude forgot, and something they call "dreaming" — a nightly job that improves the agent's memory for the next day.

9 Aug youtube.com

VideoAI Engineer

Benchmarking Coding Agents on New vs Legacy Codebases — Denys Linkov, Wisedocs

Wisedocs processed medical claims across ten different code repositories that no one wanted to touch. Denys Linkov's team spent six months consolidating them into a single repository and benchmarked whether they should have waited for better AI models instead. A refactoring task took three hours with the o3 model and resulted in ten major errors, while newer models like Sonnet 4.6 and Opus 4.8 completed it faster with fewer mistakes. However, when GPT 5.5 was given the entire job at once, it wrote 2,000 lines of code in 10 minutes that were mostly shell without actual functionality—something the model itself acknowledged. Linkov concludes that doing the work directly was worthwhile: code repositories were opened for commits, work that took months now takes a week, and developers from other parts of the company now volunteer to help.

8 Aug youtube.com

VideoAI Engineer

The New Primitives: Building AI Native Software — Kwindla Kramer, Daily

Kwindla Kramer argues that agents — AI systems capable of performing tasks autonomously — are today's primitive building blocks, much like web pages were in the 1990s. He traces technology history to show how each decade had its focus: the 1950s were about getting human intentions into computers, the 1960s brought interactivity, and the 1970s introduced abstractions for larger systems. His main point is that the real future lies in AI-embedded software that comes after agents — systems with asynchronous communication, long-lived sub-agents that share knowledge, and dynamic interfaces. He uses his own game Gradient Bang as an example of the technology he believes we should build next.

8 Aug youtube.com

VideoAI Engineer

Open Source Is Dead. Long Live Open Source. — Saoud Rizwan, Cline

Saoud Rizwan argues that the open source community has died—not because code has disappeared, but because AI-generated pull requests and security bugs have made it impossible to maintain projects. Projects like Zig, curl, and tldraw are now closing contributions or banning AI entirely. What survives instead is "open weights," meaning open AI models that companies can use themselves. Rizwan compares this to Open Compute, where Facebook opened its datacenter designs, the industry standardized on them, and everyone—including Facebook—saved billions. He urges American AI labs to release open models before the world becomes dependent on foreign alternatives.

8 Aug youtube.com

VideoAI Engineer

Anthropic's CCA Exam as a Field-Guide for Agentic Engineering — Frank Coyle, UC Berkeley

Frank Coyle reviews six production scenarios from Anthropic's Claude Certified Architect exam, demonstrating what NOT to do when building with AI agents. A critical mistake is simply using the model's response directly—you must instead check the "stop reason" to know whether the model actually succeeded in running a tool or ran out of tokens. Other common errors include giving a single agent too many tools (better to have small specialized agents with one or two tools each), and allowing agents to see each other's thinking process (which causes them to converge on the same idea). Coyle also shows practical tricks like running batch jobs for half the token cost, and isolating subtask results in separate context rows to save tokens.

7 Aug youtube.com

VideoAI Engineer

Realtime multiplayer, automation, and you! — Idan Gazit, GitHub

Idan Gazit from GitHub demonstrated how he wrote an automation description in just three lines of plain English—instructions he would give a colleague—and got Copilot to expand it into a complete workflow that updates his website from Astro 5 to Astro 7. What matters is not just that it worked, but how he built in security: the agent can only open exactly one pull request, secret keys are completely outside the system, and permitted actions are declared clearly in advance so no one can instruct around the interface. He also presented a second prototype that runs each session in an isolated virtual environment and resembles a chat app, to keep policy constraints and infrastructure decisions visible before the agent starts working.

7 Aug youtube.com

VideoAI Engineer

Always-on agents run production without the on-call tax — Justin Smith, Resolve AI

Justin Smith from Resolve AI explains how agents—AI programs that run automatically—can monitor production environments without requiring developers to be on call around the clock. When a new code version is released, the agent can read what changed, determine which monitoring metrics are relevant to that specific change, and create a customized control plan. The agent can also decide when to check again—perhaps in an hour if it's a type of error that appears intermittently, or in three days to verify everything remains stable. Smith emphasizes that approximately 70 percent of a developer's time goes to operating code rather than writing it, and that new AI tools have actually made this worse by increasing the volume of changes that must be monitored. Resolve agents work by answering three questions: when they should run, how they run (isolated in the cloud so it doesn't matter if you close your laptop), and how they know what to do.

7 Aug youtube.com

VideoAI Engineer

Codex, Behind the Harness — Dominik Kundel, OpenAI

Dominik Kundel from OpenAI presents Codex, a system for running AI models as agents—programs that can perform tasks autonomously. When the model could run fast enough, the network became the bottleneck, so they switched to websockets for a constant connection instead of resending all data. The system manages which tools the model has access to by hiding rarely-used tools from the context window and only showing them when the model searches for them. All actions—such as editing files or running commands—are executed in sandboxes for security, and when someone requests something dangerous, an exception is triggered that a completely separate, protected AI agent reviews. The entire system is open-source and written in Rust.

7 Aug youtube.com

VideoAI Engineer

The State of Model Routing — NVIDIA, Cognition, OpenRouter

The video discusses choosing between large and small AI models for different tasks. It shows that selecting the cheapest model for each individual task does not always pay off — a small model operating outside its area of competence can become very expensive by calling helper functions repeatedly. Cognition solves this by having a large model plan and delegate work to smaller models, which saved 40 percent on costs. Another finding is that it is valuable to keep the same model active throughout an entire conversation — this saves money because already processed information can be reused instead of being reprocessed.

6 Aug youtube.com

VideoAI Engineer

Compression at the Edge — Chris Alexiuk, NVIDIA

A panel discussion on how to shrink large AI models without making them less capable. The key insight is that model layers have vastly different importance — the first and last layers are critical while many middle layers contribute almost nothing. GLM 5.2 was reduced from 1.5 terabytes to 250 GB (86 percent smaller) without corresponding performance degradation. NVIDIA advocates for NVFP4, a compact numerical format that shares precision information across groups of 16 values. The panel emphasizes that real-world testing in actual use cases is more important than abstract measurements.

6 Aug youtube.com

VideoAI Engineer

Local Models: Trust, Control, Optimization — Carter Abdallah, NVIDIA

In a panel discussion from NVIDIA, experts argue that open AI models (where code can be viewed and modified) are becoming increasingly important for businesses. Lucas Atkins from Arcee AI explains that many companies chose Chinese open models not because they were better, but because availability was more secure — this is his definition of trust. Vincent Weisser from Prime Intellect demonstrates examples of how an open model can be adapted to a specific task in one or two weeks and achieve better results than more expensive closed systems. A major advantage of open models is that you own your own data and can train further on it, creating a powerful feedback loop. The panel predicts that open models will reach the same level as today's best closed systems within a year.

6 Aug youtube.com

VideoAI Engineer

Gadgets: Personal app vibe coding that is actually safe — Kenton Varda, Cloudflare

Kenton Varda from Cloudflare presents Gadgets, a way to allow AI agents to add features directly into apps without requiring developer involvement. In the example, Claude was asked to build a presentation program from a Google Document, and the agent simply added features like strikethrough, text centering, and an SVG box that the app itself lacked. Instead of today's model where user requests end up in developers' backlogs indefinitely, each user can have their own AI assistant that fixes what they need — but without creating a large security gap, because Gadgets is built on isolation: each app instance runs in a sandbox where an XSS bug cannot steal anything.

6 Aug youtube.com

VideoAI Engineer

Building Turbopuffer: Gergely Orosz (@pragmaticengineer ) × Simon Eskildsen (CEO)

# Fireside Chat: Gergely Orosz × Simon Eskildsen **Location:** Main Stage **When:** Day 2 - June 30, 2026 · 12:30pm-1:30pm ## Speakers ### Gergely Orosz Author / Founder, The Pragmatic Engineer · The Pragmatic Engineer [X/Twitter](https://twitter.com/gergelyorosz) · [LinkedIn](https://www.linkedin.com/in/gergelyorosz/) · [Website](https://pragmaticengineer.com) Software engineer, engineering leader, and author of The Software Engineer's Guidebook; best known for The Pragmatic Engineer newsletter and blog covering software engineering practices, engineering leadership, and the tech industry. Previously held engineering leadership roles at Uber and worked at companies including Skype and Skyscanner. ### Simon Eskildsen CEO and co-founder · turbopuffer [X/Twitter](https://x.com/Sirupsen) · [LinkedIn](https://www.linkedin.com/in/sirupsen/) · [Website](https://sirupsen.com) · [Blog](https://sirupsen.com/napkin) Co-founder and CEO at turbopuffer. Formerly Principal Engineer at Shopify, where he helped scale infra from 1K → 1M RPS. — [View on the schedule](https://www.ai.engineer/worldsfair/schedule?session=asn_slot_2026_06_30_main_stage_1230_2026_06_25t07_57_06_000z)

3 Aug youtube.com

VideoAI Engineer

MCP Apps: Extending the Frontier — Ido Salomon & Liad Yosef

MCP Apps is a method for AI assistants to display interactive interfaces instead of long text blocks. Liad Yosef, who created MCP UI, demonstrates how a server can return an actual button, chart, or interactive view that renders directly in chat — and when users click, responses flow back into the AI's loop. Because it is an open standard rather than a tool-specific solution, the same interface appears everywhere — from small tools to platforms with hundreds of millions of weekly users. The specification is still being developed, including how apps and chat should communicate with each other.

3 Aug youtube.com

VideoAI Engineer

MCP Tasks (async): Why Aren't Any Agents Supporting Them? — Cornelia Davis, Temporal

MCP tasks is a standard that enables AI agents to initiate long-running jobs that can be paused, reported on, and resumed without losing their work—for example, invoice processing that can wait for human approval mid-process. Cornelia Davis, an expert in distributed systems, demonstrates how this solves a real problem: when networks fail or processes crash during a long-running job, context is often lost. The specification uses a stateless core with long-running behavior as an extension, and allows the server to send updates to the client instead of the client constantly polling. Despite the technology being available, almost no AI agent supports it yet, partly because implementing long-running work correctly is difficult.

2 Aug youtube.com

VideoAI Engineer

When Will The Benchmaxxing Plague End? — Nick Heiner, Surge AI

Nick Heiner from Surge AI argues that AI models often perform better on tests than they actually do in practice—a problem he calls "benchmaxxing." He identifies several ways tests become inaccurate: some test tasks are completely broken, models may have memorized test content rather than understanding it, and companies can cheat by having systems learn to answer exactly as tests expect rather than solving the actual task. The worst cases occur when test instructions contradict each other or encounter technical bugs, making a perfect score achievable only by learning the system's weakness. Heiner argues the solution is to use domain experts, ensure test instructions and evaluation align, and pay for genuine human evaluation.

2 Aug youtube.com

VideoAI Engineer

What's Next After RLHF? — Diogo Almeida, TypeSafe AI

Diogo Almeida, who helped develop GPT-4, argues that RLHF (a training method that teaches AI models to please humans) has solved one problem but created another. The method makes models good at appearing helpful and safe, but it optimizes for making humans happy rather than for models actually doing the right thing—like how pressure to please could make a model claim a fart sound is a symphony. Almeida divides AI use into two worlds: assistance (where humans catch mistakes) and automation (where the system acts independently with real consequences). For automation, the RLHF approach is a disadvantage, since the model's built-in desire to please becomes dangerous when no one is there to oversee it. His proposal is to start optimizing for verifiable results instead of human approval.

1 Aug youtube.com

VideoAI Engineer

Emulated: The Data for Fully Autonomous Software Engineers and Companies — Joseph Wang

Emulated is building training data for AI agents designed to work as software engineers in real production environments. Their key insight is that current training environments are too simplistic—they only simulate clean code changes, not the chaotic real-world work of keeping systems running when databases crash, servers fail, and problems arise during ongoing incidents. Instead, Emulated creates simulations of entire companies and cloud infrastructure, where an AI agent must handle everything from resource allocation and cost management to scaling services when failures occur. The founders have backgrounds in network infrastructure and believe that by combining domain expertise with realistic simulations, AI agents can learn the messy, complete work that actual infrastructure engineers do.

31 Jul youtube.com

VideoAI Engineer

The Base Model Is Dead — Varun Singh, Arcee AI

Varun Singh from Arcee AI argues that the traditional view of base models—as mirrors of the internet upon which other techniques build—no longer holds true. Modern models integrate instruction data and synthetic training data much earlier in the process, and a new intermediate stage has emerged. Raw webtext takes a backseat while reinforcement learning (a method to train AI to reason better) has become so important that it changes what the base model needs to do from the start. Singh discusses practical challenges his team encountered when training the Trinity series, such as achieving the right balance across different data sources and achieving model stability early so it is prepared for what comes later. The point is that the base model's role is constantly being redefined as capacity increases.

31 Jul youtube.com

VideoAI Engineer

Verifiable Environments for AI in Biology — Kenny Workman, LatchBio

Kenny Workman from LatchBio discusses using AI for actual biological research. Biological experiments generate enormous amounts of data—two to six terabytes per experiment—too much for researchers to analyze manually. LatchBio has developed tools that adapt coding models to biology and built benchmarks to test whether AI actually solves real biological problems. The major challenge is that today's advanced AI models are not yet reliable enough for this work, and biology is harder than many other fields because experiments themselves are often messy and researchers rarely completely agree on answers.

31 Jul youtube.com

VideoAI Engineer

Ending AI Slop — Thais Castello Branco, Taste Labs

Thais Castello Branco from Taste Labs argues that AI models today are poor at subjective work such as writing and design, and that the problem is that models tend to optimize toward the average — the most likely outcome — which kills creativity. The solution is to break down design into measurable components (a logo, a headline, a color) that can be compared against the original, and then gather high-quality feedback from experts linked to specific choices. By making taste measurable and trainable in this way, AI can learn to break from the obvious when the situation requires it.

31 Jul youtube.com

VideoAI Engineer

Agents at Scale: Inside MiniMax's Model and the Infrastructure Behind It — Olive Song

Olive Song, who leads reinforcement learning at MiniMax, explains how the company built its AI model and the infrastructure behind it. MiniMax focuses on releasing open-source code so others can optimize and use the model widely. Song describes the practical work: how to train models that can build games and use computers, how to write and adjust GPU code for fast computation, and how to handle multiple modalities (text and images) without them collapsing during training. Much of the work's focus lies on technical details like KV-cache management and optimization on day one when the model is released.

31 Jul youtube.com

VideoAI Engineer

fighting slop with slop — Vaibhav Gupta, Boundary

Vaibhav Gupta from Boundary presents a strategy for handling problems with AI agents that generate incorrect code and data—he calls it "fighting garbage with garbage." His approach uses cheap AI agents themselves as tools to monitor other agents: they review transcripts, flag hallucinations (when AI invents things), and compare which solutions worked best. He combines this with stable rules and checklists that don't change, creating a solid foundation beneath the messy detection layer. The deeper strategy is to use type systems—design patterns that make certain errors impossible from the start—so agents can move quickly within defined boundaries without getting lost. He also presents BAML, a tool that lets agents work across Python, TypeScript, and Rust with strong constraints.

31 Jul youtube.com

VideoAI Engineer

Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI

Ari Morcos from DatologyAI argues that data quality is more important than raw computational resources when training AI models. Rather than simply feeding in massive amounts of data (a firehose), the focus should be on cleaning, organizing, and creating high-quality training data—like an oil refinery. A smaller model trained on carefully selected data can outperform much larger models and uses less power when running. Examples from customers like Thomson Reuters and Arcee demonstrate this strategy works in practice. It is cheaper to improve data quality than to buy more expensive computers.

31 Jul youtube.com

VideoAI Engineer

Rethinking Environments for Long-Horizon Work — Rayan Garg, Theta Software

Rayan Garg from Theta Software discusses how to measure and build environments for AI agents tackling long-horizon tasks. The challenge is that there's no clear definition of what "long-horizon" means — many use time limits where agents reach a success threshold, but this is both noisy and misleading since the same duration can mask vastly different difficulty levels. How you choose to measure significantly impacts conclusions about the model. Garg's focus is on designing environments and verification systems that make these measurements honest by understanding how poor early decisions create cascading effects throughout a task, and by verifying results from final state outcomes rather than a judge's guess.

31 Jul youtube.com

VideoAI Engineer

Teaching AI to Find Real Vulnerabilities — David Brumley, Bugcrowd

David Brumley demonstrates how to train AI models to find security vulnerabilities in code through a stepwise method—from crashing a program to reading and writing memory to full attacks. The major challenge is measurement: since many possible vulnerabilities exist, a model can always claim to have found one, so standard tests fail. Brumley built actual training pipelines with sandboxes and objective evaluators that verify whether exploits actually work. He demonstrated this on V8 (Chrome's JavaScript engine) against 41 real vulnerabilities where the best models achieved around 95 percent success, including one previously unknown vulnerability.

31 Jul youtube.com

VideoAI Engineer

Scaling to Long Horizons — Ross Taylor & Chengxi Taylor, General Reasoning

Ross Taylor and Chengxi Taylor from General Reasoning discuss training AI agents to maintain focus and perform well over extended periods—hours rather than seconds. They explain techniques such as value models, which help AI understand which choices lead to long-term success, and bootstrapping, which extracts useful signals from limited rewards. They demonstrate how frontier models failed when given real money to trade football matches, revealing that the environment where the AI was trained was not realistic enough. The key insight is that succeeding with long horizons requires not just larger context windows but, most importantly, better and more simulated environments for training.

31 Jul youtube.com

VideoAI Engineer

Data and Environment Curation for Post-Training LLMs — Mahesh Sathiamoorthy, Bespoke Labs

Mahesh Sathiamoorthy from Bespoke Labs argues that the challenge in training large language models after they are pre-trained is not the algorithms but the data and environments in which the models learn. He describes how his team built OpenThoughts, a dataset for reasoning, and learned that what matters is having many different ways to solve the same problem, not just one "correct" method. An example is how properly labeled training data about credit card compliance rules made a model much better at understanding these rules. The core idea is that careful data curation — selection and organization of good training examples — is more crucial than just more computing power.

31 Jul youtube.com

VideoAI Engineer

First Steps Toward Automated AI Research — Richard Socher, CEO Recursive AI

Richard Socher, CEO of Recursive AI, presents a vision of automating scientific research by allowing AI agents to repeatedly test ideas, identify problems, and improve solutions — without a single researcher becoming the bottleneck. He calls the goal a "Eureka machine" and demonstrates concrete early examples: an automated system that significantly improved an AI model's accuracy compared to standard training, an architecture search that finds better design choices without manual tuning, and GPU optimizations that delivered real improvements. Socher is cautious in emphasizing these are only first steps, but argues the direction is correct — that automating research in medicine, economics, astrophysics, and other fields can compress scientific progress the same way humans compressed the journey from the Enlightenment to the Moon in a few hundred years.

30 Jul youtube.com

VideoAI Engineer

Benchmarks: The Good, the Bad, and the Ugly — Ali Khial, G2i

Ali Khial from G2i examined popular coding benchmarks—tests that measure how well AI models perform at programming—alongside his top engineers. He discovered that many benchmark tasks are flawed: instructions so vague that correct answers are rejected, tests that check incorrect details like variable names, and many genuinely good solutions marked as wrong. The problem is that AI models are increasingly good at "cheating" by finding the test rather than solving the problem. Khial presents principles for benchmarks you can trust: be precise where it matters, vague where it doesn't, keep a private test set secret so it doesn't leak, and require production-quality code.

30 Jul youtube.com

VideoAI Engineer

Learning on the Job: The Future of Post-Training — Raymond Feng, Applied Compute

Raymond Feng from Applied Compute explains how AI models can continue learning after deployment through reinforcement learning. Rather than training on simple question-and-answer pairs, models are trained on real work tasks that companies need to solve. The system functions as an orchestration layer that collects model performance data, evaluates results, and uses that information to update the model. A major challenge is that models can learn to cheat — for example, by causing a tool to time out instead of actually solving the task. Another difficulty is faithfully recreating the production environment so training reflects reality, and managing data from previous interactions that cannot easily be replayed.

30 Jul youtube.com

VideoAI Engineer

Reinforcement Learning without Verifiable Rewards — Will Brown, Prime Intellect

Reinforcement learning has been easy to sell where the answer is checkable, like math or code, and Will Brown's talk is about everything else. Most valuable tasks have no clean verifier, so Prime Intellect's work is on how you build reward signal when there is no ground truth waiting. He frames RL simply first, a model acting in a harness with tools and skills, getting a reward, and nudging its weights, then asks how you keep climbing once you leave the verifiable island behind. His answer leans on environments as the anchor. You can set up judges, generate question and answer pairs grounded in real documents and repos, and use a reverse direction trick where you hide something, like a bug or a backdoor, so the model can learn to find it again, which conveniently gives you a difficulty dial to keep tasks not too easy and not too hard. He is direct about the dangers: reward hacking will find you if you are not careful, so you inspect traces, run small experiments, and bring in expert understanding. The goal he keeps returning to is making this a real science, with open models and shared benchmarks, where environments turn into new tasks and higher levels of ability.

30 Jul youtube.com

VideoAI Engineer

Your Finance Agent's Bottleneck Is You — Ramana Siddanth Emani, Auditoria AI

Ramana Siddanth Emani from Auditoria AI argues that developers themselves are often the bottleneck when scaling production agents for finance—not the models or hardware. He demonstrates how to use coding agents to automate your own developer loop: agents can run in parallel on separate code branches, pull tasks from GitHub, Jira, and QA reports, write tests, run build processes, and report results. With a human as final verifier, you can ship much more code while gradually becoming less necessary in the process. At Auditoria, they use the same pattern for financial agents that communicate with each other and reconcile data.

30 Jul youtube.com

VideoAI Engineer

Build for the Memo, Not the Demo — Shawn Chan, China Resources Holdings

Shawn Chan, who after 15 years and approximately 200 investment committee meetings makes billion-dollar decisions, says that AI-generated presentations often look polished but are "definitely wrong" in dangerous ways. He distinguishes between a demo—a beautiful presentation designed to impress—and a memo that must hold up in a room full of people whose job is to find problems. For an AI product to pass the "memo test," it requires that figures align, that facts and assumptions are kept separate, and that you can trace where each claim comes from. He provides an example of how a single incorrect sentence in a demo once cost enormous value when the market noticed it.

30 Jul youtube.com

VideoAI Engineer

Let's integrate AI Agents in Event-Sourced Systems — Divakar Kumar, FlyersSoft

Divakar Kumar demonstrates how to integrate AI agents into existing payment systems to handle complex cases that neither rule-based systems nor machine learning can solve independently—such as when a card is declined without a clear reason. Rather than rebuilding the system, he adds an agent layer on top of existing architecture based on event sourcing, where events stream through the system and agents read these to make decisions. He uses multiple agents working together—one analyzes risk, another reaches conclusions—communicating via message queues. The system is designed to prevent infinite loops and remains serverless.

30 Jul youtube.com

VideoAI Engineer

Wearing the Agent: From Group Chats to Glasses — Sai Krishna Rallabandi

Sai Krishna Rallabandi spent eight months investigating what goes wrong when AI agents transition from working for a single person to operating in group chats or smart glasses. The core challenge is that agents must track who said what during extended conversations without wasting tokens, while preventing information leakage between users. He demonstrates that two individually secure features can become insecure when combined, and proposes a solution where the agent reads everything, creates guards, then fine-tunes a smaller model with personal filters that show only relevant information and block prompt injection attacks.

30 Jul youtube.com

VideoAI Engineer

We Vetted 2000 AI Skills Before They Reached Developers — Lucas Palma, Nubank

Lucas Palma from Nubank describes how the bank built a tool called Skill Vector to control AI features — small code programs that extend what AI models can do — before they reach developers. Each feature is first scanned with automated controls that search for dangerous commands, then reviewed by an AI model to catch problems that automation misses. After checking over 2000 features this way, they found real security issues that were fed into the bank's vulnerability management system. Nubank combined automated scanning with AI review, but needed to improve guidance and prevent developers from running uncontrolled features locally before approval.

30 Jul youtube.com

VideoAI Engineer

How Kepler Built Verifiable AI for Financial Services — Vinoo Ganesh

Kepler has built a system to make AI reliable for the financial industry. The starting point is that language models are good at predicting the next word but poor at exact mathematics that the financial world actually needs. Instead of letting the model do everything itself, Kepler surrounds it with a system where every number must be traceable to its source, the model only does what it is good at, and all figures are reviewed before use. This way the system becomes reliable for financial data instead of just producing content that sounds good.

29 Jul youtube.com

VideoAI Engineer

Persona Engineering: A Field Guide to AI Synthetic Personas — Ishan Anand, InsightSciences.ai

A research team compared responses from over a thousand real people answering market surveys with AI agents playing the same role. The agents matched people well on average but lacked human variation — and small changes in the prompt (instructions) caused purchasing decisions to swing dramatically, because the model guessed at hidden patterns that were never mentioned. The problem is that a model can look correct on average while simultaneously distorting minority groups and concealing the variation that actually matters. The solution is to first compare people against people to know what agreement is even possible, then judge the synthetic personas against that standard — and most importantly: treat them as economic actors that must be validated against real results before influencing any decision.

29 Jul youtube.com

VideoAI Engineer

Why Off-the-Shelf AI Doesn't Understand Money — Udi Menkes, Intuit

Udi Menkes from Intuit explains why common AI models do not understand economics and money in a deep way. They can provide fluent answers to financial questions, but lack understanding of the unique aspects of your business or what happened when similar companies made the same decisions. Intuit's own financial AI models are trained on data from over 100 million customers and perform better than general models while being faster. Menkes discusses how true economic intelligence goes beyond just reporting what has happened — it should help you make better decisions, or even make them for you.

29 Jul youtube.com

VideoAI Engineer

SimulationMaxxing: How we ship agents 20× faster — Aman Gupta (Nubank) + Shreya Rajpal (Snowglobe)

Nubank, which serves 135 million customers, used simulated training to deploy five AI agents to production 20 times faster than before. The challenge was that testing agents took considerable time—they needed real data from customer conversations that were long and context-dependent, not just simple questions and answers. Snowglobe solved this by automatically generating simulated customer conversations with realistic details such as customer account information, tone, and purpose. This method worked well enough—simulated conversations matched real ones in approximately 80 percent of cases—to allow Nubank's team to test agents quickly and iterate without waiting for feedback from production.

29 Jul youtube.com

VideoAI Engineer

Skills are new features: Building Skill-Centric Harness — Yogendra Miraje, FactSet

A skill is a capability you provide to an AI agent — essentially a short file (skill.md) with a name and description that serves as routing signals for the agent to select the right tool. At small scales, a simple registry suffices, but as the system grows — over ten skills require search and embeddings, over a hundred require actual governance with ownership, testing, and review. Critically, skills without tests begin to malfunction when new models are released, and at enterprise scale, governing and reviewing skills is at least as important as writing them.

29 Jul youtube.com

VideoAI Engineer

Morgan Stanley's ALPHALAB: Multi-Agent Research Across Optimization Domains — Brendan Rappazzo

Morgan Stanley has built AlphaLab, a system with multiple AI agents that work together to solve research problems. You describe the problem in plain English, and the system writes code, runs tests, manages computational jobs, and performs statistical analyses autonomously. A strategy agent proposes experiments and worker agents execute them, all presented in a concise format that humans can review and approve before the system optimizes. Morgan Stanley open-sourced it and is already using it internally to find real improvements in areas such as models and credit forecasting.

29 Jul youtube.com

VideoAI Engineer

Your Agent Didn't Fail. Your Harness Did. — Vinoth Govindarajan, OpenAI

The video discusses failures that occur not in the AI model itself but in the system around it—called "harness failures". The problem is that agents can respond confidently based on outdated data while two processes overwrite each other without crashing, or when a tool is called but never receives a response. The solution is to view it as: the model proposes, the system executes, and a receipt proves it actually happened. The speaker shows common mistakes such as incorrect permissions, missing deadlines, and state that is never saved, and presents five questions to ask for every incident: what started it, what state did it inherit, what permissions were used, what ran, and what evidence remains.

29 Jul youtube.com

VideoAI Engineer

How Forward Deployed Engineering is done at Factory — Eno Reyes

Factory deploys engineers directly at its largest customers' sites to gather feedback for improving its AI agent Droid. The process functions like a factory: signals from real customer needs become development plans, are tested, and then become Droid updates. Importantly, customers own the system that connects Droid to their environment, so their data remains with them and the agent can run completely isolated from the internet. The future focus is on autonomy—how much Droid can do without human intervention—but this development must be balanced carefully.

29 Jul youtube.com

VideoAI Engineer

AI tools for Forward Deployed Engineering — Vasuman Moza, Varick Agents

Varick Agents builds AI agents that sit on top of systems companies already use, rather than forcing them to migrate. The company maps how a department actually works today—including how purchase orders are matched against invoices—and automates these processes end-to-end. Varick uses its own trained models to understand messy business data and identify when different records refer to the same entity, allowing the agent to operate autonomously with proper context.

28 Jul youtube.com