All collections

AI Engineer World's Fair 2026

Talks from AI Engineer World's Fair 2026, drafted from the AI Engineer channel's own playlists. Editorial metadata is best-effort parsed and under review.

375 videos · refreshed 23 hours ago

Event 1
Track 1
  1. 201 ↓7
  2. 202

    Are LLM Performance Benchmarks Reliable? — Ashok Chandrasekar & Jason Kramberger, Google

    4.2K views 12 likes 0 comments

  3. 203 ↓8

    Your AI Product Will Fail Unless You Can Explain It - Veronica Hylak, Hey AI

    4.1K views 213 likes 17 comments

  4. 204 ↓8
  5. 205 ↓8

    Building GTM AI Agents: Lessons from Deploying to 6,000 Users — Sait Izmit, Snowflake

    4K views 64 likes 1 comments

  6. 206 ↑161

    Operating Distributed Inference Systems at Scale — Nishant Gupta & Naman Ahuja, Meta

    4K views 22 likes 1 comments

  7. 207 ↓9

    From Agent Traces to Agent Simulations — Rustem Feyzkhanov, Snorkel AI

    3.9K views 69 likes 2 comments

  8. 208 ↓9

    Evals-Driven Development for a Mental Health AI Coach — Akele Reed & Dave Revere, SonderMind

    3.9K views 95 likes 7 comments

  9. 209 ↓9

    Building Agents Is Trivial Now, Context Is the Next Frontier — Jeff Ng, Unblocked

    3.9K views 71 likes 8 comments

  10. 210 ↓9

    Stop AI Agent Hallucinations: 5 Techniques + Production Patterns - Elizabeth Fuentes, AWS

    3.8K views 112 likes 15 comments

  11. 211 ↓8

    The Desktop Frontier — Ahmad Osman, Osmantic

    3.8K views 84 likes 6 comments

  12. 212 ↓10
  13. 213 ↓9

    The Future of Evals: From LLM as a Judge to Agent as a Judge — Aparna Dhinakaran, Arize AI

    3.7K views 74 likes 3 comments

  14. 214 ↓9

    Multiplayer agentic engineering — Arjun Singh, Superconductor

    3.7K views 19 likes 2 comments

  15. 215 ↓9

    The Missing Layer: Design Taste in AI Agents — Hassan El Mghari, Together AI

    3.7K views 81 likes 5 comments

  16. 216 ↓9
  17. 217 ↓9

    AI Agents for Performance: Ship Faster, Pay Less — Rajat Shah, Netflix

    3.7K views 69 likes 1 comments

  18. 218 ↓9

    DeepSWE: A Contamination-Resistant Coding Benchmark — James Shi, Datacurve

    3.6K views 71 likes 8 comments

  19. 219 ↓9

    The Last Human Code Review: Building Trust in AI-Generated Code — Itamar Friedman, Qodo

    3.6K views 64 likes 6 comments

  20. 220 ↓8

    The State of Model Routing — NVIDIA, Cognition, OpenRouter

    3.5K views 19 likes 7 comments

  21. 221 ↓10

    RLM: Recursive Language Models for Large Codebases - Shashi, Superagentic AI

    3.5K views 72 likes 2 comments

  22. 222 ↓9

    Benchmarking Coding Agents on New vs Legacy Codebases — Denys Linkov, Wisedocs

    3.3K views 45 likes 6 comments

  23. 223 ↓9

    How Autoresearch is changing ML research — Zhengyao Jiang, Weco

    3.3K views 74 likes 1 comments

  24. 224 ↓9

    Your Finance Agent's Bottleneck Is You — Ramana Siddanth Emani, Auditoria AI

    3.2K views 56 likes 3 comments

  25. 225 ↓9

    Lessons from Studying Every Memory System — Shlok Khemani, Independent

    3.2K views 99 likes 4 comments