What to Watch from AI Engineer World’s Fair 2026

ai
conference
youtube
agents
Author

Lawrence Wu

Published

July 9, 2026

AI Engineer World’s Fair 2026 just wrapped in San Francisco, and the AI Engineer YouTube channel has started posting talks. I wanted a quick way to decide what to watch first, so I pulled the current video metadata and ranked the posted talks by views per day. Note this post and all of the code associated with grabbing the data was written by GPT 5.5 (High) using the Codex harness.

The raw view count alone is a little unfair because the uploads are staggered. A talk posted on June 9 has had a month to accumulate views; a talk posted on July 8 has had about a day. So the main ranking below is:

views per day = views / max(1, days since upload)

As of August 1, 2026 at 9:54 PM Pacific, the early winners are:

  1. Why Agentic Systems Need Ontologies — Frank Coyle, UC Berkeley
  2. Field Guide to Fable — Thariq Shihipar, Anthropic
  3. Harness Engineering is not Enough: Why Software Factories Fail — Dex Horthy, HumanLayer
  4. Understanding is the new bottleneck — Geoffrey Litt, Notion
  5. “The engineer of the future is the person who is able to choose what is worth doing.” — Addy Osmani

Method

I included 226 posted conference-talk videos from June 8, 2026 onward. I excluded the 2026 vibe reel, the “6 Things to Know” preview, and scheduled premieres that did not have public view counts yet.

The topic and summary fields are lightweight notes derived from the title, not transcript-level summaries. The popularity ranking is the reproducible part: YouTube metadata, upload date, views, and views per day.

AI Engineer World’s Fair 2026 ran June 29 to July 2, 2026 at Moscone West in San Francisco, according to the conference schedule.

Ranked Watchlist

Rank Video Topic Summary Views Views/day
1 Why Agentic Systems Need Ontologies — Frank Coyle, UC Berkeley Ontologies for agents Frank Coyle argues that agentic systems need explicit ontologies to reason reliably over structured knowledge. 184,335 20,481.7
2 Field Guide to Fable — Thariq Shihipar, Anthropic AI-native product design A guide to Fable and the product/design patterns behind building with modern AI systems. 139,374 5,360.5
3 Harness Engineering is not Enough: Why Software Factories Fail — Dex Horthy, HumanLayer Agent harness engineering Dex Horthy argues that harness engineering alone can’t make “software factory” approaches to AI development succeed. 44,786 4,976.2
4 Understanding is the new bottleneck — Geoffrey Litt, Notion Code understanding Argues that understanding what AI-written systems do, not producing code, is the new limiting factor. 102,348 4,652.2
5 “The engineer of the future is the person who is able to choose what is worth doing.” — Addy Osmani Engineering judgment Addy Osmani argues that choosing what is worth building matters more than raw execution speed. 74,278 4,126.6
6 From fork() to Fleet: Designing an Agent Sandbox Cloud — Abhishek Bhardwaj, OpenAI Agent sandboxing Abhishek Bhardwaj describes designing a cloud sandbox system for safely running many agent processes at scale. 71,943 3,786.5
7 Building Great Agent Skills: The Missing Manual Agent skills Practical guidance for designing reusable skills that make agents more capable and dependable. 121,403 3,678.9
8 AI’s Jurassic Park Period — Aaron Stanley, dbt Labs AI risk Compares the current AI moment to Jurassic Park’s overreach and draws out the lessons. 44,099 3,674.9
9 How Forward Deployed Engineering is done at Cognition — Jia Wu Forward-deployed engineering Jia Wu describes how forward-deployed engineers work with customers at Cognition. 14,467 3,616.8
10 Everything we knew about software has changed — Theo Browne, @t3dotgg AI product strategy Theo Browne frames what is worth building now that AI has changed developer workflows and product expectations. 85,935 3,580.6
11 Building Closed-Loop Evals for a Multimodal Agent at Scale — Soumya Gupta & Jai Chopra, Uber Multimodal agent evals Uber engineers describe building closed-loop evaluations for a multimodal agent operating at scale. 28,072 3,509.0
12 The Golden Age of AI Engineering — Alexander Embiricos & Romain Huet & Peter Steinberger, OpenAI AI engineering keynote OpenAI leaders discuss why this is a golden age for engineers building with frontier models. 72,886 3,169.0
13 Every company should have a Brain — Garry Tan, Y Combinator AI business strategy Garry Tan argues that AI is rewriting the fundamental economics and dynamics of building a business. 47,428 3,161.9
14 Forward Deployed Engineering 101 — Kevin Bai, Anthropic, ex Palantir & Rippling Founding FDE Forward-deployed engineering Kevin Bai introduces how forward-deployed engineers work with customers at Anthropic. 12,458 3,114.5
15 Full Workshop: Setting Yourself Up for Success —Jason Liu, OpenAI Codex Coding agent workshop Jason Liu runs a full workshop on setting yourself up for success with OpenAI Codex. 23,946 2,993.2
16 Don’t Ship Skills Without Evals — Philipp Schmid, Google DeepMind Skill evals Philipp Schmid argues that agent skills need real evaluation coverage before they ship to production. 53,503 2,972.4
17 Thinner Agents on a Smarter Substrate: The Ontology-based Semantic Layer — Emil Eifrem, Neo4j Ontology semantic layer Emil Eifrem makes the case for thinner agents built on a smarter, ontology-based semantic substrate. 25,157 2,515.7
18 AI tools for Forward Deployed Engineering — Vasuman Moza, Varick Agents Forward-deployed engineering Vasuman Moza surveys the AI tooling that forward-deployed engineers rely on day to day. 9,155 2,288.8
19 Why We Killed Our Multi-Agent Pipeline — Subbiah Sethuraman and Abhilash Asokan, ZS Associates Multi-agent lessons A candid account of why ZS Associates scrapped their multi-agent pipeline and what replaced it. 17,217 1,913.0
20 Ending AI Slop — Thais Castello Branco, Taste Labs AI content quality Thais Castello Branco makes the case for raising the taste bar and ending AI slop. 1,760 1,760.0
21 fighting slop with slop — Vaibhav Gupta, Boundary AI content quality Vaibhav Gupta explores using generated output to detect and clean up other generated slop. 1,737 1,737.0
22 Loop Engineering from First Principles — Kyle Mistele, HumanLayer Agent loop design Kyle Mistele derives agent loop design from first principles rather than copying existing harnesses. 12,047 1,721.0
23 Claude for Long-Horizon Tasks — Lance Martin, Anthropic Long-horizon agents Lance Martin covers patterns for keeping Claude on track across long, multi-step tasks. 15,148 1,514.8
24 Reinforcement Learning without Verifiable Rewards — Will Brown, Prime Intellect Reinforcement learning Will Brown covers reinforcement learning approaches when rewards cannot be cleanly verified. 1,448 1,448.0
25 Persona Engineering: A Field Guide to AI Synthetic Personas — Ishan Anand, InsightSciences.ai Synthetic personas Ishan Anand offers a field guide to engineering AI synthetic personas that behave believably. 4,244 1,414.7
26 Learning on the Job: The Future of Post-Training — Raymond Feng, Applied Compute Post-training Raymond Feng argues models should keep learning on the job after deployment. 1,350 1,350.0
27 In the Land of AI Agents, the Verifiers Are King — Tariq Shaukat, Sonar Agent verification Makes the case that verification, not generation, is the decisive capability for reliable agents. 15,976 1,331.3
28 The Future Is Domain-Specific Agents - Justin Schroeder, StandardAgents Domain agents Argues for agents specialized around concrete domains rather than generic assistants. 43,421 1,315.8
29 CrabRAG: Why Automated Assistants Need Graph Memory, Not More Tokens — Stephen Chin, Neo4j Graph memory for RAG Stephen Chin argues automated assistants need graph-based memory rather than ever-larger token windows. 12,432 1,243.2
30 Teaching AI to Find Real Vulnerabilities — Prof. David Brumley, Bugcrowd Security research David Brumley shows how to teach AI systems to find real software vulnerabilities. 1,159 1,159.0
31 State of the Union: Why Local, Why Now — NVIDIA, Osmantic, Roboflow, EXO Labs, @matthew_berman Local AI A multi-vendor panel on the case for running AI locally rather than in the cloud. 23,947 1,140.3
32 Your Finance Agent’s Bottleneck Is You — Ramana Siddanth Emani, Auditoria AI Finance agents Argues the human reviewer, not the model, is the real bottleneck in finance agent workflows. 2,226 1,113.0
33 Agents at Scale: Inside MiniMax’s Model and the Infrastructure Behind It — Olive Song Model infrastructure Olive Song covers MiniMax’s model and the infrastructure that runs its agents at scale. 1,032 1,032.0
34 Forward Deployed Engineering at Cursor — Pauline Brunet Forward-deployed engineering Pauline Brunet describes how forward-deployed engineers work directly alongside customers at Cursor. 17,954 997.4
35 From Writing Code to Designing Systems: How the Developer Role is Changing — Chris Noring, Microsoft Developer role shift Discusses how the developer’s role is moving from writing code to designing broader systems. 20,908 995.6
36 Stop Making Models Bigger, Make Them Behave — Kobie Crawford, Snorkel Model behavior tuning Kobie Crawford argues for improving model behavior through better data rather than simply scaling model size. 50,806 977.0
37 HTML Is All Agents Need — James Russo, HeyGen Agent-generated graphics Argues HTML is the right substrate for agents to produce visual and interface output. 10,514 955.8
38 First Steps Toward Automated AI Research — Richard Socher, CEO Recursive AI Automated AI research Richard Socher lays out early steps toward AI systems that conduct their own research. 1,867 933.5
39 Skills are new features: Building Skill-Centric Harness — Yogendra Miraje, FactSet Agent skills Yogendra Miraje argues skills are the new feature unit and shows how to build a skill-centric harness. 2,800 933.3
40 The Great Loops Debate — Dex Horthy, Geoff Huntley, Ian Livingstone, Greg Pstrucha, @insecure-agents Agent loop design A panel debates how the control loops driving AI agents should be structured. 13,862 924.1
41 How Forward Deployed Engineering is done at Factory — Eno Reyes Forward-deployed engineering Eno Reyes describes how forward-deployed engineering works at Factory. 3,580 895.0
42 What’s Next After RLHF? — Diogo Almeida, TypeSafe AI Post-training Diogo Almeida looks at what comes after RLHF for aligning models. 890 890.0
43 Claude Fable, Claude Tag, and Anthropic’s Culture — Cat Wu & Thariq Shihipar ft Simon Willison Fireside conversation Simon Willison talks with Anthropic’s Cat Wu and Thariq Shihipar about building with AI today. 14,175 833.8
44 Why Large? Tiny LMs & Agents on Edge/Robotics — Cormac Brick, Google Small models on edge Cormac Brick makes the case for tiny language models and agents running on edge and robotics hardware. 5,771 824.4
45 State of Data — Sean Cai, Independent / State of Data State of data Sean Cai surveys where the data ecosystem stands for AI practitioners today. 4,832 805.3
46 Scaling to Long Horizons — Ross Taylor & Chengxi Taylor, General Reasoning Long-horizon agents Ross and Chengxi Taylor cover scaling agent training to long-horizon tasks. 777 777.0
47 Special Topics in Kernels, RL, Reward Hacking in Agents — Daniel Han, Unsloth RL and reward hacking Daniel Han covers kernels, reinforcement learning, and reward-hacking pitfalls in training agents. 11,446 763.1
48 The Dirty Secret of Forward Deployed Engineering — Natalie Meurer, Sierra Forward-deployed engineering Natalie Meurer shares the unglamorous realities behind forward-deployed engineering work. 3,041 760.2
49 AI Agents for Performance: Ship Faster, Pay Less — Rajat Shah, Netflix Performance agents Rajat Shah describes how Netflix uses AI agents to ship performance improvements faster and at lower cost. 2,978 744.5
50 Rethinking Environments for Long-Horizon Work — Rayan Garg, Theta Software Agent environments Rayan Garg rethinks the environments agents need for long-horizon work. 739 739.0
51 “Software engineering is not about writing code” — Benoit Schillings, Google DeepMind VP of Research Engineering judgment Benoit Schillings argues software engineering is about problem-solving, not writing code. 10,965 731.0
52 Emulated: The Data for Fully Autonomous Software Engineers and Companies — Joseph Wang Training data Joseph Wang discusses the data needed for fully autonomous software engineers. 719 719.0
53 Stop Burning Tokens: Why self-improvement needs domain expertise first - Annabell Schäfer, Langfuse Self-improving agents Argues agent self-improvement needs domain expertise first before it can stop wasting tokens. 9,630 687.9
54 WTF Is the Context Layer? The Missing Infrastructure for Production Agents — Prukalpa Sankar Context infrastructure Prukalpa Sankar defines the context layer as the missing infrastructure production agents actually need. 11,990 666.1
55 Verifiable Environments for AI in Biology — Kenny Workman, LatchBio Scientific agents Kenny Workman presents verifiable environments for applying AI to biology. 658 658.0
56 The Base Model Is Dead — Varun Singh, Arcee AI Model training Varun Singh argues the plain base model is no longer the useful unit. 621 621.0
57 Active Graph Agent Runtime (BabyAGI 4) — Yohei Nakajima, Untapped Capital Graph agent runtime Yohei Nakajima introduces an active graph-based agent runtime in the latest iteration of BabyAGI. 6,132 613.2
58 SimulationMaxxing: How we ship agents 20× faster — Aman Gupta (Nubank) + Shreya Rajpal (Snowglobe) Agent simulation Nubank and Snowglobe engineers describe using simulation to ship agents dramatically faster. 1,810 603.3
59 RAG is dead, right?? — Kuba Rogut, Turbopuffer RAG relevance Kuba Rogut examines whether retrieval-augmented generation is still necessary as context windows grow. 31,824 600.5
60 Benchmarks: The Good, the Bad, and the Ugly — Ali Khial, G2i AI benchmarking Ali Khial reviews what benchmarks measure well, measure poorly, and get outright wrong. 600 600.0
61 The Production AI Playbook: Deploying Agents at Enterprise Scale — Sandipan Bhaumik, Databricks Enterprise agent deployment Sandipan Bhaumik shares a playbook for deploying AI agents at enterprise scale. 26,058 592.2
62 Citation Needed: Provenance for LLM-Built Knowledge Graphs — Daniel Chalef, Zep AI Knowledge-graph provenance Daniel Chalef covers tracking provenance and citations in LLM-constructed knowledge graphs. 5,260 584.4
63 Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI Data quality Ari Morcos argues data quality multiplies the value of every unit of compute. 583 583.0
64 Your Attention Is the Bottleneck, Not Your Agents — Zack Proser, WorkOS Human attention limits Zack Proser argues the real bottleneck in agent workflows is human attention, not the agents. 29,641 581.2
65 Let’s integrate AI Agents in Event-Sourced Systems — Divakar Kumar, FlyersSoft Event-sourced agents Divakar Kumar shows how to integrate AI agents into event-sourced system architectures. 1,152 576.0
66 A Practitioner’s Guide to Graphs - Tim Ainge, Good Collective Graphs for AI A practitioner’s guide to using graph structures in AI applications. 7,970 569.3
67 We Cut 94% of AI Coding Tokens With a Local Code Index - Rajkumar Sakthivel, Tesco Code indexing Shows how local code indexing can sharply reduce token usage in AI coding workflows. 19,129 562.6
68 Your Moat Is Your Data Model — Mike Phipps, Gates Foundation Data model as moat Mike Phipps argues a company’s data model, not its models or prompts, is its durable competitive moat. 5,312 531.2
69 Turn 10,994 Notes Into Memory - Paul Iusztin, Decoding AI & Louis-François Bouchard, Towards AI Agent memory Shows how to turn a large personal note corpus into usable agent memory. 17,997 499.9
70 Your agent architecture has a half-life of 6 months — Dan Farrelly, CTO, Inngest Agent architecture Argues agent architectures decay fast and must be designed to be replaced as the field moves. 5,495 499.5
71 The Unreasonable Effectiveness of Separating the Task from the Model — Maxime Rivest & Isaac Miller Task-model separation Makes the case that separating the task definition from the underlying model yields outsized gains. 4,248 472.0
72 Data and Environment Curation for Post-Training LLMs — Mahesh Sathiamoorthy, Bespoke Labs Post-training data Mahesh Sathiamoorthy covers curating data and environments for post-training LLMs. 467 467.0
73 Evals-Driven Development for a Mental Health AI Coach — Akele Reed & Dave Revere, SonderMind Healthcare evals Akele Reed and Dave Revere describe evals-driven development for a mental health AI coach. 3,227 461.0
74 Recursive Model Improvement — Lee Robinson, Cursor, SpaceXAI Recursive model improvement Lee Robinson discusses using models to recursively improve themselves and the tools around them. 7,596 446.8
75 Every Harness Will Become A Claw — Sam Bhagwat, Mastra Agent harnesses Argues every agent harness is converging toward a Claw-like general-purpose tool. 4,879 443.5
76 From Agent Traces to Agent Simulations — Rustem Feyzkhanov, Snorkel AI Agent simulation Rustem Feyzkhanov describes moving from recorded agent traces to full agent simulations. 3,012 430.3
77 DeepSWE: A Contamination-Resistant Coding Benchmark — James Shi, Datacurve Coding benchmarks James Shi introduces DeepSWE, a coding benchmark designed to resist training-data contamination. 2,544 424.0
78 How Forward Deployed Engineering is done at Kepler — Vinoo Ganesh Forward-deployed engineering Vinoo Ganesh describes how forward-deployed engineering is practiced at Kepler. 1,678 419.5
79 Build for the Memo, Not the Demo — Shawn Chan, China Resources Holdings AI product strategy Shawn Chan argues teams should build for the decision memo rather than the flashy demo. 836 418.0
80 Modern Post-Training: A Deep Dive — Will Brown, Prime Intellect Open AI infrastructure Will Brown presents the Prime Intellect stack for open, decentralized AI training and infrastructure. 7,856 413.5
81 HTML is All You Need (for Agents to Make Graphics) - Amol Kapoor, Nori Agent-generated graphics Argues that HTML is a strong substrate for agents producing visual outputs. 13,985 411.3
82 We Vetted 2000 AI Skills Before They Reached Developers — Lucas Palma, Nubank Skill vetting Lucas Palma describes vetting two thousand AI skills before releasing them to developers. 1,227 409.0
83 Build Evals That Actually Matter - Nick Ung & Akshay Sharma, Lyft AI evals Nick Ung shares how to design evaluations that meaningfully measure agent quality instead of vanity metrics. 5,263 404.8
84 “The biggest challenge in your stack? Evals, Evals, Evals” - 2026 State of AI Engineering results State of AI engineering A data-driven survey of where AI engineering stands in 2026. 4,398 399.8
85 How Kepler Built Verifiable AI for Financial Services — Vinoo Ganesh Verifiable AI in finance Vinoo Ganesh explains how Kepler built verifiable AI systems for financial services. 1,181 393.7
86 What if the harness mattered more than the model? - Aditya Bhargava, Etsy Evaluation harnesses Makes the case that the surrounding harness can matter more than model choice for real results. 9,824 393.0
87 Your Agent Didn’t Fail. Your Harness Did. — Vinoth Govindarajan, OpenAI Agent harness design Vinoth Govindarajan argues most agent failures trace back to the harness, not the model. 1,151 383.7
88 The Messy Reality of Scale: Synthetic Data and Pre-Training — Marah Abdin & Robert McHardy, poolside Synthetic data at scale poolside shares the messy realities of synthetic data generation and pre-training at scale. 2,206 367.7
89 Design Patterns for AI Trust: Juries, Libraries, and Agent Tiers — Alex Bauer, Upside.tech AI trust patterns Presents design patterns like juries, libraries, and agent tiers for building trustworthy AI systems. 7,612 362.5
90 The Future of Evals: From LLM as a Judge to Agent as a Judge — Aparna Dhinakaran, Arize AI Agent-as-judge evals Aparna Dhinakaran traces the shift from LLM-as-judge to agent-as-judge evaluation. 2,782 347.8
91 Imagination Engineering: “Live in the future and then build what’s missing.” AI product design Eve Bouffard explores how AI expands what designers can imagine and prototype. 5,350 334.4
92 Morgan Stanley’s ALPHALAB: Multi-Agent Research Across Optimization Domains — Brendan Rappazzo Multi-agent research Brendan Rappazzo presents Morgan Stanley’s ALPHALAB multi-agent system for optimization research. 958 319.3
93 Beyond the Harness: A Journey Towards Adaptative Engineering - Rajiv Chandegra, Annicha Labs Adaptive engineering Explores engineering systems that adapt beyond static evaluation harnesses. 7,704 308.2
94 From Signal to PR: Anatomy of a Self-Improving Agent — Jason Lopatecki, Arize Self-improving agents Jason Lopatecki dissects a self-improving agent that turns production signals into pull requests. 2,430 303.8
95 How Forward Deployed Engineering is done at Ramp — Leo Mehr Forward-deployed engineering Leo Mehr describes how forward-deployed engineering works at Ramp. 1,204 301.0
96 The UX of AI: Making AI-Powered Apps Your Users Don’t Hate - Kathryn Grayson Nanz, Progress Software AI UX design Offers UX guidance for building AI-powered apps that users actually enjoy using. 3,945 281.8
97 “I’ve never seen anything scarier than an LLM with tool calls.” — Erik Meijer aka @HeadinTheBox Formal verification for agents Erik Meijer argues for grounding agent-written code in formal proofs rather than trusting agent behavior directly. 5,233 275.4
98 Recursive Coding Agents - Raymond Weitekamp, OpenProse Recursive coding agents Explores coding agents that recursively decompose and improve their own work. 10,029 271.1
99 Serving 2 Million Models Without Melting: Scaling the Hugging Face Hub — Arek Borucki, Hugging Face ML infrastructure at scale Arek Borucki explains how the Hugging Face Hub serves two million models without falling over. 1,081 270.2
100 How Evals and Prompts Shape Agent Behavior — Preetika Bhateja & Daniel Bump, YouTube Ads Evals and prompting YouTube Ads engineers show how evaluations and prompts jointly shape agent behavior. 2,053 256.6
101 The agent-ready web: Simplify user actions with WebMCP — Tara Agyemang, Google Agent-ready web Tara Agyemang introduces WebMCP for making web actions easier for agents to perform. 12,196 239.1
102 The Desktop Frontier — Ahmad Osman, Osmantic Desktop AI Explores running capable AI on the desktop as the next frontier. 2,525 229.5
103 Computer-Use 2.0: Agents Just Got Multi-Cursor — Francesco Bonacci, Cua Computer-use agents Francesco Bonacci introduces multi-cursor computer-use agents that can act across several interfaces at once. 3,835 225.6
104 How Forward Deployed Engineering is done at Decagon — Sunny Rekhi Forward-deployed engineering Sunny Rekhi describes how forward-deployed engineering is done at Decagon. 895 223.8
105 Why Off-the-Shelf AI Doesn’t Understand Money — Udi Menkes, Intuit Financial domain AI Udi Menkes argues general-purpose AI lacks the domain grounding that money-related tasks demand. 671 223.7
106 Build Systems, Not Code - Angie Jones, Agentic AI Foundation Autonomous engineering Angie Jones argues engineers should build systems rather than hand-write code in the agent era. 7,716 208.5
107 Evaling Video Slop — Maor Bril, Character.ai Video generation evals Maor Bril covers how to evaluate low-quality AI-generated video at scale. 1,447 206.7
108 AI on Your Lakehouse: Context Comes in Shapes, Not Queries — Zach Blumenfeld, Neo4j Lakehouse context Zach Blumenfeld argues AI context on a lakehouse comes in graph shapes, not flat queries. 1,798 199.8
109 Using Spec-Driven Development for Production Workflows - Erik Hanchett, AWS Spec-driven development Shows how specifications can drive more reliable AI-assisted production workflows. 6,721 197.7
110 Training Frontier Models to Out-Think Hackers — Uri Rolls, Arithmetic & Thom Wolf, Hugging Face Security-focused training Covers training frontier models to anticipate and outmaneuver attackers. 1,577 197.1
111 Wearing the Agent: From Group Chats to Glasses — Sai Krishna Rallabandi Wearable agents Sai Krishna Rallabandi traces agents moving from group chats onto wearable glasses. 575 191.7
112 Perception Agents — Antje Barth, Amazon AGI Lab Perception agents Antje Barth introduces agents focused on perception across multimodal inputs. 1,724 191.6
113 Notion’s Token Town — Sarah Sachs, Notion LLM token economics Sarah Sachs walks through how Notion reasons about token usage and cost at scale. 1,722 191.3
114 Should AI Engineers Still Read Code in 2026? The Z/L Continuum — Alex Volkov, ThursdAI AI coding practice Asks how much code AI engineers should still read as agents write more of it. 4,204 191.1
115 The Pipeline Is Dead - Iris ten Teije, Sky Valley Ambient Computing Agentic workflows Reframes traditional software pipelines around more dynamic, ambient AI workflows. 4,741 189.6
116 AI System Design: From Idea to Production - Apoorva Joshi, MongoDB AI system design Walks through moving an AI system from concept to production. 6,345 186.6
117 How Autoresearch is changing ML research — Zhengyao Jiang, Weco Autonomous coding agents Zhengyao Jiang recounts how an AI agent topped the leaderboard of OpenAI’s hiring challenge. 2,909 181.8
118 From Blind Spots to Merged PRs: Continuous Agentic Performance Optimization - May Walter, Hud Agentic optimization May Walter covers continuously optimizing agent performance from blind spots through to merged pull requests. 2,254 173.4
119 Why Eval++ Is the Next Great Compute Primitive — Sunil Pai & Matt Carey, Cloudflare Evaluation infrastructure Cloudflare engineers argue that richer evaluation is becoming a core compute primitive. 9,315 172.5
120 Teaching Coding Agents to do Spreadsheets - Nuno Campos, Witan Labs Spreadsheet agents Covers how to make coding agents useful for spreadsheet-heavy work. 4,105 171.0
121 Stop AI Agent Hallucinations: 5 Techniques + Production Patterns - Elizabeth Fuentes, AWS Agent hallucinations Presents five production-tested techniques for reducing AI agent hallucinations. 3,580 170.5
122 Building an Autonomous Engineering Org - Angie Jones, Agentic AI Foundation Autonomous engineering Discusses the organizational patterns behind autonomous AI-assisted engineering. 5,783 170.1
123 Building an ACP-Compatible Agent Live — Bennet Fenner, Zed Agent protocols A live build showing how an agent can integrate with the Agent Client Protocol ecosystem. 4,074 169.8
124 Why MCP and ChatGPT Apps Use Double Iframes — Frédéric Barthelet, Alpic MCP app architecture Frederic Barthelet explains the double-iframe pattern behind MCP and ChatGPT apps. 7,882 167.7
125 Build AI Systems for Discernment, Not Approval - Angel Ortmann Lee, Duolingo AI product judgment Advocates building AI systems that help users judge well instead of merely agreeing with them. 4,185 167.4
126 Why Can’t Anyone Answer Questions About the Business? — Garrett Galow, WorkOS Business intelligence agents Garrett Galow examines why business questions stay hard to answer and how AI can help. 8,354 163.8
127 Your Agent’s Biggest Lie: “I Searched the Web” — Rafael Levi, Bright Data Agent web access Rafael Levi exposes how often agents fake web searches and how to give them real access. 7,206 160.1
128 Why More Context Makes Your Agent Dumber and What to Do About It — Nupur Sharma, Qodo Context management Nupur Sharma explains why piling on context degrades agents and how to curate it instead. 8,483 157.1
129 Full Workshop: Better Auth — Paola Estefania, Better Auth Agent authentication Covers authentication patterns purpose-built for AI agents. 1,710 155.5
130 The Agentic AI Engineer - Benedikt Sanftl, Mutagent Agentic engineering Defines what changes when AI engineers work through agentic systems. 5,048 153.0
131 MCP Apps: Primitives, discovery, and the Future of Software - Pietro Zullo, Manufact, Inc MCP apps Covers MCP app primitives, discovery, and how software interfaces may evolve. 4,055 150.2
132 Everything Is a Rollout — Alex Shaw + Ryan Marten, Terminal-Bench, Harbor, Laude Institute Agent benchmarking Reframes agent evaluation and training around the idea that everything is a rollout. 1,140 142.5
133 Why Your Agent Disagrees With Itself (And What To Do About It) - Diane Lin, Datadog Agent consistency Diane Lin examines why an agent can contradict its own earlier answers and how to make its behavior consistent. 1,707 142.2
134 Skills are the New SDKs - Elvin Aghammadzada, DataRobot Agent skills Elvin Aghammadzada argues that reusable skills are becoming the new packaging unit for agent capabilities, replacing SDKs. 1,707 142.2
135 Every Solo Agent Builder Eventually Reinvents a Worse Version of CI/CD - Sumaiya Shrabony Agent tooling Argues solo builders end up rebuilding CI/CD-like infrastructure for their agents anyway. 2,979 141.9
136 Your AI Product Will Fail Unless You Can Explain It - Veronica Hylak, Hey AI Explainability Positions explanation as a core product requirement for successful AI systems. 3,794 140.5
137 On AI and Knowledge — Pablo Castro, Distinguished Engineer & CVP for AI Knowledge, Microsoft AI knowledge systems Pablo Castro reflects on how AI reshapes the creation and use of organizational knowledge. 2,107 140.5
138 Content Is Code - Matt Palmer, Conductor Content as code Frames content itself as a form of code in AI-driven workflows. 1,920 137.1
139 RLM: Recursive Language Models for Large Codebases - Shashi, Superagentic AI Recursive language models Introduces recursive language models as a technique for reasoning over very large codebases. 2,723 136.2
140 Frontier results, on device - RL Nabors, Arize On-device AI Covers frontier-quality AI results running closer to the device. 4,247 128.7
141 The AI bugpocalypse is here. Now what? - Jack Cable, Corridor AI-generated code security Argues that AI-written code is fueling a surge of security bugs and outlines how to respond. 2,573 128.7
142 Autonomous Agents for Scientific Tasks - Sina Shahandeh, Radicait Scientific agents Sina Shahandeh describes building autonomous agents that carry out scientific research tasks. 1,779 127.1
143 Continual Learning for AI Agents: From Failures to Durable Improvements - Soheil Feizi, RELAI Continual learning Looks at how agents can convert failures into durable improvements over time. 3,386 125.4
144 The Log Is The Agent - Ishaan Sehgal, Omnara Agent observability Argues the log is the primary interface for understanding and driving an agent. 4,630 125.1
145 Agents Building Agents - Alfonso Graziano, Nearform Meta-agents Explores agents that construct and orchestrate other agents. 4,245 124.9
146 Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs Long-horizon evals Lukas Petersson introduces Vending-Bench for evaluating agents on long-horizon tasks. 995 124.4
147 Agentic Security: Permissions, Provenance, and the Agent Supply Chain — Steve Yegge, Gas Town Agent security Steve Yegge covers permissions, provenance, and supply-chain risk for AI agents. 1,489 124.1
148 Stop Renting Your Cognitive Infrastructure - Thiyagarajan Maruthavanan, Kalmantic Labs AI ownership Advocates owning rather than renting the cognitive infrastructure behind AI systems. 1,648 117.7
149 Enterprise Agents Have a Structure Problem - Ishita Daga, Tesla Enterprise agents Ishita Daga argues enterprise agents fail on structure and organization rather than model quality. 1,411 117.6
150 Self Driving Products: Product Signals to Pull Requests — Joshua Snyder, PostHog Product automation Joshua Snyder shows how product signals can drive automated pull requests. 6,106 117.4
151 From Systems of Record to Systems of Context — Omri Bruchim & Tomer Ast, monday.com Systems of context Argues software is shifting from systems of record to systems of context built for AI. 1,137 113.7
152 Road to 5 Million Tokens: Breaking Barriers in Long Context Training — Max Ryabinin, Together AI Long-context training Max Ryabinin covers the engineering behind training models with multi-million-token context. 5,918 109.6
153 Video Has No Memory. Here’s How We Built One. — James Le, TwelveLabs Video memory James Le describes how TwelveLabs built persistent memory for video understanding. 986 109.6
154 Your Agents Need a Save Button - Hamza Tahir, ZenML Agent state persistence Argues agents need durable checkpointing so their work can be saved and resumed reliably. 1,526 109.0
155 Deterministic Infra for Non-Deterministic AI Agents - Nishant Gupta, Meta Superintelligence Labs Agent infrastructure Discusses infrastructure patterns for making stochastic agents more controllable. 3,560 107.9
156 Local Agentic Theory For Mobile Games — Shafik Quoraishee & Joanne Song, The New York Times On-device game agents Covers running local agentic reasoning for mobile games at The New York Times. 938 104.2
157 Your coding agent doesn’t always follow your rules — Talha Sheikh, Checkout.com Agent reliability Discusses why coding agents drift from instructions and what that means for reliability. 2,403 100.1
158 Learned Execution Graphs for Anomaly Detection & Drift in APIs — Ritvik Pandya, JP Morgan Chase API anomaly detection Ritvik Pandya applies learned execution graphs to detect anomalies and drift in APIs. 899 99.9
159 What Does Done Even Mean? Agents and Paperclip’s Liveness Model - Dotta, Paperclip Agent task completion Proposes a liveness model for defining when an agent’s task is actually complete. 1,978 98.9
160 The Agentic Web and the Bazaar Era of AI - Ramesh Raskar, MIT Media Lab Agentic web Frames the emerging agentic web as a decentralized “bazaar” era for AI-driven commerce and services. 1,934 96.7
161 Agents Need Receipts, Not More Tool Calls - Armanas Povilionis, Alithea Bio Agent auditability Argues agents should produce verifiable receipts of their actions rather than simply making more tool calls. 1,331 95.1
162 ReviewDebt: a practical framework for scoring every pull request — Sachin Gupta, Ebay Code review process Introduces a scoring framework for tracking review debt across pull requests. 1,882 94.1
163 The Missing Layer After Launch - Raphael Kalandadze, Wandero AI Post-launch AI systems Discusses the operational/product layer needed after an AI product ships. 2,539 94.0
164 The Prompt Is Still a Punch Card - Ted Johnson, JoinIn AI Prompting interfaces Compares prompts to old low-level interfaces and suggests better abstractions are needed. 2,813 93.8
165 Through the AI Fog: The Architectural Decision Agentic Security Depends On — Manoj Nair, Snyk Agentic security architecture Argues one architectural decision determines whether agentic security holds up. 1,124 93.7
166 I Run a Fleet of AI Agents Across Three Machines. Here’s What Broke. - Kyle Jaejun Lee, KRAFTON Multi-agent operations Lessons from running multiple agents across machines and dealing with the operational failures. 2,219 92.5
167 From Tokens to Cells: Foundation Models for Single-Cell Biology - Akram Baharlouei, Altos Labs Single-cell biology models Akram Baharlouei presents foundation models that treat single-cell biology data much like tokens in a language model. 1,162 89.4
168 Think You Can Build a Game with AI? Think Again! - Danielle An & David Hoe, Meta AI game development A reality check on the limits and challenges of building games with AI. 2,006 83.6
169 500 people vibe-coded for 30 days. I was one of them. - Sanja Grbic, Automattic Vibe coding Lessons from a month-long large-scale vibe-coding experiment. 2,062 82.5
170 Don’t Let the LLM Drive - Ornella Bahidika & Joel Allou, Microsoft Agent control flow Makes the case for keeping deterministic control around the LLM rather than letting the model drive the whole workflow. 985 82.1
171 The Factory That Dreams: 39 AI Agents, No Framework - Rushabh Doshi, Machinecraft Multi-agent systems Describes running 39 AI agents together without relying on an agent framework. 1,697 80.8
172 Your LLM Stack Is a 2008 Database With Better Marketing — Lovina Dmello, NVIDIA LLM infrastructure Argues today’s LLM stacks resemble dated database systems dressed up with new marketing. 964 80.3
173 Semantic Blindness: 500,000 Sensors Confused an LLM - Raahul Singh & Vanč Levstik, Phaidra Sensor data / LLM limits Explores how a massive sensor dataset exposed blind spots in an LLM’s semantic understanding. 1,603 80.2
174 GTM Is You - Victoria Melnikova, Evil Martians Go-to-market A builder-oriented view of go-to-market where technical creators carry more of the motion. 1,999 80.0
175 Chat and citations won’t save your vertical AI - Atul Ramachandran, Filed Inc Vertical AI product Argues chat interfaces and citations alone aren’t enough to make vertical AI products succeed. 1,672 79.6
176 Agents Need Feature Flags - Sachin Gupta Agent rollout control Makes the case for feature-flagging agent behavior so rollouts can be controlled safely. 1,092 78.0
177 Stop Writing Tone Instructions. Layer Them. - Isadora Martin-Dye, Isadora & Co Prompt tone design Advocates layering tone instructions rather than writing them as one monolithic block. 2,760 76.7
178 You Might Not Need 50 Diffusion Steps — Ziv Ilan, Nvidia Diffusion efficiency Ziv Ilan shows how to cut diffusion sampling steps without sacrificing quality. 3,414 74.2
179 How we taught agents to use good retrieval - Hanna Lichtenberg, Mixedbread AI Agent retrieval Explains approaches for making agents use retrieval more effectively. 1,719 68.8
180 remobi.app: Don’t change your terminal workflow for mobile Mobile dev tools Demos remobi.app, which brings your existing terminal workflow to mobile without changing it. 1,365 68.2
181 A Song of Types and Agents - Roberto Stagi, Ratel Type systems for agents Looks at how strong typing can make agent-built software more reliable. 1,344 67.2
182 Structuring the Unstructured - Cedric Clyburn, Red Hat Data structuring Covers techniques for turning unstructured information into usable structured data. 2,277 67.0
183 Agents in Production: How OpenGov Built and Scaled OG Assist - Gabe De Mesa, OpenGov Agents in production A case study on building and scaling OpenGov’s OG Assist agent. 2,231 62.0
184 You Didn’t Ship a Bug. You Just Wrote It for a Human. - Ravi Madabhushi, Scalekit Human-facing bugs Argues that agent mistakes shipped to users are still bugs, just paid for by a human instead of a compiler. 789 60.7
185 Stop Evaluating Models Like It’s the 50s - Alejandro Vidal, Mindmakers Model evaluation Argues for modernizing how AI models are evaluated beyond outdated benchmarks. 1,146 60.3
186 Your Agent Failed in Prod. Good Luck Reproducing It. - Tisha Chawla & Susheem Koul, Microsoft Production debugging Covers the reproducibility challenge when production agents fail. 1,990 60.3
187 We Gave an Agent Production Code Access and Then Tried to Sleep at Night — Moritz Johner, Form3 Agents in production A candid account of granting an agent production code access and managing the risk. 713 59.4
188 It’s 10pm. Do You Know Where Your Agents Are? — Kim Maida, Keycard Agent observability Argues teams need real visibility and control over what their agents are doing. 712 59.3
189 Build the AI GTM Agent That Knows the Buyer - Dr. Sajjan Kanukolanu, Position2 (Position Squared) Go-to-market agents Shows how to build a go-to-market agent that understands the buyer well enough to drive real sales motions. 707 58.9
190 A Genius With Amnesia - Victor Savkin, Nx Agent memory Frames LLMs as brilliant but memoryless and explores how to fix that. 2,102 58.4
191 Your Voice Agent Doesn’t Need a Frontier Model - Joel Allou & Ornella Bahidika, Microsoft Voice agents Argues that voice agents can run effectively on smaller models rather than requiring a frontier model. 699 58.2
192 Develop at Idea Velocity - Jeffrey Lee-Chan, Snapchat Development speed Makes the case for building fast enough to keep pace with ideas rather than process. 1,222 58.2
193 Your agent is blindfolded — Johan Lajili, Poolside AI Agent observability Argues agents need better visibility into their environment to act effectively. 1,356 56.5
194 SWE-Marathon: Evaluating Coding Agents at Billion-Token Scale - Rishi Desai, Abundant AI Coding-agent evals Presents large-scale evaluation of coding agents over very long token horizons. 1,360 54.4
195 Claws Out: Securing and Building with OpenClaw - Nick Taylor, Pomerium OpenClaw security Covers security considerations for building on and with the OpenClaw ecosystem. 1,138 54.2
196 Agent Output Is Not UX: Rendering Layer Your LLM Pipeline Is Missing - Bala Ramdoss, Amazon Lens Agent UX rendering Argues raw agent output is not UX and that a rendering layer is the missing piece in most LLM pipelines. 641 53.4
197 Can Oncology Workflows Run Without Human Touch? - Anant Shankhdhar, Risa Labs Healthcare agents Anant Shankhdhar asks how far oncology workflows can be automated before a human must stay in the loop. 641 53.4
198 Your Agent Is Wasting Tokens and You Don’t Know It - Erik Hanchett, AWS Token efficiency Highlights hidden token waste in agent systems and ways to reduce it. 1,765 51.9
199 You Can’t Prompt the Room: The Last Skill AI Won’t Replace - Balázs Horváth, VisualLabs Human skills Argues for the continued importance of human facilitation and social skill. 1,673 50.7
200 The Prompt is the Platform - Dominik Tornow, Resonate HQ Prompt platforms Frames prompts as a platform layer for AI applications and workflows. 1,665 50.5
201 Using RL Agent to Detect and Remediate ETL Pipeline Failures - Anna Marie Benzon RL for data pipelines Shows reinforcement-learning agents applied to ETL failure detection and remediation. 1,656 50.2
202 Privacy-Preserving Intelligence — Steve Korshakov, Bee (acq. Amazon) Privacy-preserving AI Covers building intelligent systems that protect user privacy. 581 48.4
203 Production Evals For Agentic AI Systems - Nishant Gupta, Meta Superintelligence Labs Agent evals Covers building production-grade evaluations for agentic AI systems. 1,782 48.2
204 Browser Agents Don’t Need Better Models. They Need Better Eyes. - Kushan Raj, ARK Browser agents Argues browser agents are bottlenecked by perception, not model quality. 1,627 47.9
205 Running a Chess YouTube Channel entirely by AI — Stephan Steinfurt, TNG AI media automation Describes using AI to automate an entire chess YouTube channel workflow. 1,110 46.2
206 Security Track Intro — Randall Degges, Snyk Security track intro Opens the security track with an overview of AI security themes. 552 46.0
207 Agentic Development Security — Ezra Tanzer, Snyk Agentic development security Covers securing development workflows that rely on AI agents. 547 45.6
208 Voice Agents That Handle Interrupts - Chintan Agrawal and Daniel Wirjo, AWS Voice agents Covers how to build voice agents that gracefully handle user interruptions in real time. 516 43.0
209 Voice In, Visuals Out: The Agony and the Ecstasy - Allen Pike, Forestwalk Labs Multimodal UX Covers challenges and promise in voice-to-visual AI workflows. 1,378 40.5
210 When Agents Meet Physical Data: The Other Physics of Agent Harnesses - Dmitry Petrov, DataChain Agent data infrastructure Dmitry Petrov explores what happens when agent harnesses meet large, physical datasets. 483 40.2
211 Respect The Process - Andrew Dumit, Watershed Technology Inc. Engineering process Emphasizes process discipline in AI-enabled engineering workflows. 962 38.5
212 The 100-Tool Agent Is a Trap - Sohail Shaikh & Ankush Rastogi, Prosodica Agent tool design Argues loading an agent with too many tools backfires. 1,278 37.6
213 AI-Driven Multi-Document Correlation for Financial Compliance - Varsha Shah, Independent Compliance AI Applies AI to correlating multiple documents for financial compliance workflows. 1,230 36.2
214 OpenClaw in Your Hand: Building a Physical AI Terminal - Lech Kalinowski, Callstack Physical AI devices Shows a handheld or physical AI terminal concept built around OpenClaw. 1,223 36.0
215 GPU Cloud Deployment Without Leaving Your IDE — Audry Hsu, RunPod GPU deployment tooling Audry Hsu demos deploying to GPU cloud directly from the IDE. 1,864 35.2
216 From Transcription to Live Music: Gemini’s Audio Stack — Thor Schaeff, Google DeepMind Audio AI Thor Schaeff walks through Gemini’s audio stack from transcription to live music. 1,856 35.0
217 The Miranda Hypothesis: How Hamilton Poisoned Persona Evals - Jacob E. Thomas, Results Gen Persona evals Examines how a flawed persona corrupted evaluation results. 1,244 33.6
218 Sovereign Escape Velocity: Ownership w Open Models — Gus Martins, & Ian Ballantyne, Google DeepMind Open models Google DeepMind advocates for owning your AI stack through open models. 1,729 33.2
219 Agents Need Receipts, Not More Tool Calls - Armanas Povilionis, Alithea Bio Agent auditability Argues agents should produce verifiable receipts of their actions rather than simply making more tool calls. 372 31.0
220 Your LLM Deception Monitor Is Broken. The Fix Is in the Training Data - Sachin Kumar, LexisNexis Safety monitoring Argues that deception monitoring depends heavily on training-data quality. 711 29.6
221 User Signal Dies at the Retrieval Boundary - Sonam Pankaj, StarlightSearch Retrieval signals Discusses preserving user intent and signals across retrieval system boundaries. 945 27.8
222 When All Context Matters: Extended Cache Augmented Generation - Luis Romero-Sevilla, Orbis Cache-augmented generation Covers using extended cache techniques when broad context matters. 910 26.8
223 Research to Reality: Bringing Frontier ML Research to Production - Vaidas Razgaitis, Higharc ML productionization Discusses translating frontier ML research into production systems. 865 25.4
224 Bypassing the Multimodal Tax: Hybrid RAG, SQL RRF & UI Telemetry - Abed Matini, Ogilvy Hybrid retrieval Combines RAG, SQL, reciprocal-rank fusion, and UI telemetry for multimodal workflows. 862 25.4
225 Medic for Apache Spark - First Aid for Failing Jobs - Drasko Profirovic, Pinterest Data pipeline debugging Drasko Profirovic presents a tool that diagnoses and repairs failing Apache Spark jobs. 299 24.9
226 Using LLMs to Secure Source Code — Eugene Yan, Anthropic Code security Eugene Yan shows how LLMs can find and fix source-code vulnerabilities. 334 22.3

Code to Replicate the Ranking

This script uses yt-dlp to fetch the channel’s current video metadata, filters to recent posted conference talks, and ranks by views per day. Install yt-dlp first if needed:

uv tool install yt-dlp

Then run:

import datetime as dt
import json
import subprocess

CHANNEL_URL = "https://www.youtube.com/@aiDotEngineer/videos"
REFERENCE_DATE = dt.date(2026, 8, 1)
PLAYLIST_END = 250

EXCLUDE_TITLE_SUBSTRINGS = [
    "Vibe Reel",
    "Things to Know about AIE",
]

cmd = [
    "yt-dlp",
    "--playlist-end",
    str(PLAYLIST_END),
    "--skip-download",
    "--ignore-errors",
    "--print",
    "%()j",
    CHANNEL_URL,
]

proc = subprocess.run(
    cmd,
    check=False,
    text=True,
    capture_output=True,
    timeout=420,
)

rows = []
for line in proc.stdout.splitlines():
    if not line.startswith("{"):
        continue

    video = json.loads(line)
    title = video.get("title") or ""
    upload_date = video.get("upload_date")
    view_count = video.get("view_count")

    if not upload_date or view_count is None:
        continue
    if upload_date < "20260608":
        continue
    if any(skip in title for skip in EXCLUDE_TITLE_SUBSTRINGS):
        continue

    uploaded = dt.datetime.strptime(upload_date, "%Y%m%d").date()
    days_since_upload = max(1, (REFERENCE_DATE - uploaded).days)
    views_per_day = view_count / days_since_upload

    rows.append(
        {
            "title": title,
            "url": video.get("webpage_url")
            or f"https://www.youtube.com/watch?v={video.get('id')}",
            "upload_date": uploaded.isoformat(),
            "views": view_count,
            "days_since_upload": days_since_upload,
            "views_per_day": views_per_day,
            "duration_minutes": round((video.get("duration") or 0) / 60, 1),
        }
    )

rows.sort(key=lambda row: row["views_per_day"], reverse=True)

print("| Rank | Video | Uploaded | Views | Days | Views/day |")
print("|---:|---|---:|---:|---:|---:|")

for rank, row in enumerate(rows, start=1):
    print(
        "| {rank} | [{title}]({url}) | {upload_date} | "
        "{views:,} | {days_since_upload} | {views_per_day:,.1f} |".format(
            rank=rank,
            **row,
        )
    )

A couple of notes:

  • --ignore-errors lets the script skip scheduled premieres that do not have watchable metadata yet.
  • PLAYLIST_END = 250 was enough for this snapshot because the recent conference batch appeared in the first 250 channel videos.
  • The views are a snapshot. Re-running the script later will produce different view counts and a different views/day ranking.