What to Watch from AI Engineer World’s Fair 2026
AI Engineer World’s Fair 2026 just wrapped in San Francisco, and the AI Engineer YouTube channel has started posting talks. I wanted a quick way to decide what to watch first, so I pulled the current video metadata and ranked the posted talks by views per day. Note this post and all of the code associated with grabbing the data was written by GPT 5.5 (High) using the Codex harness.
The raw view count alone is a little unfair because the uploads are staggered. A talk posted on June 9 has had a month to accumulate views; a talk posted on July 8 has had about a day. So the main ranking below is:
views per day = views / max(1, days since upload)
As of August 1, 2026 at 9:54 PM Pacific, the early winners are:
- Why Agentic Systems Need Ontologies — Frank Coyle, UC Berkeley
- Field Guide to Fable — Thariq Shihipar, Anthropic
- Harness Engineering is not Enough: Why Software Factories Fail — Dex Horthy, HumanLayer
- Understanding is the new bottleneck — Geoffrey Litt, Notion
- “The engineer of the future is the person who is able to choose what is worth doing.” — Addy Osmani
Method
I included 226 posted conference-talk videos from June 8, 2026 onward. I excluded the 2026 vibe reel, the “6 Things to Know” preview, and scheduled premieres that did not have public view counts yet.
The topic and summary fields are lightweight notes derived from the title, not transcript-level summaries. The popularity ranking is the reproducible part: YouTube metadata, upload date, views, and views per day.
AI Engineer World’s Fair 2026 ran June 29 to July 2, 2026 at Moscone West in San Francisco, according to the conference schedule.
Ranked Watchlist
| Rank | Video | Topic | Summary | Views | Views/day |
|---|---|---|---|---|---|
| 1 | Why Agentic Systems Need Ontologies — Frank Coyle, UC Berkeley | Ontologies for agents | Frank Coyle argues that agentic systems need explicit ontologies to reason reliably over structured knowledge. | 184,335 | 20,481.7 |
| 2 | Field Guide to Fable — Thariq Shihipar, Anthropic | AI-native product design | A guide to Fable and the product/design patterns behind building with modern AI systems. | 139,374 | 5,360.5 |
| 3 | Harness Engineering is not Enough: Why Software Factories Fail — Dex Horthy, HumanLayer | Agent harness engineering | Dex Horthy argues that harness engineering alone can’t make “software factory” approaches to AI development succeed. | 44,786 | 4,976.2 |
| 4 | Understanding is the new bottleneck — Geoffrey Litt, Notion | Code understanding | Argues that understanding what AI-written systems do, not producing code, is the new limiting factor. | 102,348 | 4,652.2 |
| 5 | “The engineer of the future is the person who is able to choose what is worth doing.” — Addy Osmani | Engineering judgment | Addy Osmani argues that choosing what is worth building matters more than raw execution speed. | 74,278 | 4,126.6 |
| 6 | From fork() to Fleet: Designing an Agent Sandbox Cloud — Abhishek Bhardwaj, OpenAI | Agent sandboxing | Abhishek Bhardwaj describes designing a cloud sandbox system for safely running many agent processes at scale. | 71,943 | 3,786.5 |
| 7 | Building Great Agent Skills: The Missing Manual | Agent skills | Practical guidance for designing reusable skills that make agents more capable and dependable. | 121,403 | 3,678.9 |
| 8 | AI’s Jurassic Park Period — Aaron Stanley, dbt Labs | AI risk | Compares the current AI moment to Jurassic Park’s overreach and draws out the lessons. | 44,099 | 3,674.9 |
| 9 | How Forward Deployed Engineering is done at Cognition — Jia Wu | Forward-deployed engineering | Jia Wu describes how forward-deployed engineers work with customers at Cognition. | 14,467 | 3,616.8 |
| 10 | Everything we knew about software has changed — Theo Browne, @t3dotgg | AI product strategy | Theo Browne frames what is worth building now that AI has changed developer workflows and product expectations. | 85,935 | 3,580.6 |
| 11 | Building Closed-Loop Evals for a Multimodal Agent at Scale — Soumya Gupta & Jai Chopra, Uber | Multimodal agent evals | Uber engineers describe building closed-loop evaluations for a multimodal agent operating at scale. | 28,072 | 3,509.0 |
| 12 | The Golden Age of AI Engineering — Alexander Embiricos & Romain Huet & Peter Steinberger, OpenAI | AI engineering keynote | OpenAI leaders discuss why this is a golden age for engineers building with frontier models. | 72,886 | 3,169.0 |
| 13 | Every company should have a Brain — Garry Tan, Y Combinator | AI business strategy | Garry Tan argues that AI is rewriting the fundamental economics and dynamics of building a business. | 47,428 | 3,161.9 |
| 14 | Forward Deployed Engineering 101 — Kevin Bai, Anthropic, ex Palantir & Rippling Founding FDE | Forward-deployed engineering | Kevin Bai introduces how forward-deployed engineers work with customers at Anthropic. | 12,458 | 3,114.5 |
| 15 | Full Workshop: Setting Yourself Up for Success —Jason Liu, OpenAI Codex | Coding agent workshop | Jason Liu runs a full workshop on setting yourself up for success with OpenAI Codex. | 23,946 | 2,993.2 |
| 16 | Don’t Ship Skills Without Evals — Philipp Schmid, Google DeepMind | Skill evals | Philipp Schmid argues that agent skills need real evaluation coverage before they ship to production. | 53,503 | 2,972.4 |
| 17 | Thinner Agents on a Smarter Substrate: The Ontology-based Semantic Layer — Emil Eifrem, Neo4j | Ontology semantic layer | Emil Eifrem makes the case for thinner agents built on a smarter, ontology-based semantic substrate. | 25,157 | 2,515.7 |
| 18 | AI tools for Forward Deployed Engineering — Vasuman Moza, Varick Agents | Forward-deployed engineering | Vasuman Moza surveys the AI tooling that forward-deployed engineers rely on day to day. | 9,155 | 2,288.8 |
| 19 | Why We Killed Our Multi-Agent Pipeline — Subbiah Sethuraman and Abhilash Asokan, ZS Associates | Multi-agent lessons | A candid account of why ZS Associates scrapped their multi-agent pipeline and what replaced it. | 17,217 | 1,913.0 |
| 20 | Ending AI Slop — Thais Castello Branco, Taste Labs | AI content quality | Thais Castello Branco makes the case for raising the taste bar and ending AI slop. | 1,760 | 1,760.0 |
| 21 | fighting slop with slop — Vaibhav Gupta, Boundary | AI content quality | Vaibhav Gupta explores using generated output to detect and clean up other generated slop. | 1,737 | 1,737.0 |
| 22 | Loop Engineering from First Principles — Kyle Mistele, HumanLayer | Agent loop design | Kyle Mistele derives agent loop design from first principles rather than copying existing harnesses. | 12,047 | 1,721.0 |
| 23 | Claude for Long-Horizon Tasks — Lance Martin, Anthropic | Long-horizon agents | Lance Martin covers patterns for keeping Claude on track across long, multi-step tasks. | 15,148 | 1,514.8 |
| 24 | Reinforcement Learning without Verifiable Rewards — Will Brown, Prime Intellect | Reinforcement learning | Will Brown covers reinforcement learning approaches when rewards cannot be cleanly verified. | 1,448 | 1,448.0 |
| 25 | Persona Engineering: A Field Guide to AI Synthetic Personas — Ishan Anand, InsightSciences.ai | Synthetic personas | Ishan Anand offers a field guide to engineering AI synthetic personas that behave believably. | 4,244 | 1,414.7 |
| 26 | Learning on the Job: The Future of Post-Training — Raymond Feng, Applied Compute | Post-training | Raymond Feng argues models should keep learning on the job after deployment. | 1,350 | 1,350.0 |
| 27 | In the Land of AI Agents, the Verifiers Are King — Tariq Shaukat, Sonar | Agent verification | Makes the case that verification, not generation, is the decisive capability for reliable agents. | 15,976 | 1,331.3 |
| 28 | The Future Is Domain-Specific Agents - Justin Schroeder, StandardAgents | Domain agents | Argues for agents specialized around concrete domains rather than generic assistants. | 43,421 | 1,315.8 |
| 29 | CrabRAG: Why Automated Assistants Need Graph Memory, Not More Tokens — Stephen Chin, Neo4j | Graph memory for RAG | Stephen Chin argues automated assistants need graph-based memory rather than ever-larger token windows. | 12,432 | 1,243.2 |
| 30 | Teaching AI to Find Real Vulnerabilities — Prof. David Brumley, Bugcrowd | Security research | David Brumley shows how to teach AI systems to find real software vulnerabilities. | 1,159 | 1,159.0 |
| 31 | State of the Union: Why Local, Why Now — NVIDIA, Osmantic, Roboflow, EXO Labs, @matthew_berman | Local AI | A multi-vendor panel on the case for running AI locally rather than in the cloud. | 23,947 | 1,140.3 |
| 32 | Your Finance Agent’s Bottleneck Is You — Ramana Siddanth Emani, Auditoria AI | Finance agents | Argues the human reviewer, not the model, is the real bottleneck in finance agent workflows. | 2,226 | 1,113.0 |
| 33 | Agents at Scale: Inside MiniMax’s Model and the Infrastructure Behind It — Olive Song | Model infrastructure | Olive Song covers MiniMax’s model and the infrastructure that runs its agents at scale. | 1,032 | 1,032.0 |
| 34 | Forward Deployed Engineering at Cursor — Pauline Brunet | Forward-deployed engineering | Pauline Brunet describes how forward-deployed engineers work directly alongside customers at Cursor. | 17,954 | 997.4 |
| 35 | From Writing Code to Designing Systems: How the Developer Role is Changing — Chris Noring, Microsoft | Developer role shift | Discusses how the developer’s role is moving from writing code to designing broader systems. | 20,908 | 995.6 |
| 36 | Stop Making Models Bigger, Make Them Behave — Kobie Crawford, Snorkel | Model behavior tuning | Kobie Crawford argues for improving model behavior through better data rather than simply scaling model size. | 50,806 | 977.0 |
| 37 | HTML Is All Agents Need — James Russo, HeyGen | Agent-generated graphics | Argues HTML is the right substrate for agents to produce visual and interface output. | 10,514 | 955.8 |
| 38 | First Steps Toward Automated AI Research — Richard Socher, CEO Recursive AI | Automated AI research | Richard Socher lays out early steps toward AI systems that conduct their own research. | 1,867 | 933.5 |
| 39 | Skills are new features: Building Skill-Centric Harness — Yogendra Miraje, FactSet | Agent skills | Yogendra Miraje argues skills are the new feature unit and shows how to build a skill-centric harness. | 2,800 | 933.3 |
| 40 | The Great Loops Debate — Dex Horthy, Geoff Huntley, Ian Livingstone, Greg Pstrucha, @insecure-agents | Agent loop design | A panel debates how the control loops driving AI agents should be structured. | 13,862 | 924.1 |
| 41 | How Forward Deployed Engineering is done at Factory — Eno Reyes | Forward-deployed engineering | Eno Reyes describes how forward-deployed engineering works at Factory. | 3,580 | 895.0 |
| 42 | What’s Next After RLHF? — Diogo Almeida, TypeSafe AI | Post-training | Diogo Almeida looks at what comes after RLHF for aligning models. | 890 | 890.0 |
| 43 | Claude Fable, Claude Tag, and Anthropic’s Culture — Cat Wu & Thariq Shihipar ft Simon Willison | Fireside conversation | Simon Willison talks with Anthropic’s Cat Wu and Thariq Shihipar about building with AI today. | 14,175 | 833.8 |
| 44 | Why Large? Tiny LMs & Agents on Edge/Robotics — Cormac Brick, Google | Small models on edge | Cormac Brick makes the case for tiny language models and agents running on edge and robotics hardware. | 5,771 | 824.4 |
| 45 | State of Data — Sean Cai, Independent / State of Data | State of data | Sean Cai surveys where the data ecosystem stands for AI practitioners today. | 4,832 | 805.3 |
| 46 | Scaling to Long Horizons — Ross Taylor & Chengxi Taylor, General Reasoning | Long-horizon agents | Ross and Chengxi Taylor cover scaling agent training to long-horizon tasks. | 777 | 777.0 |
| 47 | Special Topics in Kernels, RL, Reward Hacking in Agents — Daniel Han, Unsloth | RL and reward hacking | Daniel Han covers kernels, reinforcement learning, and reward-hacking pitfalls in training agents. | 11,446 | 763.1 |
| 48 | The Dirty Secret of Forward Deployed Engineering — Natalie Meurer, Sierra | Forward-deployed engineering | Natalie Meurer shares the unglamorous realities behind forward-deployed engineering work. | 3,041 | 760.2 |
| 49 | AI Agents for Performance: Ship Faster, Pay Less — Rajat Shah, Netflix | Performance agents | Rajat Shah describes how Netflix uses AI agents to ship performance improvements faster and at lower cost. | 2,978 | 744.5 |
| 50 | Rethinking Environments for Long-Horizon Work — Rayan Garg, Theta Software | Agent environments | Rayan Garg rethinks the environments agents need for long-horizon work. | 739 | 739.0 |
| 51 | “Software engineering is not about writing code” — Benoit Schillings, Google DeepMind VP of Research | Engineering judgment | Benoit Schillings argues software engineering is about problem-solving, not writing code. | 10,965 | 731.0 |
| 52 | Emulated: The Data for Fully Autonomous Software Engineers and Companies — Joseph Wang | Training data | Joseph Wang discusses the data needed for fully autonomous software engineers. | 719 | 719.0 |
| 53 | Stop Burning Tokens: Why self-improvement needs domain expertise first - Annabell Schäfer, Langfuse | Self-improving agents | Argues agent self-improvement needs domain expertise first before it can stop wasting tokens. | 9,630 | 687.9 |
| 54 | WTF Is the Context Layer? The Missing Infrastructure for Production Agents — Prukalpa Sankar | Context infrastructure | Prukalpa Sankar defines the context layer as the missing infrastructure production agents actually need. | 11,990 | 666.1 |
| 55 | Verifiable Environments for AI in Biology — Kenny Workman, LatchBio | Scientific agents | Kenny Workman presents verifiable environments for applying AI to biology. | 658 | 658.0 |
| 56 | The Base Model Is Dead — Varun Singh, Arcee AI | Model training | Varun Singh argues the plain base model is no longer the useful unit. | 621 | 621.0 |
| 57 | Active Graph Agent Runtime (BabyAGI 4) — Yohei Nakajima, Untapped Capital | Graph agent runtime | Yohei Nakajima introduces an active graph-based agent runtime in the latest iteration of BabyAGI. | 6,132 | 613.2 |
| 58 | SimulationMaxxing: How we ship agents 20× faster — Aman Gupta (Nubank) + Shreya Rajpal (Snowglobe) | Agent simulation | Nubank and Snowglobe engineers describe using simulation to ship agents dramatically faster. | 1,810 | 603.3 |
| 59 | RAG is dead, right?? — Kuba Rogut, Turbopuffer | RAG relevance | Kuba Rogut examines whether retrieval-augmented generation is still necessary as context windows grow. | 31,824 | 600.5 |
| 60 | Benchmarks: The Good, the Bad, and the Ugly — Ali Khial, G2i | AI benchmarking | Ali Khial reviews what benchmarks measure well, measure poorly, and get outright wrong. | 600 | 600.0 |
| 61 | The Production AI Playbook: Deploying Agents at Enterprise Scale — Sandipan Bhaumik, Databricks | Enterprise agent deployment | Sandipan Bhaumik shares a playbook for deploying AI agents at enterprise scale. | 26,058 | 592.2 |
| 62 | Citation Needed: Provenance for LLM-Built Knowledge Graphs — Daniel Chalef, Zep AI | Knowledge-graph provenance | Daniel Chalef covers tracking provenance and citations in LLM-constructed knowledge graphs. | 5,260 | 584.4 |
| 63 | Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI | Data quality | Ari Morcos argues data quality multiplies the value of every unit of compute. | 583 | 583.0 |
| 64 | Your Attention Is the Bottleneck, Not Your Agents — Zack Proser, WorkOS | Human attention limits | Zack Proser argues the real bottleneck in agent workflows is human attention, not the agents. | 29,641 | 581.2 |
| 65 | Let’s integrate AI Agents in Event-Sourced Systems — Divakar Kumar, FlyersSoft | Event-sourced agents | Divakar Kumar shows how to integrate AI agents into event-sourced system architectures. | 1,152 | 576.0 |
| 66 | A Practitioner’s Guide to Graphs - Tim Ainge, Good Collective | Graphs for AI | A practitioner’s guide to using graph structures in AI applications. | 7,970 | 569.3 |
| 67 | We Cut 94% of AI Coding Tokens With a Local Code Index - Rajkumar Sakthivel, Tesco | Code indexing | Shows how local code indexing can sharply reduce token usage in AI coding workflows. | 19,129 | 562.6 |
| 68 | Your Moat Is Your Data Model — Mike Phipps, Gates Foundation | Data model as moat | Mike Phipps argues a company’s data model, not its models or prompts, is its durable competitive moat. | 5,312 | 531.2 |
| 69 | Turn 10,994 Notes Into Memory - Paul Iusztin, Decoding AI & Louis-François Bouchard, Towards AI | Agent memory | Shows how to turn a large personal note corpus into usable agent memory. | 17,997 | 499.9 |
| 70 | Your agent architecture has a half-life of 6 months — Dan Farrelly, CTO, Inngest | Agent architecture | Argues agent architectures decay fast and must be designed to be replaced as the field moves. | 5,495 | 499.5 |
| 71 | The Unreasonable Effectiveness of Separating the Task from the Model — Maxime Rivest & Isaac Miller | Task-model separation | Makes the case that separating the task definition from the underlying model yields outsized gains. | 4,248 | 472.0 |
| 72 | Data and Environment Curation for Post-Training LLMs — Mahesh Sathiamoorthy, Bespoke Labs | Post-training data | Mahesh Sathiamoorthy covers curating data and environments for post-training LLMs. | 467 | 467.0 |
| 73 | Evals-Driven Development for a Mental Health AI Coach — Akele Reed & Dave Revere, SonderMind | Healthcare evals | Akele Reed and Dave Revere describe evals-driven development for a mental health AI coach. | 3,227 | 461.0 |
| 74 | Recursive Model Improvement — Lee Robinson, Cursor, SpaceXAI | Recursive model improvement | Lee Robinson discusses using models to recursively improve themselves and the tools around them. | 7,596 | 446.8 |
| 75 | Every Harness Will Become A Claw — Sam Bhagwat, Mastra | Agent harnesses | Argues every agent harness is converging toward a Claw-like general-purpose tool. | 4,879 | 443.5 |
| 76 | From Agent Traces to Agent Simulations — Rustem Feyzkhanov, Snorkel AI | Agent simulation | Rustem Feyzkhanov describes moving from recorded agent traces to full agent simulations. | 3,012 | 430.3 |
| 77 | DeepSWE: A Contamination-Resistant Coding Benchmark — James Shi, Datacurve | Coding benchmarks | James Shi introduces DeepSWE, a coding benchmark designed to resist training-data contamination. | 2,544 | 424.0 |
| 78 | How Forward Deployed Engineering is done at Kepler — Vinoo Ganesh | Forward-deployed engineering | Vinoo Ganesh describes how forward-deployed engineering is practiced at Kepler. | 1,678 | 419.5 |
| 79 | Build for the Memo, Not the Demo — Shawn Chan, China Resources Holdings | AI product strategy | Shawn Chan argues teams should build for the decision memo rather than the flashy demo. | 836 | 418.0 |
| 80 | Modern Post-Training: A Deep Dive — Will Brown, Prime Intellect | Open AI infrastructure | Will Brown presents the Prime Intellect stack for open, decentralized AI training and infrastructure. | 7,856 | 413.5 |
| 81 | HTML is All You Need (for Agents to Make Graphics) - Amol Kapoor, Nori | Agent-generated graphics | Argues that HTML is a strong substrate for agents producing visual outputs. | 13,985 | 411.3 |
| 82 | We Vetted 2000 AI Skills Before They Reached Developers — Lucas Palma, Nubank | Skill vetting | Lucas Palma describes vetting two thousand AI skills before releasing them to developers. | 1,227 | 409.0 |
| 83 | Build Evals That Actually Matter - Nick Ung & Akshay Sharma, Lyft | AI evals | Nick Ung shares how to design evaluations that meaningfully measure agent quality instead of vanity metrics. | 5,263 | 404.8 |
| 84 | “The biggest challenge in your stack? Evals, Evals, Evals” - 2026 State of AI Engineering results | State of AI engineering | A data-driven survey of where AI engineering stands in 2026. | 4,398 | 399.8 |
| 85 | How Kepler Built Verifiable AI for Financial Services — Vinoo Ganesh | Verifiable AI in finance | Vinoo Ganesh explains how Kepler built verifiable AI systems for financial services. | 1,181 | 393.7 |
| 86 | What if the harness mattered more than the model? - Aditya Bhargava, Etsy | Evaluation harnesses | Makes the case that the surrounding harness can matter more than model choice for real results. | 9,824 | 393.0 |
| 87 | Your Agent Didn’t Fail. Your Harness Did. — Vinoth Govindarajan, OpenAI | Agent harness design | Vinoth Govindarajan argues most agent failures trace back to the harness, not the model. | 1,151 | 383.7 |
| 88 | The Messy Reality of Scale: Synthetic Data and Pre-Training — Marah Abdin & Robert McHardy, poolside | Synthetic data at scale | poolside shares the messy realities of synthetic data generation and pre-training at scale. | 2,206 | 367.7 |
| 89 | Design Patterns for AI Trust: Juries, Libraries, and Agent Tiers — Alex Bauer, Upside.tech | AI trust patterns | Presents design patterns like juries, libraries, and agent tiers for building trustworthy AI systems. | 7,612 | 362.5 |
| 90 | The Future of Evals: From LLM as a Judge to Agent as a Judge — Aparna Dhinakaran, Arize AI | Agent-as-judge evals | Aparna Dhinakaran traces the shift from LLM-as-judge to agent-as-judge evaluation. | 2,782 | 347.8 |
| 91 | Imagination Engineering: “Live in the future and then build what’s missing.” | AI product design | Eve Bouffard explores how AI expands what designers can imagine and prototype. | 5,350 | 334.4 |
| 92 | Morgan Stanley’s ALPHALAB: Multi-Agent Research Across Optimization Domains — Brendan Rappazzo | Multi-agent research | Brendan Rappazzo presents Morgan Stanley’s ALPHALAB multi-agent system for optimization research. | 958 | 319.3 |
| 93 | Beyond the Harness: A Journey Towards Adaptative Engineering - Rajiv Chandegra, Annicha Labs | Adaptive engineering | Explores engineering systems that adapt beyond static evaluation harnesses. | 7,704 | 308.2 |
| 94 | From Signal to PR: Anatomy of a Self-Improving Agent — Jason Lopatecki, Arize | Self-improving agents | Jason Lopatecki dissects a self-improving agent that turns production signals into pull requests. | 2,430 | 303.8 |
| 95 | How Forward Deployed Engineering is done at Ramp — Leo Mehr | Forward-deployed engineering | Leo Mehr describes how forward-deployed engineering works at Ramp. | 1,204 | 301.0 |
| 96 | The UX of AI: Making AI-Powered Apps Your Users Don’t Hate - Kathryn Grayson Nanz, Progress Software | AI UX design | Offers UX guidance for building AI-powered apps that users actually enjoy using. | 3,945 | 281.8 |
| 97 | “I’ve never seen anything scarier than an LLM with tool calls.” — Erik Meijer aka @HeadinTheBox | Formal verification for agents | Erik Meijer argues for grounding agent-written code in formal proofs rather than trusting agent behavior directly. | 5,233 | 275.4 |
| 98 | Recursive Coding Agents - Raymond Weitekamp, OpenProse | Recursive coding agents | Explores coding agents that recursively decompose and improve their own work. | 10,029 | 271.1 |
| 99 | Serving 2 Million Models Without Melting: Scaling the Hugging Face Hub — Arek Borucki, Hugging Face | ML infrastructure at scale | Arek Borucki explains how the Hugging Face Hub serves two million models without falling over. | 1,081 | 270.2 |
| 100 | How Evals and Prompts Shape Agent Behavior — Preetika Bhateja & Daniel Bump, YouTube Ads | Evals and prompting | YouTube Ads engineers show how evaluations and prompts jointly shape agent behavior. | 2,053 | 256.6 |
| 101 | The agent-ready web: Simplify user actions with WebMCP — Tara Agyemang, Google | Agent-ready web | Tara Agyemang introduces WebMCP for making web actions easier for agents to perform. | 12,196 | 239.1 |
| 102 | The Desktop Frontier — Ahmad Osman, Osmantic | Desktop AI | Explores running capable AI on the desktop as the next frontier. | 2,525 | 229.5 |
| 103 | Computer-Use 2.0: Agents Just Got Multi-Cursor — Francesco Bonacci, Cua | Computer-use agents | Francesco Bonacci introduces multi-cursor computer-use agents that can act across several interfaces at once. | 3,835 | 225.6 |
| 104 | How Forward Deployed Engineering is done at Decagon — Sunny Rekhi | Forward-deployed engineering | Sunny Rekhi describes how forward-deployed engineering is done at Decagon. | 895 | 223.8 |
| 105 | Why Off-the-Shelf AI Doesn’t Understand Money — Udi Menkes, Intuit | Financial domain AI | Udi Menkes argues general-purpose AI lacks the domain grounding that money-related tasks demand. | 671 | 223.7 |
| 106 | Build Systems, Not Code - Angie Jones, Agentic AI Foundation | Autonomous engineering | Angie Jones argues engineers should build systems rather than hand-write code in the agent era. | 7,716 | 208.5 |
| 107 | Evaling Video Slop — Maor Bril, Character.ai | Video generation evals | Maor Bril covers how to evaluate low-quality AI-generated video at scale. | 1,447 | 206.7 |
| 108 | AI on Your Lakehouse: Context Comes in Shapes, Not Queries — Zach Blumenfeld, Neo4j | Lakehouse context | Zach Blumenfeld argues AI context on a lakehouse comes in graph shapes, not flat queries. | 1,798 | 199.8 |
| 109 | Using Spec-Driven Development for Production Workflows - Erik Hanchett, AWS | Spec-driven development | Shows how specifications can drive more reliable AI-assisted production workflows. | 6,721 | 197.7 |
| 110 | Training Frontier Models to Out-Think Hackers — Uri Rolls, Arithmetic & Thom Wolf, Hugging Face | Security-focused training | Covers training frontier models to anticipate and outmaneuver attackers. | 1,577 | 197.1 |
| 111 | Wearing the Agent: From Group Chats to Glasses — Sai Krishna Rallabandi | Wearable agents | Sai Krishna Rallabandi traces agents moving from group chats onto wearable glasses. | 575 | 191.7 |
| 112 | Perception Agents — Antje Barth, Amazon AGI Lab | Perception agents | Antje Barth introduces agents focused on perception across multimodal inputs. | 1,724 | 191.6 |
| 113 | Notion’s Token Town — Sarah Sachs, Notion | LLM token economics | Sarah Sachs walks through how Notion reasons about token usage and cost at scale. | 1,722 | 191.3 |
| 114 | Should AI Engineers Still Read Code in 2026? The Z/L Continuum — Alex Volkov, ThursdAI | AI coding practice | Asks how much code AI engineers should still read as agents write more of it. | 4,204 | 191.1 |
| 115 | The Pipeline Is Dead - Iris ten Teije, Sky Valley Ambient Computing | Agentic workflows | Reframes traditional software pipelines around more dynamic, ambient AI workflows. | 4,741 | 189.6 |
| 116 | AI System Design: From Idea to Production - Apoorva Joshi, MongoDB | AI system design | Walks through moving an AI system from concept to production. | 6,345 | 186.6 |
| 117 | How Autoresearch is changing ML research — Zhengyao Jiang, Weco | Autonomous coding agents | Zhengyao Jiang recounts how an AI agent topped the leaderboard of OpenAI’s hiring challenge. | 2,909 | 181.8 |
| 118 | From Blind Spots to Merged PRs: Continuous Agentic Performance Optimization - May Walter, Hud | Agentic optimization | May Walter covers continuously optimizing agent performance from blind spots through to merged pull requests. | 2,254 | 173.4 |
| 119 | Why Eval++ Is the Next Great Compute Primitive — Sunil Pai & Matt Carey, Cloudflare | Evaluation infrastructure | Cloudflare engineers argue that richer evaluation is becoming a core compute primitive. | 9,315 | 172.5 |
| 120 | Teaching Coding Agents to do Spreadsheets - Nuno Campos, Witan Labs | Spreadsheet agents | Covers how to make coding agents useful for spreadsheet-heavy work. | 4,105 | 171.0 |
| 121 | Stop AI Agent Hallucinations: 5 Techniques + Production Patterns - Elizabeth Fuentes, AWS | Agent hallucinations | Presents five production-tested techniques for reducing AI agent hallucinations. | 3,580 | 170.5 |
| 122 | Building an Autonomous Engineering Org - Angie Jones, Agentic AI Foundation | Autonomous engineering | Discusses the organizational patterns behind autonomous AI-assisted engineering. | 5,783 | 170.1 |
| 123 | Building an ACP-Compatible Agent Live — Bennet Fenner, Zed | Agent protocols | A live build showing how an agent can integrate with the Agent Client Protocol ecosystem. | 4,074 | 169.8 |
| 124 | Why MCP and ChatGPT Apps Use Double Iframes — Frédéric Barthelet, Alpic | MCP app architecture | Frederic Barthelet explains the double-iframe pattern behind MCP and ChatGPT apps. | 7,882 | 167.7 |
| 125 | Build AI Systems for Discernment, Not Approval - Angel Ortmann Lee, Duolingo | AI product judgment | Advocates building AI systems that help users judge well instead of merely agreeing with them. | 4,185 | 167.4 |
| 126 | Why Can’t Anyone Answer Questions About the Business? — Garrett Galow, WorkOS | Business intelligence agents | Garrett Galow examines why business questions stay hard to answer and how AI can help. | 8,354 | 163.8 |
| 127 | Your Agent’s Biggest Lie: “I Searched the Web” — Rafael Levi, Bright Data | Agent web access | Rafael Levi exposes how often agents fake web searches and how to give them real access. | 7,206 | 160.1 |
| 128 | Why More Context Makes Your Agent Dumber and What to Do About It — Nupur Sharma, Qodo | Context management | Nupur Sharma explains why piling on context degrades agents and how to curate it instead. | 8,483 | 157.1 |
| 129 | Full Workshop: Better Auth — Paola Estefania, Better Auth | Agent authentication | Covers authentication patterns purpose-built for AI agents. | 1,710 | 155.5 |
| 130 | The Agentic AI Engineer - Benedikt Sanftl, Mutagent | Agentic engineering | Defines what changes when AI engineers work through agentic systems. | 5,048 | 153.0 |
| 131 | MCP Apps: Primitives, discovery, and the Future of Software - Pietro Zullo, Manufact, Inc | MCP apps | Covers MCP app primitives, discovery, and how software interfaces may evolve. | 4,055 | 150.2 |
| 132 | Everything Is a Rollout — Alex Shaw + Ryan Marten, Terminal-Bench, Harbor, Laude Institute | Agent benchmarking | Reframes agent evaluation and training around the idea that everything is a rollout. | 1,140 | 142.5 |
| 133 | Why Your Agent Disagrees With Itself (And What To Do About It) - Diane Lin, Datadog | Agent consistency | Diane Lin examines why an agent can contradict its own earlier answers and how to make its behavior consistent. | 1,707 | 142.2 |
| 134 | Skills are the New SDKs - Elvin Aghammadzada, DataRobot | Agent skills | Elvin Aghammadzada argues that reusable skills are becoming the new packaging unit for agent capabilities, replacing SDKs. | 1,707 | 142.2 |
| 135 | Every Solo Agent Builder Eventually Reinvents a Worse Version of CI/CD - Sumaiya Shrabony | Agent tooling | Argues solo builders end up rebuilding CI/CD-like infrastructure for their agents anyway. | 2,979 | 141.9 |
| 136 | Your AI Product Will Fail Unless You Can Explain It - Veronica Hylak, Hey AI | Explainability | Positions explanation as a core product requirement for successful AI systems. | 3,794 | 140.5 |
| 137 | On AI and Knowledge — Pablo Castro, Distinguished Engineer & CVP for AI Knowledge, Microsoft | AI knowledge systems | Pablo Castro reflects on how AI reshapes the creation and use of organizational knowledge. | 2,107 | 140.5 |
| 138 | Content Is Code - Matt Palmer, Conductor | Content as code | Frames content itself as a form of code in AI-driven workflows. | 1,920 | 137.1 |
| 139 | RLM: Recursive Language Models for Large Codebases - Shashi, Superagentic AI | Recursive language models | Introduces recursive language models as a technique for reasoning over very large codebases. | 2,723 | 136.2 |
| 140 | Frontier results, on device - RL Nabors, Arize | On-device AI | Covers frontier-quality AI results running closer to the device. | 4,247 | 128.7 |
| 141 | The AI bugpocalypse is here. Now what? - Jack Cable, Corridor | AI-generated code security | Argues that AI-written code is fueling a surge of security bugs and outlines how to respond. | 2,573 | 128.7 |
| 142 | Autonomous Agents for Scientific Tasks - Sina Shahandeh, Radicait | Scientific agents | Sina Shahandeh describes building autonomous agents that carry out scientific research tasks. | 1,779 | 127.1 |
| 143 | Continual Learning for AI Agents: From Failures to Durable Improvements - Soheil Feizi, RELAI | Continual learning | Looks at how agents can convert failures into durable improvements over time. | 3,386 | 125.4 |
| 144 | The Log Is The Agent - Ishaan Sehgal, Omnara | Agent observability | Argues the log is the primary interface for understanding and driving an agent. | 4,630 | 125.1 |
| 145 | Agents Building Agents - Alfonso Graziano, Nearform | Meta-agents | Explores agents that construct and orchestrate other agents. | 4,245 | 124.9 |
| 146 | Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs | Long-horizon evals | Lukas Petersson introduces Vending-Bench for evaluating agents on long-horizon tasks. | 995 | 124.4 |
| 147 | Agentic Security: Permissions, Provenance, and the Agent Supply Chain — Steve Yegge, Gas Town | Agent security | Steve Yegge covers permissions, provenance, and supply-chain risk for AI agents. | 1,489 | 124.1 |
| 148 | Stop Renting Your Cognitive Infrastructure - Thiyagarajan Maruthavanan, Kalmantic Labs | AI ownership | Advocates owning rather than renting the cognitive infrastructure behind AI systems. | 1,648 | 117.7 |
| 149 | Enterprise Agents Have a Structure Problem - Ishita Daga, Tesla | Enterprise agents | Ishita Daga argues enterprise agents fail on structure and organization rather than model quality. | 1,411 | 117.6 |
| 150 | Self Driving Products: Product Signals to Pull Requests — Joshua Snyder, PostHog | Product automation | Joshua Snyder shows how product signals can drive automated pull requests. | 6,106 | 117.4 |
| 151 | From Systems of Record to Systems of Context — Omri Bruchim & Tomer Ast, monday.com | Systems of context | Argues software is shifting from systems of record to systems of context built for AI. | 1,137 | 113.7 |
| 152 | Road to 5 Million Tokens: Breaking Barriers in Long Context Training — Max Ryabinin, Together AI | Long-context training | Max Ryabinin covers the engineering behind training models with multi-million-token context. | 5,918 | 109.6 |
| 153 | Video Has No Memory. Here’s How We Built One. — James Le, TwelveLabs | Video memory | James Le describes how TwelveLabs built persistent memory for video understanding. | 986 | 109.6 |
| 154 | Your Agents Need a Save Button - Hamza Tahir, ZenML | Agent state persistence | Argues agents need durable checkpointing so their work can be saved and resumed reliably. | 1,526 | 109.0 |
| 155 | Deterministic Infra for Non-Deterministic AI Agents - Nishant Gupta, Meta Superintelligence Labs | Agent infrastructure | Discusses infrastructure patterns for making stochastic agents more controllable. | 3,560 | 107.9 |
| 156 | Local Agentic Theory For Mobile Games — Shafik Quoraishee & Joanne Song, The New York Times | On-device game agents | Covers running local agentic reasoning for mobile games at The New York Times. | 938 | 104.2 |
| 157 | Your coding agent doesn’t always follow your rules — Talha Sheikh, Checkout.com | Agent reliability | Discusses why coding agents drift from instructions and what that means for reliability. | 2,403 | 100.1 |
| 158 | Learned Execution Graphs for Anomaly Detection & Drift in APIs — Ritvik Pandya, JP Morgan Chase | API anomaly detection | Ritvik Pandya applies learned execution graphs to detect anomalies and drift in APIs. | 899 | 99.9 |
| 159 | What Does Done Even Mean? Agents and Paperclip’s Liveness Model - Dotta, Paperclip | Agent task completion | Proposes a liveness model for defining when an agent’s task is actually complete. | 1,978 | 98.9 |
| 160 | The Agentic Web and the Bazaar Era of AI - Ramesh Raskar, MIT Media Lab | Agentic web | Frames the emerging agentic web as a decentralized “bazaar” era for AI-driven commerce and services. | 1,934 | 96.7 |
| 161 | Agents Need Receipts, Not More Tool Calls - Armanas Povilionis, Alithea Bio | Agent auditability | Argues agents should produce verifiable receipts of their actions rather than simply making more tool calls. | 1,331 | 95.1 |
| 162 | ReviewDebt: a practical framework for scoring every pull request — Sachin Gupta, Ebay | Code review process | Introduces a scoring framework for tracking review debt across pull requests. | 1,882 | 94.1 |
| 163 | The Missing Layer After Launch - Raphael Kalandadze, Wandero AI | Post-launch AI systems | Discusses the operational/product layer needed after an AI product ships. | 2,539 | 94.0 |
| 164 | The Prompt Is Still a Punch Card - Ted Johnson, JoinIn AI | Prompting interfaces | Compares prompts to old low-level interfaces and suggests better abstractions are needed. | 2,813 | 93.8 |
| 165 | Through the AI Fog: The Architectural Decision Agentic Security Depends On — Manoj Nair, Snyk | Agentic security architecture | Argues one architectural decision determines whether agentic security holds up. | 1,124 | 93.7 |
| 166 | I Run a Fleet of AI Agents Across Three Machines. Here’s What Broke. - Kyle Jaejun Lee, KRAFTON | Multi-agent operations | Lessons from running multiple agents across machines and dealing with the operational failures. | 2,219 | 92.5 |
| 167 | From Tokens to Cells: Foundation Models for Single-Cell Biology - Akram Baharlouei, Altos Labs | Single-cell biology models | Akram Baharlouei presents foundation models that treat single-cell biology data much like tokens in a language model. | 1,162 | 89.4 |
| 168 | Think You Can Build a Game with AI? Think Again! - Danielle An & David Hoe, Meta | AI game development | A reality check on the limits and challenges of building games with AI. | 2,006 | 83.6 |
| 169 | 500 people vibe-coded for 30 days. I was one of them. - Sanja Grbic, Automattic | Vibe coding | Lessons from a month-long large-scale vibe-coding experiment. | 2,062 | 82.5 |
| 170 | Don’t Let the LLM Drive - Ornella Bahidika & Joel Allou, Microsoft | Agent control flow | Makes the case for keeping deterministic control around the LLM rather than letting the model drive the whole workflow. | 985 | 82.1 |
| 171 | The Factory That Dreams: 39 AI Agents, No Framework - Rushabh Doshi, Machinecraft | Multi-agent systems | Describes running 39 AI agents together without relying on an agent framework. | 1,697 | 80.8 |
| 172 | Your LLM Stack Is a 2008 Database With Better Marketing — Lovina Dmello, NVIDIA | LLM infrastructure | Argues today’s LLM stacks resemble dated database systems dressed up with new marketing. | 964 | 80.3 |
| 173 | Semantic Blindness: 500,000 Sensors Confused an LLM - Raahul Singh & Vanč Levstik, Phaidra | Sensor data / LLM limits | Explores how a massive sensor dataset exposed blind spots in an LLM’s semantic understanding. | 1,603 | 80.2 |
| 174 | GTM Is You - Victoria Melnikova, Evil Martians | Go-to-market | A builder-oriented view of go-to-market where technical creators carry more of the motion. | 1,999 | 80.0 |
| 175 | Chat and citations won’t save your vertical AI - Atul Ramachandran, Filed Inc | Vertical AI product | Argues chat interfaces and citations alone aren’t enough to make vertical AI products succeed. | 1,672 | 79.6 |
| 176 | Agents Need Feature Flags - Sachin Gupta | Agent rollout control | Makes the case for feature-flagging agent behavior so rollouts can be controlled safely. | 1,092 | 78.0 |
| 177 | Stop Writing Tone Instructions. Layer Them. - Isadora Martin-Dye, Isadora & Co | Prompt tone design | Advocates layering tone instructions rather than writing them as one monolithic block. | 2,760 | 76.7 |
| 178 | You Might Not Need 50 Diffusion Steps — Ziv Ilan, Nvidia | Diffusion efficiency | Ziv Ilan shows how to cut diffusion sampling steps without sacrificing quality. | 3,414 | 74.2 |
| 179 | How we taught agents to use good retrieval - Hanna Lichtenberg, Mixedbread AI | Agent retrieval | Explains approaches for making agents use retrieval more effectively. | 1,719 | 68.8 |
| 180 | remobi.app: Don’t change your terminal workflow for mobile | Mobile dev tools | Demos remobi.app, which brings your existing terminal workflow to mobile without changing it. | 1,365 | 68.2 |
| 181 | A Song of Types and Agents - Roberto Stagi, Ratel | Type systems for agents | Looks at how strong typing can make agent-built software more reliable. | 1,344 | 67.2 |
| 182 | Structuring the Unstructured - Cedric Clyburn, Red Hat | Data structuring | Covers techniques for turning unstructured information into usable structured data. | 2,277 | 67.0 |
| 183 | Agents in Production: How OpenGov Built and Scaled OG Assist - Gabe De Mesa, OpenGov | Agents in production | A case study on building and scaling OpenGov’s OG Assist agent. | 2,231 | 62.0 |
| 184 | You Didn’t Ship a Bug. You Just Wrote It for a Human. - Ravi Madabhushi, Scalekit | Human-facing bugs | Argues that agent mistakes shipped to users are still bugs, just paid for by a human instead of a compiler. | 789 | 60.7 |
| 185 | Stop Evaluating Models Like It’s the 50s - Alejandro Vidal, Mindmakers | Model evaluation | Argues for modernizing how AI models are evaluated beyond outdated benchmarks. | 1,146 | 60.3 |
| 186 | Your Agent Failed in Prod. Good Luck Reproducing It. - Tisha Chawla & Susheem Koul, Microsoft | Production debugging | Covers the reproducibility challenge when production agents fail. | 1,990 | 60.3 |
| 187 | We Gave an Agent Production Code Access and Then Tried to Sleep at Night — Moritz Johner, Form3 | Agents in production | A candid account of granting an agent production code access and managing the risk. | 713 | 59.4 |
| 188 | It’s 10pm. Do You Know Where Your Agents Are? — Kim Maida, Keycard | Agent observability | Argues teams need real visibility and control over what their agents are doing. | 712 | 59.3 |
| 189 | Build the AI GTM Agent That Knows the Buyer - Dr. Sajjan Kanukolanu, Position2 (Position Squared) | Go-to-market agents | Shows how to build a go-to-market agent that understands the buyer well enough to drive real sales motions. | 707 | 58.9 |
| 190 | A Genius With Amnesia - Victor Savkin, Nx | Agent memory | Frames LLMs as brilliant but memoryless and explores how to fix that. | 2,102 | 58.4 |
| 191 | Your Voice Agent Doesn’t Need a Frontier Model - Joel Allou & Ornella Bahidika, Microsoft | Voice agents | Argues that voice agents can run effectively on smaller models rather than requiring a frontier model. | 699 | 58.2 |
| 192 | Develop at Idea Velocity - Jeffrey Lee-Chan, Snapchat | Development speed | Makes the case for building fast enough to keep pace with ideas rather than process. | 1,222 | 58.2 |
| 193 | Your agent is blindfolded — Johan Lajili, Poolside AI | Agent observability | Argues agents need better visibility into their environment to act effectively. | 1,356 | 56.5 |
| 194 | SWE-Marathon: Evaluating Coding Agents at Billion-Token Scale - Rishi Desai, Abundant AI | Coding-agent evals | Presents large-scale evaluation of coding agents over very long token horizons. | 1,360 | 54.4 |
| 195 | Claws Out: Securing and Building with OpenClaw - Nick Taylor, Pomerium | OpenClaw security | Covers security considerations for building on and with the OpenClaw ecosystem. | 1,138 | 54.2 |
| 196 | Agent Output Is Not UX: Rendering Layer Your LLM Pipeline Is Missing - Bala Ramdoss, Amazon Lens | Agent UX rendering | Argues raw agent output is not UX and that a rendering layer is the missing piece in most LLM pipelines. | 641 | 53.4 |
| 197 | Can Oncology Workflows Run Without Human Touch? - Anant Shankhdhar, Risa Labs | Healthcare agents | Anant Shankhdhar asks how far oncology workflows can be automated before a human must stay in the loop. | 641 | 53.4 |
| 198 | Your Agent Is Wasting Tokens and You Don’t Know It - Erik Hanchett, AWS | Token efficiency | Highlights hidden token waste in agent systems and ways to reduce it. | 1,765 | 51.9 |
| 199 | You Can’t Prompt the Room: The Last Skill AI Won’t Replace - Balázs Horváth, VisualLabs | Human skills | Argues for the continued importance of human facilitation and social skill. | 1,673 | 50.7 |
| 200 | The Prompt is the Platform - Dominik Tornow, Resonate HQ | Prompt platforms | Frames prompts as a platform layer for AI applications and workflows. | 1,665 | 50.5 |
| 201 | Using RL Agent to Detect and Remediate ETL Pipeline Failures - Anna Marie Benzon | RL for data pipelines | Shows reinforcement-learning agents applied to ETL failure detection and remediation. | 1,656 | 50.2 |
| 202 | Privacy-Preserving Intelligence — Steve Korshakov, Bee (acq. Amazon) | Privacy-preserving AI | Covers building intelligent systems that protect user privacy. | 581 | 48.4 |
| 203 | Production Evals For Agentic AI Systems - Nishant Gupta, Meta Superintelligence Labs | Agent evals | Covers building production-grade evaluations for agentic AI systems. | 1,782 | 48.2 |
| 204 | Browser Agents Don’t Need Better Models. They Need Better Eyes. - Kushan Raj, ARK | Browser agents | Argues browser agents are bottlenecked by perception, not model quality. | 1,627 | 47.9 |
| 205 | Running a Chess YouTube Channel entirely by AI — Stephan Steinfurt, TNG | AI media automation | Describes using AI to automate an entire chess YouTube channel workflow. | 1,110 | 46.2 |
| 206 | Security Track Intro — Randall Degges, Snyk | Security track intro | Opens the security track with an overview of AI security themes. | 552 | 46.0 |
| 207 | Agentic Development Security — Ezra Tanzer, Snyk | Agentic development security | Covers securing development workflows that rely on AI agents. | 547 | 45.6 |
| 208 | Voice Agents That Handle Interrupts - Chintan Agrawal and Daniel Wirjo, AWS | Voice agents | Covers how to build voice agents that gracefully handle user interruptions in real time. | 516 | 43.0 |
| 209 | Voice In, Visuals Out: The Agony and the Ecstasy - Allen Pike, Forestwalk Labs | Multimodal UX | Covers challenges and promise in voice-to-visual AI workflows. | 1,378 | 40.5 |
| 210 | When Agents Meet Physical Data: The Other Physics of Agent Harnesses - Dmitry Petrov, DataChain | Agent data infrastructure | Dmitry Petrov explores what happens when agent harnesses meet large, physical datasets. | 483 | 40.2 |
| 211 | Respect The Process - Andrew Dumit, Watershed Technology Inc. | Engineering process | Emphasizes process discipline in AI-enabled engineering workflows. | 962 | 38.5 |
| 212 | The 100-Tool Agent Is a Trap - Sohail Shaikh & Ankush Rastogi, Prosodica | Agent tool design | Argues loading an agent with too many tools backfires. | 1,278 | 37.6 |
| 213 | AI-Driven Multi-Document Correlation for Financial Compliance - Varsha Shah, Independent | Compliance AI | Applies AI to correlating multiple documents for financial compliance workflows. | 1,230 | 36.2 |
| 214 | OpenClaw in Your Hand: Building a Physical AI Terminal - Lech Kalinowski, Callstack | Physical AI devices | Shows a handheld or physical AI terminal concept built around OpenClaw. | 1,223 | 36.0 |
| 215 | GPU Cloud Deployment Without Leaving Your IDE — Audry Hsu, RunPod | GPU deployment tooling | Audry Hsu demos deploying to GPU cloud directly from the IDE. | 1,864 | 35.2 |
| 216 | From Transcription to Live Music: Gemini’s Audio Stack — Thor Schaeff, Google DeepMind | Audio AI | Thor Schaeff walks through Gemini’s audio stack from transcription to live music. | 1,856 | 35.0 |
| 217 | The Miranda Hypothesis: How Hamilton Poisoned Persona Evals - Jacob E. Thomas, Results Gen | Persona evals | Examines how a flawed persona corrupted evaluation results. | 1,244 | 33.6 |
| 218 | Sovereign Escape Velocity: Ownership w Open Models — Gus Martins, & Ian Ballantyne, Google DeepMind | Open models | Google DeepMind advocates for owning your AI stack through open models. | 1,729 | 33.2 |
| 219 | Agents Need Receipts, Not More Tool Calls - Armanas Povilionis, Alithea Bio | Agent auditability | Argues agents should produce verifiable receipts of their actions rather than simply making more tool calls. | 372 | 31.0 |
| 220 | Your LLM Deception Monitor Is Broken. The Fix Is in the Training Data - Sachin Kumar, LexisNexis | Safety monitoring | Argues that deception monitoring depends heavily on training-data quality. | 711 | 29.6 |
| 221 | User Signal Dies at the Retrieval Boundary - Sonam Pankaj, StarlightSearch | Retrieval signals | Discusses preserving user intent and signals across retrieval system boundaries. | 945 | 27.8 |
| 222 | When All Context Matters: Extended Cache Augmented Generation - Luis Romero-Sevilla, Orbis | Cache-augmented generation | Covers using extended cache techniques when broad context matters. | 910 | 26.8 |
| 223 | Research to Reality: Bringing Frontier ML Research to Production - Vaidas Razgaitis, Higharc | ML productionization | Discusses translating frontier ML research into production systems. | 865 | 25.4 |
| 224 | Bypassing the Multimodal Tax: Hybrid RAG, SQL RRF & UI Telemetry - Abed Matini, Ogilvy | Hybrid retrieval | Combines RAG, SQL, reciprocal-rank fusion, and UI telemetry for multimodal workflows. | 862 | 25.4 |
| 225 | Medic for Apache Spark - First Aid for Failing Jobs - Drasko Profirovic, Pinterest | Data pipeline debugging | Drasko Profirovic presents a tool that diagnoses and repairs failing Apache Spark jobs. | 299 | 24.9 |
| 226 | Using LLMs to Secure Source Code — Eugene Yan, Anthropic | Code security | Eugene Yan shows how LLMs can find and fix source-code vulnerabilities. | 334 | 22.3 |
Code to Replicate the Ranking
This script uses yt-dlp to fetch the channel’s current video metadata, filters to recent posted conference talks, and ranks by views per day. Install yt-dlp first if needed:
uv tool install yt-dlpThen run:
import datetime as dt
import json
import subprocess
CHANNEL_URL = "https://www.youtube.com/@aiDotEngineer/videos"
REFERENCE_DATE = dt.date(2026, 8, 1)
PLAYLIST_END = 250
EXCLUDE_TITLE_SUBSTRINGS = [
"Vibe Reel",
"Things to Know about AIE",
]
cmd = [
"yt-dlp",
"--playlist-end",
str(PLAYLIST_END),
"--skip-download",
"--ignore-errors",
"--print",
"%()j",
CHANNEL_URL,
]
proc = subprocess.run(
cmd,
check=False,
text=True,
capture_output=True,
timeout=420,
)
rows = []
for line in proc.stdout.splitlines():
if not line.startswith("{"):
continue
video = json.loads(line)
title = video.get("title") or ""
upload_date = video.get("upload_date")
view_count = video.get("view_count")
if not upload_date or view_count is None:
continue
if upload_date < "20260608":
continue
if any(skip in title for skip in EXCLUDE_TITLE_SUBSTRINGS):
continue
uploaded = dt.datetime.strptime(upload_date, "%Y%m%d").date()
days_since_upload = max(1, (REFERENCE_DATE - uploaded).days)
views_per_day = view_count / days_since_upload
rows.append(
{
"title": title,
"url": video.get("webpage_url")
or f"https://www.youtube.com/watch?v={video.get('id')}",
"upload_date": uploaded.isoformat(),
"views": view_count,
"days_since_upload": days_since_upload,
"views_per_day": views_per_day,
"duration_minutes": round((video.get("duration") or 0) / 60, 1),
}
)
rows.sort(key=lambda row: row["views_per_day"], reverse=True)
print("| Rank | Video | Uploaded | Views | Days | Views/day |")
print("|---:|---|---:|---:|---:|---:|")
for rank, row in enumerate(rows, start=1):
print(
"| {rank} | [{title}]({url}) | {upload_date} | "
"{views:,} | {days_since_upload} | {views_per_day:,.1f} |".format(
rank=rank,
**row,
)
)A couple of notes:
--ignore-errorslets the script skip scheduled premieres that do not have watchable metadata yet.PLAYLIST_END = 250was enough for this snapshot because the recent conference batch appeared in the first 250 channel videos.- The views are a snapshot. Re-running the script later will produce different view counts and a different views/day ranking.