How Lyft Builds Evals That Actually Matter in Production | Interrupt 26
LangChain · 2026-06-15 · 17м 11с · 3 778 просмотров · YouTube ↗
Топики: ai-agent-orchestration
Аудио ещё не скачано.
📝 Summary
Summary ещё не сгенерён.
📜 Transcript
Transcript ещё не сделан.
⚙️ Pipeline jobs
Нет job'ов в очереди.
📄 Описание YouTube
Показать
Nick Ung, ML lead for safety and customer care at Lyft, breaks down how his team built an eval system that actually keeps pace with production AI agents at Interrupt, the agent conference by LangChain. From offline simulation with mocked MCP outputs to LLM-as-judge rubrics tied to real task success criteria, Nick shares the mistakes they made early and the framework they built to close the loop between evaluation, annotation, and continuous improvement across 270K AI interactions per month. How Lyft Builds Evals That Actually Matter in Production | Interrupt 26 0:00 Introduction and Lyft AI Assist overview 1:47 AI Assist journey and agent examples 3:35 Why eval matters at scale 4:18 Offline eval as a quality gate 6:03 Building a lightweight simulator 8:17 LLM as judge and the problem with scalar metrics 9:52 Task-based rubrics and actionable failure modes 11:28 Trusting your LLM judge 12:17 The rude awakening — when offline evals lie 14:14 How LangSmith manages the eval workflow 15:46 What's next for AI Assist Extra resources: • Everything we shipped at Interrupt: https://www.langchain.com/blog/interrupt-2026-overview • Meet LangSmith Engine: https://www.langchain.com/blog/introducing-langsmith-engine • About LangChain: https://www.langchain.com/