← все видео

How Lyft Builds Evals That Actually Matter in Production | Interrupt 26

LangChain · 2026-06-15 · 17м 11с · 3 778 просмотров · YouTube ↗

Топики: ai-agent-orchestration

Аудио ещё не скачано.

📝 Summary

Summary ещё не сгенерён.

📜 Transcript

Transcript ещё не сделан.

⚙️ Pipeline jobs

Нет job'ов в очереди.

📄 Описание YouTube

Показать
Nick Ung, ML lead for safety and customer care at Lyft, breaks down how his team built an eval system that actually keeps pace with production AI agents at Interrupt, the agent conference by LangChain.

From offline simulation with mocked MCP outputs to LLM-as-judge rubrics tied to real task success criteria, Nick shares the mistakes they made early  and the framework they built to close the loop between evaluation, annotation, and continuous improvement across 270K AI interactions per month.

How Lyft Builds Evals That Actually Matter in Production | Interrupt 26
0:00 Introduction and Lyft AI Assist overview
1:47 AI Assist journey and agent examples
3:35 Why eval matters at scale
4:18 Offline eval as a quality gate
6:03 Building a lightweight simulator
8:17 LLM as judge and the problem with scalar metrics
9:52 Task-based rubrics and actionable failure modes
11:28 Trusting your LLM judge
12:17 The rude awakening — when offline evals lie
14:14 How LangSmith manages the eval workflow
15:46 What's next for AI Assist

Extra resources:
• Everything we shipped at Interrupt: https://www.langchain.com/blog/interrupt-2026-overview
• Meet LangSmith Engine: https://www.langchain.com/blog/introducing-langsmith-engine
• About LangChain: https://www.langchain.com/