From Prototype to Production: Evaluating AI Products with Aman Kahn: ProductTank SF
Mind the Product · 2025-12-23 · 1ч 23м · 318 просмотров · YouTube ↗
Топики: product-discovery-loop
Аудио ещё не скачано.
📝 Summary
Summary ещё не сгенерён.
📜 Transcript
Transcript ещё не сделан.
⚙️ Pipeline jobs
Нет job'ов в очереди.
📄 Описание YouTube
Показать
In this Product Tank SF talk, AI product strategist Aman Kahn breaks down how product teams can move from prototyping AI agents to confidently shipping them. Drawing on experience at Cruise, Spotify and now Arize, Aman outlines a practical framework for observability and evaluation (eval) of LLM-powered products. He demonstrates how PMs can use traces and spans to understand agent behaviour, build evaluation workflows using LLMs as judges, and align cross-functional teams around quality metrics. No code needed. Chapters 00:00 – Intro and background 03:00 – What kind of AI PM are you? 08:00 – Observability fundamentals 15:00 – Live demo: setting up tracing 24:00 – Observability to usefulness 26:00 – What is eval and why it matters 30:00 – Building your first eval workflow 45:00 – Running eval on agent data 55:00 – Improving your eval loop 01:05:00 – Role of teams in eval 01:15:00 – Dashboards and multi-agent orchestration 01:20:00 – Closing thoughts and resources Key takeaways Observability is foundational to AI product development System-level evals matter more than model benchmarks PMs can lead eval strategy without deep technical skills Start small with real data — iterate before you scale Use LLMs to label data, but always verify with humans Evaluation is continuous and works both offline and online Discrete labels (e.g. correct/incorrect) work better than scoring Strong evals come from cross-functional collaboration