← все видео

From Prototype to Production: Evaluating AI Products with Aman Kahn: ProductTank SF

Mind the Product · 2025-12-23 · 1ч 23м · 318 просмотров · YouTube ↗

Топики: product-discovery-loop

Аудио ещё не скачано.

📝 Summary

Summary ещё не сгенерён.

📜 Transcript

Transcript ещё не сделан.

⚙️ Pipeline jobs

Нет job'ов в очереди.

📄 Описание YouTube

Показать
In this Product Tank SF talk, AI product strategist Aman Kahn breaks down how product teams can move from prototyping AI agents to confidently shipping them. Drawing on experience at Cruise, Spotify and now Arize, Aman outlines a practical framework for observability and evaluation (eval) of LLM-powered products.

He demonstrates how PMs can use traces and spans to understand agent behaviour, build evaluation workflows using LLMs as judges, and align cross-functional teams around quality metrics. No code needed. 

Chapters
00:00 – Intro and background
03:00 – What kind of AI PM are you?
08:00 – Observability fundamentals
15:00 – Live demo: setting up tracing
24:00 – Observability to usefulness
26:00 – What is eval and why it matters
30:00 – Building your first eval workflow
45:00 – Running eval on agent data
55:00 – Improving your eval loop
01:05:00 – Role of teams in eval
01:15:00 – Dashboards and multi-agent orchestration
01:20:00 – Closing thoughts and resources

Key takeaways 
Observability is foundational to AI product development
System-level evals matter more than model benchmarks
PMs can lead eval strategy without deep technical skills
Start small with real data — iterate before you scale
Use LLMs to label data, but always verify with humans
Evaluation is continuous and works both offline and online
Discrete labels (e.g. correct/incorrect) work better than scoring
Strong evals come from cross-functional collaboration