Product Discovery Meets AI Evals with Teresa Torres
Hamel Husain · 2025-11-17 · 50м 41с · 18 151 просмотров · YouTube ↗
Топики: product-discovery-loop
🎧 Аудио
📝 Summary
model=deepseek-v4-flash · prompt=summary-v7 · 13 071→2 495 tokens · 2026-07-20 14:10:01
🎯 Главная суть
Continuous Discovery Habits и AI evals работают по одному научному методу: оба проверяют гипотезы через наблюдения и итерации. Самая частая ошибка AI-команд — начинать с eval'ов, не выбрав правильную проблему клиента. Evals — это просто ещё один discovery habit, способ проверить, что решение реально отвечает потребностям пользователя. Без глубокого понимания customer journey любые eval'ы будут garbage-in-garbage-out.
🔍 Главная ошибка AI-команд: решать не ту проблему
Самое страшное, что может случиться при создании AI-продукта — продукт с полем ввода «спроси меня что угодно». Если вы выбрали не ту задачу клиента, никакие хорошие eval'ы не помогут. Сейчас из-за хайпа FOMO и давления «добавить AI в роадмап» команды срезают углы. Строить AI-продукты объективно трудно, но лёгкость срезания углов делает эту ошибку массовой. Напрасно вкладывать энергию в eval'ы, если бизнес решает проблему, которая никому не нужна.
🧬 Evals как следующий discovery habit
Evals — это естественное продолжение assumption testing. В discovery работа идёт в цикле: индуктивное построение теории о том, как устроен клиент (opportunity space), и дедуктивная проверка этой теории через решения и эксперименты (assumption testing). Evals встраиваются в этот цикл: вы смотрите на traces и проверяете, соответствует ли ваша модель потребностям клиента. Это быстрая обратная связь: после каждого изменения (промпт, chunking, температура) вы видите, прошли traces или нет. По сути, это тот же способ оценивать гипотезы, просто в AI-контексте.
📊 Синтетические данные: начинать нужно с понимания клиента, а не с догадок
Если вы на старте (zero-to-one) и у вас нет production-данных, Hamel и Shreya в своём курсе предлагают генерировать синтетические traces через dimensions и tuples. Тереза поддерживает эту идею, но подчёркивает: нельзя просто гадать, какие dimensions взять — например, какие функции поддерживать, какие сегменты клиентов, какие сценарии. Всё это должно опираться на реальное знание клиентов. Прежде чем тратить LLM-запросы на синтез, нужно пойти и поговорить с пользователями: выяснить их потребности, понять, как они формулируют запросы. Если начать с угаданных dimensions, traces не будут отражать реальное поведение в production — и всё зря.
🧑💼 Кто должен быть benevolent dictator? Клиент
При error analysis часть ошибок очевидна (например, ответ не учёл питомца, хотя пользователь о нём написал; или предлагает субботу, хотя суббота недоступна). Но есть неочевидные случаи, требующие доменных знаний и понимания конкретной стратегии компании. Например, в ответе для инвестора не тот тон и не те объекты (starter homes). Кто решает, что это ошибка? Часто внутри компании назначается «benevolent dictator», который произвольно решает, что «хорошо». Тереза предлагает альтернативу: пусть этим диктатором будет клиент. Можно показывать пользователям traces в удобной форме и просить их оценить, правильно ли ответил AI. Это легко делать прямо в production через thumbs up/thumbs down с возможностью аннотации. Для случаев high ambiguity это даёт гораздо более валидную обратную связь, чем внутренние споры.
🎯 Метрики: не только бенчмарки, но custom metrics под ценность
При выборе метрик для eval'ов легко опереться на отраслевые бенчмарки: conciseness, helpfulness. Но они не всегда измеряют то, что реально важно для вашего клиента. Пример из рекрутинга: job boards меряют success как «подал ли заявку», хотя для соискателя и работодателя важнее «был ли нанят». Измерять «нанят» сложнее, но именно это — правильная метрика ценности. Аналогично в evals: для core-ценности вашего продукта нужна custom metric (например, match между тоном ответа и персоной клиента). Generic eval'ы хороши на старте, но быстро становятся недостаточными, как MAU → DAU/MAU → activation metric. Эволюция метрик — обязательная часть зрелого подхода.
🔬 Научный метод: инструментирование продукта и интервью о прошлом поведении
Тереза даёт практический совет тем, кто хочет улучшить discovery. Первый способ — инструментировать продукт. Даже просто для одного фичера: записать ожидаемый эффект (сколько клиентов будут использовать, сколько дойдут до «момента ценности»), а через 1–3 месяца сравнить с реальностью. Это улучшает интуицию и выявляет неверные предположения. Второй способ — регулярные customer interviews (хотя бы 20–30 минут в неделю). Ключевое правило: не спрашивать «что вы думаете?», «что вам нравится?», «что вы обычно делаете?». Вместо этого задавать вопрос о конкретном прошлом случае: «расскажите о последнем разе, когда вы использовали мой продукт / столкнулись с проблемой, которую он решает». Такая формулировка задействует систему 2 (медленное осознанное мышление) и даёт гораздо более надёжные данные, чем абстрактные мнения.
🔗 Параллели между trauma-анализом и customer interviews
В evals вы смотрите на конкретный trace — единичный пример того, что произошло. В customer interviews вы просите рассказать конкретную историю прошлого поведения. Оба подхода основаны на эмпиризме: наблюдение реального события, а не спекуляция. Это повышает точность гипотез и позволяет быстрее корректировать курс. Тереза отмечает: «смотрите на данные» и «говорите с клиентами» — это одно и то же по духу. Оба метода — часть научного мышления, которое должно быть в основе и product discovery, и AI evals.
📜 Transcript
en · 9 386 слов · 114 сегментов · clean
Показать текст транскрипта
The things that gives me nightmare most when working with teams building AI products is when the product says ask me anything. It doesn't matter how good your evals are if you picked the wrong customer problem to solve. I want you to think about this as we are just having an initial conversation about what this could look like. How do we get our teams to get excited about evals and what I say is don't talk about evals talk about the results of all the things you fixed my goal is hopefully this starts a conversation about what does discovery look like with AI products how does discovery inform evals you're gonna see I'm gonna frame evals as the next discovery habit because I do think it is a really important piece AI products because it's so trendy and there's so much FOMO and we're being all asked to add AI features to our roadmaps it's really easy to just cut corners but Building AI products is really hard. Don't spend all this time on evals if you're solving the wrong problem. Why are we letting somebody arbitrarily inside our building walls be the benevolent dictator of what good looks like when ultimately we're trying to serve a customer? Why don't we let the customer... Welcome everybody. So today we have a really special guest, Teresa Torres. Many of you already know her as a product discovery coach and the author of the book, Continuous Discovery Habits. Other product managers have described her as a national treasure, and I can see why. Her work has been really influential, helping teams move from feature factories to building products that create value. And it's one of the things that she really dives into in her teaching. So Teresa was a student in her first cohort. And in her last session, that she did with us, Teresa demonstrated her end-to-end evals workflow for an AI discovery coach that she's building, all from within a notebook. And she used AI assistants to do that, and she found important errors and used evals to systematically improve her app. So this course has focused heavily on the how of evals. So you've learned about the analyze, measure, and improve lifecycle. Today, Teresa is here to connect this technical process to the question of what we should be evaluating and she'll be explaining how her discovery habits can inform evals but more importantly she's going to show how the eval process can improve discovery enabling like faster assumption testing things like that so this is the session that bridges the gap from being an engineer who measures ai to being a product leader who can drive it towards value really excited about this please give a warm welcome to teresa Thank you, Hamill. That was a very nice intro. I'm curious what, I know Hamill said that this cohort is a lot of product managers, thanks to Lenny. I'm gonna start with a really high level overview of the discovery habits. I'm gonna try to connect it to what you're learning in class. And then what I'm gonna do is I'm gonna go through some of the ideas that you've learned in this class and show you where your discovery work can help make it even better. This is pretty, this is gonna be pretty informal. This is not like a polished talk that I've given a million times. You're just gonna see my slides are pretty rough. I want you to think about this as we are just having an initial conversation about what this could look like. I want you to push back on some of these ideas. So think about this. I always tell the teams that I work with, you're always iterating on crummy first drafts. Hopefully it's not crummy, but consider this kind of my rough first draft. And then my goal is hopefully this starts a conversation about how do we, what does discovery look like with AI products? How does discovery inform evals? You're going to see I'm going to frame evals as the next discovery habit because I do think it is a really important piece. And we'll just go from there. So I'm going to share my slides and we're going to dive right in. All right, let's do it. So if you're not familiar with my work, I've spent the last 14 years as a product discovery coach. Some of the ideas that we'll talk through, like I always share this slide at the beginning of my talks because product people tend to hear ideas and they think that'll never work in my environment. It's worked in a lot of countries, in a lot of organizations, big and small, regulated consumer, B2B. It really is battle tested and we're going to talk about why this works in so many broad contexts because I think this is going to be relevant to evals as well. I always start with what in the world is product discovery and like what is this jargon that we use? And so in the product world, discovery is just the work that we're doing when we decide what to build. And we contrast that with delivery, which is the engineering work that we're doing to write code, to ship production quality products, to maintain those products over time. And the reason why this split is important to talk about is a lot of companies overemphasize delivery and underemphasize discovery. And we're seeing that companies are finally starting to realize these are equally important domains. We could build the best version of the wrong product and we're not going to succeed. I did write a book, Continuous Discovery Habits. I hate it when people define terms because I feel like we get into these opinionated wars about what words mean. But in my coaching, I met a lot of teams that were like, Teresa, we already do continuous discovery. And really they did discovery activities on a periodic basis, but not from a continuous mindset. And so I did define continuous discovery in my book as weekly touch points with customers by the team building the product where they conduct small research activities in pursuit of a desired product outcome. And this is a mouthful. So to help folks understand it, I visualize it. And so this visual, I'm gonna spend three minutes walking through this just so you have the context and then we're gonna dive right in to how does this relate to evals. So the whole world that I created around continuous discovery habits was really meant for empowered product teams. So teams being asked to deliver an outcome. Just a little bit of history, especially for the younger folks in the group, and maybe even some of you work in these environments, because it's still very common. A lot of product teams were being asked to deliver just a feature list. Here's a roadmap. Your job is to go deliver this feature list. Go build, right? And there wasn't a lot of discovery in that process. A lot more companies, thanks to things like COVID, actually generative AI, these huge external trends that are disrupting roadmaps, teams are shifting to be outcome focused. Right? They're saying, companies are saying, look, we don't know what you should build. That's your job to figure out. But we know we need you to have this impact on the business. And the challenge with starting with an outcome is this is a wide open, unstructured problem. How do we know how we're going to reach that outcome? And so with the continuous discovery habits, we're looking at small research activities. The first is we're interviewing customers. The second is we're assumption testing. And this is the piece that I wanna talk through, because this is really just the scientific method. And so I wanna highlight this in contrast with evals. So when you start with an outcome, we need to reduce customer churn, we need to increase retention. How do we know what's gonna impact that number? We don't. And so we have to go interview customers. As we interview customers, we're not asking them, what should we build? We're interviewing them to learn about their world, their goals, their context, the environment in which they're trying to reach those goals, the tools that they're using. And as we do that, we've mapped the opportunity space. And I want you to think about this as we're developing a hypothesis. This is literally just the scientific method. We're developing a hypothesis of how our customers work. So we're developing a mental model of who they are, what their needs are, what they're doing. And then when we start to design solutions, I want you to think about your solutions as an experiment of does it match our understanding of our customer? If it does, it should work for them, right? And so as we develop a solution, and we assumption test to evaluate that solution, what we're really testing is our understanding of the opportunity space. And so if you're familiar with the discovery habits, you might not have heard me talk about it this way, but it really like the reason why these tactics work in so many environments is because it's really grounded in good scientific thinking and the scientific method and this idea of inductive and deductive thinking. So if you're not familiar with those terms, Inductive thinking is this idea of generating a theory of the world and then deductive reasoning is, okay, now let's look at a particular and test it. And there's this back and forth movement between induction and deduction. And so if you're familiar with opportunity solution trees, that's the visual in the middle. The opportunity space is our inductive theory of our customer's mental model and what they need and how they work. And our solutions are deductive tests of our understanding. And when we find problems with our solutions, we have to revise our understanding of opportunity space. We have to revise that inductive theory. Okay, so the reason why I'm framing it that way is because in this class, you learned that error analysis and evals and analyzing and measuring, this is just another flavor of the scientific method, right? There's like an underlying pattern here that's consistent across both methods. It's really how do we be good critical thinkers? How do we really make sure that what we're building matters to our customer? And with like traditional discovery, we're looking at a really close match between our understanding of our customers and what we're building for them. And the whole point of continuous discovery is how do we make sure we're building the right solution that matches that target opportunity, that customer need in a way that's going to create business value by driving that outcome. And so a lot of the discovery habits are how do we align these things? And that's why we have this visual that's a decision tree, right? We're like, as we do our discovery work, we're trying to keep these things really aligned. This is really similar to eval work. So I really think about evals when done well as really just a discovery, part of the discovery habits. So this is a slide I shared in my original talk. This is a grid of my first eval results. We're looking at one, two, three, four, my first five initial evals. The column down the left is traces, and this is just whether or not they passed or failed. And by having this visibility, I have a fast feedback loop, right? Every time I make a model change, a prompt change, a temperature change, a change to my chunking strategy, or really any other change, I can now get a fast feedback loop and make sure my changes are working for my customers. And that's no different than like the traditional assumption testing we've been doing to evaluate, does my solution address this customer need in a way that's going to drive my outcome for the business? And so with evals, we're asking, is my solution any good? is actually meeting the customer's need. So for me, evals are just another way of evaluating our assumptions, evaluating our solutions, evaluating if they're a good fit for our customer. The challenge is they only do this if they're designed well, right? So this in parentheses, when done well, is really critical. And so that's what I want to talk about. And this is where I'm going to connect the discovery habits. two evals. So now we're going to get into some of the things you learned in class. But before I do that, I want to talk about the biggest mistakes I see AI teams do. And I'm now seeing this a lot because I have a new podcast called Just Now Possible, where I'm interviewing product teams about the AI products that they're building. And I'm learning there are some teams that are really good at choosing the right customer problem to solve and are building pretty successful, amazing AI products. There's a lot of teams that are just sprinkling AI across their whole product platform. And it's just table stakes. Think about all these summarization and translation and sort of the low hanging fruit of AI. And then unfortunately I've had some false positives where I've spent 90 minutes interviewing teams and what came out of it, like I didn't publish these episodes because they didn't really choose the right customer problem. And they're building something that just doesn't make any sense. One thing I want to highlight is it doesn't matter how good your evals are if you picked the wrong customer problem to solve. I actually think AI products because it's so trendy and there's so much FOMO and we're being all asked to add AI features to our roadmaps, it's really easy to just cut corners. But building AI products is really hard. I hope this class is exposing how hard that is. Don't spend all this time on evals if you're solving the wrong problem. So this is just your reminder to just go back to the basics, make sure you really have an understanding of the opportunity space, you really understand what your customers need, and you're using that to inform what AI products to build. And one trend that has come out in my interviews, and I know Hamill will, this will probably resonate because I've heard him say similar things. When you're building a chat bot and people can type in anything, one of the most important decisions you need to make is what use cases are you going to cover? And how are you going to handle when someone puts something in that isn't a use case you can handle? And literally every team I've talked to, the very first step in their AI workflow is should my agent even respond to this input? And this is where the only way you're going to get this right is if you have an understanding of the opportunity space and what use cases your customers need you to address. And I saw Hamill came on video. No, I just came on video because I thought you may be saying something else. But no, I agree with you. The things that gives me nightmare most when working with teams building AI products is when the product says, ask me anything. Yeah. It gives me the feeling that, oh, maybe they don't know. what they really want to build here. And I need to refer them to you, I think, at that point. One of the things we talk a lot about on Just Now Possible, this is part of my list of questions, is, okay, you have a big vision for your AI product. Where did you start? Like, you can't on day one address every use case. So, like, how did you know where to start and how did you... iterate your way through a big, all these use cases. And in my head, the way that I visualize this is what was your understanding of this really big opportunity space? How did you choose where to start? What small use case did you start with? And how did you iterate your way through the opportunity space? Which by the way, good product teams have always been doing this. This is what good iterative continuous development looks like, right? So I just want to call this out because AI is so hot right now. We're forgetting some of these fundamentals. Okay, so now let's get into your course content. So I stole some images from the course reader. This is not a criticism of this course. I cannot tell you how much of a fan I am of this course and the work that Hamill and Shreya are doing. I just see lots of opportunities where understanding your customer makes this way better. So one of the first things you learn before you can do any error analysis is you have to generate some initial traces. And I noticed in this version of the course reader, they hammer home if you have real data, that is the best case scenario, real evidence of what your customers are doing. But if you don't have that, like you're a zero to one product, nothing's in production, you need somewhere to start. And there's this idea of generating your initial traces with synthetic data. And when I learned this, I was really excited because I was like, I don't have an initial set of traces and I need to generate some. And I loved this idea of identifying dimensions. coming up with values for each of those dimensions and then creating your tuples. Like amazing. I loved how structured it was. But in the back of my head was, yeah, but we can't just guess what these dimensions should be. What features should we support? Are we just guessing? And like, what client personas should we support? How do we know who our clients are? And that we have first-time buyers and experienced investors. And what scenarios should we support? And especially like... how clearly the user expresses intent. How do we know that they're going to type in something as lazy as show me homes? We know this by talking to our customers. If you're building an AI product and especially a zero to one product, don't just make up dimensions. You have to rely on your discovery foundation of understanding. These are the opportunities we're going to go after. By the way, those are your use cases. Your AI is going to address those opportunities. Those are your features. And then you have to be interviewing people that have those needs so that you understand the variation in customer segments. And you can even run assumption tests about, you can show someone a chat window or whatever the interface of your AI is going to be and ask them like, hey, you're looking for a new apartment. What would you type in here? And start to get a sense for what's that variation in how people express their intent and their interest. So if... Here's the deal, like your initial synthetic data is going to generate your initial traces. And this is the foundation for your error analysis and how you're going to evaluate quality. Don't guess here, like ground it in something you actually know about your customers. I really love this, by the way. One thing that we say in our course is you need to come to synthetic generation with a hypothesis. But I could leave it at that. Yeah. And I think it's a little bit. vague to some people. And I really like the way that you are couching it. And you know, that you're going to be explaining right now. So I think you're going to give a lot more meat to this. What is what do you mean hypothesis? Yeah, okay. Let's say you follow Hamill and Shreya's process perfectly. And you've built all these evals, but you started with a guess on your synthetic data. Like what's the risk of that? You're going to end up with an initial data set of traces that doesn't reflect what happens in production. Which means like, yeah, you evaluated quality, but you evaluated quality on the wrong thing. So this process that you're learning in this course is only as good as the data set you're using to drive your evals, like period. And so if you just do a little bit of discovery to start, and like maybe you do start with some guesses, go ahead and write down what you think the right feature set is and the right client persona is and how you think they're going to express. intent. Go ahead and be explicit about what that hypothesis is but now go out and talk to customers and run some assumption tests and learn about are those the use cases they care about? Is that how they express their intent? Are you seeing those segments in the people that you're talking to? And I would do that before you even start generating your initial traces. Like why pay for that LLM activity and don't start doing error analysis. Error analysis is a lot of work. First make sure you have a really rich understanding of these, like what this represents, what these dimensions represent is how well do you understand your customer and how they might want to use your product. So I would spend some upfront time making sure that I really understand this. And what's hard about this is on this call, like works on their product and has worked on their product for a while. Like you already know something about your customer. It's going to feel really easy. Like I can just make a good guess and it's probably pretty good. But I'm gonna remind you, you build a lot of things that your customers don't use. And that's not a criticism. We all build a lot of things our customers don't use. That's just a reminder that like your first guess is probably wrong in some way. So let's just do some iterations on that guess with a customer feedback loop before we invest all this time and energy in the evals. Because evals are hard, they take a lot of work. This is really, this is the foundation of your eval. So let's start and make sure that foundation is really strong. All right. What do you suggest to people who are building chat bots? And this is really is the case that there's so many chat bots out there where the input box has a ghost text in it that says, ask me anything. Yeah. Do you, what do you suggest? Like people take a hard look at their scope and say, yeah, what do you do? I think the first thing. This has come up a lot in my podcast where I'm again, so many teams went through this learning curve of default to a chat box because we're used to using chat GPT. And like that might be the right interface, but a lot of people, they see a chat box and they don't know what to type in. So like you have to give people, you have to remove the blank page problem, right? If you just say, ask me anything, you're making me the user do all the work. If you understand your customer. and you have a really good hypothesis of this feature dimension, you can have a better default text there. You can say, ask me about any of the following things, right? And now you're steering them towards an area where you might be successful. And you see even ChatGPT over time has started doing this. When you go to the ChatGPT interface right below the open chat box where you literally can type in anything, there's suggestions on what ChatGPT can help you with. But these suggestions have to be grounded in what your customer needs. Because if you suggest things they don't care about, you're not really solving the problem. And then I'll tell you every team I've interviewed, literally their first step, like they do this scrubbing step before they even send it to the agent, of is the intent behind what the user entered, something we can adequately respond to. And only if it is, do they pass it on to the agent. If it's not, then they respond with something generic. Here's let's connect you with a human or go visit our knowledge base or whatever, because they don't want the agent hallucinating, making something up. So I think this is a theme of if you want to build a good general agent, you have to be really clear about what you can adequately address. And then you have to steer people towards those use cases. I like that. Okay, let's look at the next step. So let's assume we've gotten to the point where we're confident in our dimensions. We've gotten some feedback from our customers. We're generating our initial traces based on our tuples. It's now time to do error analysis. Now I also took this from the course reader and it is true. Some errors are obvious. You don't need to know very much about your customer to do error analysis. In this first situation, like the user put in, they have a pet and the response didn't include the pet. I don't need to have any domain knowledge. I don't need to have any customer knowledge. That's clearly wrong to me. And like, we're all gonna agree that's wrong, right? Same with this second one. There's no availability on Saturday and it's suggesting Saturday, obviously wrong. But you're gonna have error analysis cases where if you don't know your customer and what they need, it's not gonna look wrong and you're not gonna catch the error. And this is why like a lot of this work is being moved from engineers to product managers because engineers don't always have the domain knowledge to identify the error in the first place. And I think this third example is a good example of this. Send a property list to my investor client in San Mateo. And the error analysis is there's a persona mismatch and there's a tone and property mismatch. How do we know that? How do we know what investors want? How do we know the right tone to use with them? How do we know they don't want starter homes? This is the domain knowledge that's required to do good error analysis. And what's hard is we can go find a domain expert, like maybe a realtor. But is that realtor, like they have good intuitive expert knowledge, but do they really know who your customers are and your customer segments? And no two real estate firms are going after the exact same customer segments. This is really unique to your company and your company strategy. of which customers we're going after. And so it's not just domain knowledge about real estate, it's domain knowledge plus company strategy knowledge plus your specific customer's knowledge. And so this is where like discovery has to play a role. Like we have to know who our customer is, not just people who buy houses, the people who buy houses that our company is trying to reach, right? We talk about domain knowledge, like domain expertise, how do we know what good looks like? And I love like Hamill's term of the benevolent dictator. But because like I've been in those meetings where people quibble over wording and we can take forever. And it's really nice to have a benevolent dictator just say, no, this is the decision. This is not good. We're going to mark this in error. The challenge is who should be that benevolent dictator? And this is where I really want to push everybody on the call to think about. maybe it should be your customers. What does that look like? Why are we letting somebody arbitrarily inside our building walls be the benevolent dictator of what good looks like when ultimately we're trying to serve a customer? Now, this is going to sound a little ideal. I realize when we're going from zero to one, we may not have a ton of time to have a customer look at every trace and there probably is a need for internal benevolent dictator. But here's the thing, like in discovery, we have super fast feedback loops now. So what if we showed some traces in a really customer friendly form, not all the tool calls and whatnot, but imagine you were trying to do this and you got this response, what would you think? And we let a customer annotate it. We can do that with our assumption testing tools. We can use unmoderated testing platforms to get like hundreds of traces annotated by customers in a single day. And so for our really tricky ambiguous cases, Why don't we let the customer be the benevolent dictator? What does that look like? And I'll tell you, this is rough for me. Like I haven't experimented with a lot of these tactics yet. My nature is I want to push as much. I want to get my feedback loop as close to the customer as possible. And this seems like an area where doing this with customers makes a lot of sense to me. I really love this idea. And I've been talking with friends about even making some tools that... When you onboard your customer to your product, you tell them, like, help us train the AI to you and have them annotate some things deliberately. Yeah. With an incentive, we're going to make it better for you. And we're going to align the AI to you and, like, show them a progress bar. and then you can keep surfacing things to them like as they discover errors. So I think this is an excellent idea. I love that you're bringing this up. Yeah, and when we have a production product, this is actually really easy to do, right? Almost all of us are already integrating these like thumbs up, thumbs down feedback. I think the key is we want them to also annotate it. So like in my interview coach, when you get feedback, you can give a thumbs up and thumbs down, but you can also tell me why. And I think that's really key. You're actually getting a customer annotation on your trace. When we're defining the right metrics, it's the same exact thing. I'm not going to belabor this point because it's very similar to here. Like how do we know what good looks like? When you get to the point where you're defining the right metrics and you're deciding what to measure, it's really easy to rely on these like industry benchmarks. One of the hard things with metrics is are you measuring what you think you're measuring? Like in research there's this idea of validity. Does our instrument actually measure the concept we're trying to measure? And this it's a hard concept to wrap your head around because like it's so easy to make jumps as a human thinker like it was for reasoning. We think about, we don't always think about the ideal thing to measure. We're really influenced by what we think we can measure. And what we think we can measure isn't always the right thing to measure. And so there's this, I'll give an example of this. Like I used to work at recruiting companies like job boards and almost every job board, their measure of success is did you apply for a job? That's a really crappy measure of success because it leads to a lot of bad applications to jobs people are irrelevant for. The job seeker doesn't care if they applied, they care if they got the job. The employer doesn't care if you applied, they care if they want to hire you. It's like a better measure of success is, did we help somebody find a job? Did they get hired? But that measure is hard to measure. That metric is hard to measure. And so people are afraid of it, but it is the right metric of value. And so I think the evals have the same exact problem. If we just limit our measurements to like, these benchmark metrics or these traditional ML metrics, there are going to be cases where those are the right metrics. But the better you understand your customer, the better you're going to be able to come up with what is the right metric. And there are going to be times when you're designing an eval that's like for the core part of your product, the core part of the value that it delivers, that you have to come up with a custom metric. Yeah, maybe there's benchmarks on a concise tone or a helpful tone, but is that for your customer? Like what's the right persona tone match for your specific customer? And that's probably a custom metric. And so that really requires that you understand your customers and what they need. And this is why like all these off the shelf eval tools are tough. Like these out of the box evals might get you started, but you're gonna outgrown pretty quickly. And we see this with product metrics, right? Like when we're measuring engagement. We start with like monthly active users and then we get a little more sophisticated and we do DAO over MAL. So daily active users divided by monthly active users. Then we get a little more sophisticated and we say, look, not all engagement is the same. What are the behaviors in our product that are actually really important? And we start to get to an activation metric and that's very unique to our product. I think evals are the same. Like you might start with a generic eval, but you better iterate your way to something that's really tightly tied to the value your product delivers. Okay, so this is my point. Evals are the next discovery habit. They really fall into this category of assumption testing. How do I evaluate if my solution is actually meeting my customer's needs? But if your evals aren't grounded in your understanding of customer needs, it's garbage in, garbage out. Like we can measure a lot of the wrong stuff and it looks like our product is doing great, but it really only is as strong as the inputs. And so this is where I just wanna encourage you and especially as you launch. Build in customer feedback loops. Make sure they can rate your product thumbs up, thumbs down. Cloud Code is doing this really right now. I constantly get asked, how is Cloud doing? One, two, three, four, five. It's like fine, good. I don't remember what else. Cloud Code never asked me to annotate it though. So I feel like there's a missed opportunity there. So definitely be thinking about even once you go into production, where is that feedback loop coming from? All right. And then if you are interested in diving deeper on this, I will share. I didn't create a slide for this, but I do blog about discovery and now more AI products at producttalk.org. And then I do have the new podcast just now possible where I'm interviewing product teams and we go deep. Like we get into the architecture maps, how people are orchestrating, what's an agent, how they're using rag steps. Like we really get into the nitty detail. nitty gritty detail, including evals. We go deep on evals. Fun, so definitely check that out. And then we've got some time for questions. I just want to also say, you should check out Teresa's newsletter. I read every single post she does very carefully. And it's excellent. If this talk has, you could just tell from this talk, this talk should convince you in itself. Yeah, I couldn't be a bigger fan of Teresa. I was just saying in the chat, this is my favorite. This is now my favorite evals talk. And probably my second favorite evals talk is Teresa's last talk. If you really want to learn more about evals as well, please check out everything she's doing. I'll share on that front. I just yesterday, I kicked off this new series that I'm really excited about cloud code. And it doesn't have to be cloud code specific. Literally everything I'm writing applies to a command line interface tool. So if you prefer Gemini's or Codex, it's all the same things are going to apply. It's specifically for product managers and I'll share the story behind this. Over the last four months, I've gone deep on command line interface tools and not just for coding. I am using them for coding, but I now do my task management out of Cloud Code. I built my own little roll-your-own task management tool. I have this post coming up about how to set up context files so you can write really lazy prompts and Cloud still knows everything it needs to know to do a good job. I'm going to get into safety and how to run it safely on your computer and how to evaluate what code to let it run and what packages are safe and how to evaluate if it's okay to install packages. I really think that engineers are pushing the envelope on how we collaborate with AI and it's time for this to spill over into the non-technical world. I'm kicking off this series. It's probably going to be four to six posts. The very first one went live yesterday and I'll tell you what my goal with the post was. You can be a total beginner, never even heard of a terminal. I'm going to walk you through why you should care about Cloud Code, how to install it, and by the end of the blog post, you're going to have a competitive analysis of all of your competitors where you get a pricing comparison table and a feature comparison table. It's going to happen in two minutes. Cloud Code is going to do all of the work and it's going to be set up in a way that if you want to add a competitor tomorrow, you're just literally adding it to a text file and then rerunning a slash command. If some of that makes no sense to you, the article explains all of it. My goal with that blog post was to go from beginner to magic in one blog post. Like it was very ambitious, but I think I pulled it off and I think you should check it out. And then there's going to be much more coming in the next few weeks. I'm going to talk about, I now get a research report every morning of every academic article that's related to anything that I do. When I save a PDF, it automatically gets, of one of those articles, it automatically gets summarized and the summary gets added to my to-do list the next day. Like I've just built out, I'm framing this as I've built out my own personal operating system. And I now want to teach other people how to do it. So check that out. And what's the best way for people to find this? Is it to go to producttalk.org? Producttalk.org. Yeah. And then sign up to your newsletter and they'll get all this information? Or how do you find? Yeah. The articles are just on producttalk.org. Like you don't have to sign up to get them. For the Cloud Code series, there is a paywall at some point in the article. Here's my goal. My goal is to give the concept away for free. And then if you become a supporting member, what you get is literally step-by-step instructions on how to put it into practice. So if you're one of those learners where like just understanding the concept of is enough and you're going to run with it, that is completely free. If you want like a little more hand-holding, it's on par with like most Substack subscriptions. It's not crazy expensive. That's great. I'm going to sign up right after this if I haven't already. I think I may already be. uh sub stack subscriber but we'll see i'm gonna make sure okay we can take questions now ask oh by the way have you started using skills the new kind of thing from yeah so you'll see in the blog post i released yesterday i announced as i was writing the skills came out and it might change things a little bit skill i'm excited about skills i'm also a little worried about skills what skills do is they allow you to package sort of context files plus slash commands plus scripts, so deterministic code together, and then you can share it with other people. One of my concerns with Cloud Code is that if you don't have a technical background, there are some security things you need to be aware of. You're giving Cloud access to your entire computer, and if you're letting it run code and especially install packages, there is this package hacking thing that happens where People are like overloading package names and you end up not getting the package you think you will get. And you could end up installing malicious code on your computer. And so this is particularly could be a problem with skills and Anthropic even acknowledges this. So like what I would recommend if you don't have a coding background, don't run code on your computer that you don't understand. That's just rule number one. With skills in particular, you may not even know there's code in there. You need to know a skill could have code in there and you're not really deciding when it gets executed. As an engineer, someone with an engineering background, I like the idea of skills because I want to package things and create things. I plan to release skills, but I do think this is the Wild West. This is a frontier. You have to understand the dangers. Even Anthropic acknowledges this. Anthropic says, don't install skills from people you don't trust. There's a little bit of danger here. People are putting out some pretty cool skills, but there are going to be malicious actors. And when it's running on your local machine, you got to be careful about that. And I'm going to write a blog post about security too. And I am excited about skills. Just be cautious is my right now answer. Okay. So one question from Eileen is, for teams you've seen doing the nightmare of trying something AI without doing enough product discovery, what have you found to like... How do you get them back on track? What are like some first steps you can take the team? Yeah, how do you, what do you do? Yeah, I side because this is a hard question. This is not specific to AI products. Like the sad reality is most product teams don't do enough discovery or they don't do discovery in reliable ways, right? They're not getting reliable feedback. I think what's hard, like we already waste a ton. This is going to sound really depressing, but I'm going to try to bring it back to something optimistic. What's hard is like today and today's reality, product teams already waste a ton of engineering time on the wrong product to build. And actually AI making it faster and easier to build is gonna make this worse, right? Like I'm reading about companies where like the marketing team is pushing code into production. Okay, that's awesome. And it's also a nightmare, right? Like we don't want... anybody and everybody to be able to build and release their idea. Like this is what leads to terrible products. And for me, like the easier delivery gets, I think the more important discovery gets. We should not build every idea. Our AI product should not cover every use case. Like, and I think a really emerging skill is like understanding not just what use cases we should address, but like where do we need AI and where do we need to determine a stick code? And we always need some of both. And what does that interplay look like? And I think really understanding like the end goal is creating customer value. Like we need to create something for our customer. And we're not, to be honest, we're not great at this. Now we weren't great at this 20 years ago and we're getting better, but we're not great at this. And I think with AI products, because of FOMO and all the hype and are we in a bubble and companies over investing in really shallow ways where they're like release something next quarter even though we've never done this before. This is gonna get worse before it gets better. But I also, here's where it's gonna get positive. I firmly believe that organizations change and get better at discovery when individuals change and individuals get better at discovery. So I'm a big fan of high agency product people because we can influence change. And so I think the real answer to your question is not how do I get a company back on track? It's how do I get myself back on track? How do I get my team back on track? And then we influence by showing what good looks like. And I think that's, for me, that's really empowering because I can change the way I work. So especially for product people- How do you show what good looks like with discovery? How do you say, hey, this is good? All right, listen to me. How does that manifest? I think we increase our hit rate. So there's a few things that have to be in place. If you're not instrumenting your product and you have no idea if it's working or not, start instrumenting your product. That's step one, right? And that doesn't mean stop everything and instrument every part of your product. It could be as simple as for the feature you're building right now, instrument it. What do you expect to happen when it releases? How do you measure that? Are people using it? Are they finding it? Are they using it the way that you expect? Are they using it all the way through to the value creation moment? The other thing I'd recommend if you're going to start, if you're new to instrumenting your product, before you release that feature, as a team, write down what impact you expect it to have. Write it down. How many customers are going to use it? How many are going to get all the way through to the value creation moment? This is going to feel really uncomfortable. You're going to be like, how the hell do I know? You're making assumptions. Your assumptions are really idealistic. Everybody's gonna use it. They're gonna use it exactly the way that we wanted them to use it. They're all gonna have the value creation moment. And as a result, our retention is gonna go up by 10% and we're all gonna get rewarded and it's gonna be amazing. Write it down. Then a month, two months, three months, whatever the right cadence is, take a measurement compared to what you wrote down. This is where it's hard because there's gonna be a giant gap between what you thought and what actually happened. It doesn't mean you did something wrong. Every single product team, 100% of us, there's a gap between what we think happened and what actually happened. By building this habit, what you're doing is you're training your brain. You're improving your intuition. So the next time you build a feature, if you're way too optimistic, what you expect to happen is going to come down a little bit. It's also a feedback loop on your discovery. When our expectations fall short, it means there's an assumption. that our idea depended upon that wasn't quite true. So it'll help you uncover that assumption, which will make you better at your next round of assumption testing. Which by the way, just the scientific method, right? This is what scientists do. And it's just, but what's hard is that we're not all trained scientists and we have to build this muscle of take a measurement, adjust, take a measurement, adjust. Your best feedback loop on how good your discovery is, is that gap closing. Are you getting closer to understanding the impact of what you're building? And then as you get closer, you build the right stuff more often. You're never always going to have 100% batting average. You're never always going to build the right thing. There's always going to be some gap, but we should be able to close that gap. How do we get our teams to get excited about evals? And what I say is don't talk about evals. Talk about the results of all the things you fixed and all the bugs you squashed and whatnot. and like show the evidence. So that's, it's very interesting how it mirrors each other. Another genre of question we're getting is, okay, so this class has a mixture of people, product managers, engineers. Ever since I got to know you, Teresa, I've actually become a lot more interested myself in product management. I think it's you got me excited about product management in a way that no one else has. Awesome, because you'd be excited about data science. Yes. Yeah. I think it's because like I see the connection and I see the power of threading the needle all the way through. And a lot of questions are okay. For example, is, or in evals, like a really good entry point is error analysis. So we pick that on purpose as an entry point as people can like get into evals for all the engineers and other people. Do you have any intuition on good entry points or where to start? if you want to branch into this, like doing better discovery and all the things you talk about, where should you go? What should you do? Yeah. So I think error analysis is a good entry because it exposes what's not working. And so I think the same is true in product management. You got to expose what's not working. And I think there's two ways to do this. We talked about one, you can instrument your product. Here's the deal. I know that 50% of you do not have an instrumented product. The way I know is because we run a survey every two years. And you tell me only 50% of you have access to behavioral analytics in your product. That is insane to me. That number seems high to me, actually. Yeah. Well, any instrumentation. Google Analytics, any instrumentation. So 50% of us are flying completely blind. Okay, if you can get any instrumentation in place, do that. But I also know this is hard. Like a lot of us work in organizations that are not quantitative. They don't care about measurement. What they care about is stakeholders telling you what to build and they want their pet feature built. So I'm going to give you another entry that's probably easier, which is talk to your customers. So interview customers. I literally, I think you can do this in as little as one customer conversation every week. And a customer conversation can be 20 to 30 minutes. The key here is you can't just go talk to a customer like I'm going to call up Hamill and have a human-to-human conversation. It is true that customer interviews are just human-to-human conversations. We don't need to be afraid of them. But if you're going to use them as a feedback loop, you have to learn a little bit about how to get effective feedback from another human. This is where we have to have a basic understanding of cognitive biases and what questions to ask. I'm going to tell you this one secret that you need to remember. when you're interviewing a customer, don't ask them what they think, don't ask them what they like, don't ask them what they do. These are all really unreliable questions. They're all very speculative. The human will give you a response. The response will not match what they do in reality. What you want to do instead is you want to ask me, tell me about a specific time when you did a thing. And that could be, tell me about the last time you used my product. It could be tell me about the last time you had a problem my product was designed to solve. It could be tell me about the last time you used my product on the go if you're on the mobile team, right? But you want to keep your interview grounded in a specific story about past behavior. The key is you want to invoke their memory. When we invoke their memory, if you're familiar with Daniel Kahneman and Amos Tversky's work, we're now engaging system two. System two is our slow deliberate brain. far more reliable responses. All those other questions, system one responses, rife with cognitive biases. So if you can't instrument your product, a different feedback loop, and actually you should do both of these ideally, just talk to a customer on a regular basis, but ask for past stories. Wow, that's really amazing. There's so many parallels with evals because in evals, you want to look at very specific data and you don't want these like generic. kind of analysis. And that's really great that, okay, ask for specificity and specific experience. Experiences is like a trace in a way. This is what happened. Your look at the data mantra is exactly equivalent to my talk to your customers mantra. Like it, it all starts with what's really happening. Right? Like You can't, don't just make a guess like what is actually happening in reality. So like looking at a trace, this is telling you what's actually happening in reality. Talking to a customer and collecting a specific story about past behavior, this is just empiricism. Like we're observing, still it's all grounded in the scientific method, we're observing what happened and we're using that as a finding to then inform our next hypothesis. There's so many analogies here. This is why I loved this class. I was like, oh, somebody else that cares about rigorous scientific thinking in the product world. Yay. That's really great. So yeah, thank you so much for coming, Teresa. This is very, this is always a pleasure to have you. We should team up at some point somehow because really love your thinking and everybody loves your thinking. And yeah, I really recommend everyone to check you out and learn more about all the things that you do. Yeah, thank you everybody. And I really am super excited about my new podcast. So if you're familiar with my discovery work, but you don't know that I have a podcast, we geek out on all things AI. And I am, it is the one thing that brings me so much joy to like nerd out with builders. So I feel like this whole AI wave has helped a lot of people reconnect with being makers and builders. That's great. Thanks so much. And thank you everyone for coming. Thanks everybody.
⚙️ Pipeline jobs
| Stage | Status | Att. | Updated | Error |
|---|---|---|---|---|
| download | done | 2/3 | 2026-07-20 14:08:57 | |
| transcribe | done | 1/3 | 2026-07-20 14:09:26 | |
| summarize | done | 1/3 | 2026-07-20 14:10:01 | |
| embed | done | 1/3 | 2026-07-20 14:10:02 |
📄 Описание YouTube
Показать
Join the AI Evals September 2026 cohort: https://maven.com/parlance-labs/evals?promoCode=yt-2026 Teresa Torres, product discovery coach and "national treasure" to PMs worldwide, sits down with Hamel to connect the technical world of AI evaluation to the strategic question: are you even solving the right problem? If you're building AI products, this conversation will fundamentally change how you think about evals. You'll discover why "ask me anything" AI products keep product managers up at night, how assumption testing from discovery maps directly to evals workflows, and the one interview technique that gets you reliable customer feedback (hint: stop asking what they think). This is the missing bridge between being an engineer who measures AI and being a product leader who drives it toward real value. Timestamps: 00:00 Introduction 01:33 Welcome Teresa Torres - From Discovery Coach to Evals Student 11:43 What is Product Discovery? Discovery vs Delivery 18:00 The Continuous Discovery Habits Framework 24:15 Opportunity Solution Trees and Assumption Testing 31:20 How Discovery Informs AI Evals 38:45 The Scientific Method: Discovery + Evals 42:10 Instrumenting Your Product vs Flying Blind 48:43 The One Secret to Customer Interviews 49:41 Empiricism: Traces, Stories, and the Scientific Method 49:51 Final Thoughts and Teresa's AI Podcast Follow Teresa Torres: Website ► https://www.producttalk.org/ LinkedIn ► https://www.linkedin.com/in/teresatorres/ Follow Hamel: YouTube ► https://www.youtube.com/@hamelhusain7140 LinkedIn ► https://www.linkedin.com/in/hamelhusain/ X ► https://twitter.com/HamelHusain Website ► https://hamel.dev/ Resources Mentioned: - Continuous Discovery Habits (book by Teresa Torres) - AI Evals Course: https://maven.com/parlance-labs/evals?promoCode=tf-yt-c4 - Teresa's Podcast on AI and product discovery