How Clay runs 350 million GTM agents a month | Interrupt 26
LangChain · 2026-06-24 · 11м 54с · 4 552 просмотров · YouTube ↗
Топики: ai-agent-orchestration
🎧 Аудио
📝 Summary
model=deepseek-v4-flash · prompt=summary-v7 · 4 314→1 261 tokens · 2026-07-20 13:40:46
🎯 Главная суть
Clay — платформа для go-to-market (GTM) автоматизации, которая запускает 350 миллионов GTM-агентов в месяц. Основная идея: в современных условиях холодные email-кампании теряют эффективность, поэтому побеждает тот, кто быстрее всего итерирует стратегии. Чтобы получить GTM-альфу (превосходство над конкурентами), нужно масштабировать агентов на весь адресный рынок, а это инженерный вызов, включающий инфраструктуру, rate limits, стоимость и качество.
🏗️ Инфраструктура: от Lambda к durable execution
Большинство агентов проводят время в ожидании — браузеров, API или инференса. Изначально Clay запускал Claygent на AWS Lambda, но это оказалось prohibitively дорого из-за оплаты за wall time. Переход на ECS решил проблему стоимости, но породил проблемы с надёжностью — требовалась защита от случайных отказов хостов. Правильная архитектура — durable workflow execution: очереди, чекпоинтинг агента на периодических шагах. Для этого подходят инструменты вроде LangGraph или LangSmith Deployments.
⚡ Rate limits: адаптивное троттлирование и справедливость
Нагрузка на Clay сильно неравномерна (spiky). Чтобы максимально использовать выделенную мощность инференса, построена система с back pressure, которая напоминает алгоритм перегрузки TCP/IP: отправляется максимум трафика, а при достижении лимитов трафик прогрессивно снижается. Внутренние эксперименты показали, что такой подход даёт в 4–10 раз больше пропускной способности по сравнению с наивной системой. Дополнительно внедрён механизм fairness: один клиент, запускающий миллионы агентов, не должен вытеснять нового пользователя с десятью агентами.
💰 Стоимость: кэширование и ограничение шагов
Масштаб Clay (триллионы токенов в неделю) делает стоимость критической. Собственный агентский харнесс позволил внедрить стратегии кэширования, которые для таких провайдеров, как Anthropic, дают до 70% экономии. Второй приём — ограничение повторных попыток и количества вызовов инструментов: часто принудительный возврат агента после определённого числа шагов даёт лучшие результаты, чем полный прогон (требует проверки через eval). Третий — измерение стоимости в привязке к качеству и результатам.
✅ Качество: контекст, evals и продуктовая разработка
Качество Claygent строится на трёх столпах: 1) отличный контекст — собственные GTM-датасеты (40 млн компаний, 900 млн контактов) плюс веб-данные; 2) офлайн- и онлайн-evals для настройки харнесса под GTM-задачи; 3) продуктовая составляющая — в Clay встроен Agent Builder, где пользователи тестируют и итерируют агентов до запуска на рынок. Это повышает уверенность в результатах и, как следствие, качество на масштабе.
🔁 Наблюдаемость как цикл обратной связи
Чтобы агенты становились лучше со временем, необходима наблюдаемость (observability). Инструменты вроде LangSmith помогают понять, что именно оптимизируется и почему. Сочетание офлайн- и онлайн-evals позволяет непрерывно улучшать агентов, замыкая цикл: запуск → сбор данных → анализ → итерация.
🧠 Будущее: Audiences и agent memory
Clay анонсировал продукт Audiences — агрегатор всех GTM-данных в одном месте (Snowflake, Salesforce, Gong и др.) с наложением сторонних сигналов (новости, раунды финансирования). Это становится фундаментом для agent memory: агенты смогут рекомендовать стратегии на основе предыдущих попыток и контекста, образуя маховик улучшения. Дополнительные вызовы — виртуальные файловые системы для рассуждения над контекстом и песочницы.
📜 Transcript
en · 1 964 слов · 24 сегментов · clean
Показать текст транскрипта
Hi, everyone. I'm Jeff Bard. I'm the head of AI at Clay. And today I'm going to talk about scaling go-to-market agents. So not just running a single agent productively, but what happens when you need to run agents across your entire addressable market at production scale. So first, I wanted to give some brief background on Clay. We like to think of Clay as the creative tool for growth. So put simply, we help you build lists of companies and people from our go-to-market datasets. We help you enrich those lists with 150-plus data integration providers and AI agents, and we help you orchestrate those lists into things like CRM enrichment, outbound campaigns, and more. We do this at quite high scale, and so we run over 350 million go-to-market agents every month. We have a proprietary data set of over 40 million companies and 900 million contacts that our agent researches over. And so you can think of Clay as go-to-market infrastructure for running these workflows. So why is go-to-market a hard problem? And why are we running so many agents? Well, we think in go-to-market that no creative advantage lasts forever. And you can think about this from the lens of cold email. And cold email deliverability rates have been going down for the past couple of years for many reasons. But one of them is just the floor has continued to rise. After GPT-4... You can write human-sounding emails, and you can think about this from the lens of your own inbox, where you're probably drowning in a bunch of outbound emails that may or may not be targeted for the things that you care about. So how do you actually win in this environment? And we believe that the fastest to iterate wins. So you actually need to continuously evolve and build new outbound strategies and plays to be able to actually do better than your competitors. And we call this go-to-market alpha. So similar to in finance, where alpha is outperformance against the market, we believe there's a similar concept in go-to-market. So better audiences, better timing, better signals, better positioning than your competitors can yield to actually great results. So how do you actually get to that go-to-market alpha? We believe there's three levels. So level one is individual AI access and literacy building. So deploying tools like ChatDVT or Claude to your sellers for things like call analysis or outbound copywriting. That's great. But level two is actually centralizing that and deploying it across your sellers. So using flawed skills at the workspace level or after every call, generating post-call notes. Level three is creating advantages that your competitors can't copy. So think about a company. We work with a lot of AI coding companies and many of them build outbound campaigns where they're looking for people. we're looking for companies that are hiring for a head of engineering and have a lot of engineers that have starred their GitHub repo. So these are plays that are not transferable to their competitors and they're unique to them, so reaching that go-to-market alpha. We find that many teams get stuck here at level one. So their sellers might be using tools like Claude to analyze call transcripts or write outbound emails, but it's fairly low leverage because you can write the best outbound email But if someone doesn't want to buy your product or service, a creative email isn't going to actually change that. Much higher leverage is actually fixing targeting. So finding customers who already want to buy your product or service, you can actually get much more meaningful results. So our best customers really do this using a loop like this. Where they'll scan their entire addressable market, layer on signals like news articles, fundraising announcements, or bespoke data points like that GitHub stars metric that I talked about. They'll use agents to score those accounts to find out when is the right time to reach out to them and act at that time. Finally, our customers will learn from those outcomes and iterate on those plays over time. So this looks a lot like an engineering challenge, because you need to run agents across your entire addressable market. And so that's why many of our customers use tools like Clay to orchestrate this. And at Clay, we have our agent, Claygent, which does a lot of these workflows. So it will do things like company research in order to find out, is this account a good company to reach out to at this time? We run this over 350 million times a month. It processes over trillions of tokens every week. And I'm going to talk about four challenges that we've encountered and lessons that we've learned on deploying this agent at scale. The first challenge is on infrastructure, where we actually deploy this in a reliable way. Second is on rate limits and throughput. So being able to maximize our inference capacity without negative impact. Third is on cost. As much as we'd like them to be, trillions of tokens are not free. And fourth is on quality. If our agents don't yield meaningful results, then none of the other points really matter here. So we need to make sure that our agents are high quality. First challenge on infrastructure. If you were to profile Klagent, or probably many of the agents that you all are building today here, most of our agents are actually just spending their time waiting. So they're waiting on, in our case, browsers or APIs or inference. And so we used to run Klagent on Lambda, and Lambda was prohibitively expensive because Lambda charges for wall time. So we moved that to ECS, but ECS is, we traded costs for reliability. So we needed to re-architect our system to be able to recover from things like random host failure, things like that. And so the right architecture looks actually much more like a durable workflow execution. So using things like queues, checkpointing your agent at periodic steps. So using a tool like LangGraph or LangSmith deployments would help here. The second challenge that we've run into is rate limits. We have a lot of dedicated inference capacity at Clay, but our workloads are fairly spiky. And so we need to be able to maximize the inference that we have. in order to productively run our agents. There's so much effort at the inference layer to make sure that GPUs are always hot, and a lot of that gets lost at the application layer unless you're actually maximizing the inference that's available to you. So we've actually built this system with back pressure to be able to adaptively throttle against our downstream inference providers. And it looks a lot like the TCP IP congestion algorithm, where we basically will send as much traffic as we can, and as soon as we run into rate limit issues, we'll... progressively dial back that traffic. We've found from some of the experiments that we've run internally that this can yield four to ten times as much throughput as a more naive system. So it's actually quite meaningful, especially at Clay's scale. We also had to build fairness across our customers because we don't want a single customer who's running millions of agents across their market to crowd out the customer who signs up for Clay and is running their first ten agents. The third challenge that we've had to deal with is cost. And cost is meaningful at our scale. We've built our own agent harness at Clay for a variety of reasons, but one of the learnings that we've found from building our own agent harness is that caching strategies have really meaningful impact on the cost of your agents. And you actually can build agents against those caching strategies to make sure that you're maximizing that. For providers like Anthropic, this can yield to up to 70% cost savings. quite high. The second strategy on cost that we found is actually bounding retries and tool calls before they sprawl. So you have to do this in conjunction with your evals, but we found that many times if you force an agent to return after a certain number of steps or a certain amount of research, it will actually yield better results than if you were to let it run to completion. And so again, you have to do this in conjunction with your evals, but use case specific, this can be quite effective. And the third point is actually measuring cost tied to quality and outcomes, which leads to our fourth challenge on quality. We spend a lot of time on Klagent quality, and we think it starts with great context. So for us, we give Klagent access to great web data and proprietary go-to-market data sets. We have an entire team dedicated to making sure that data set is accessible to agents in a great way. We also tune our agent harness specifically for go-to-market use cases. So we have offline evals. But we also have online evals to make sure that our harness is really targeted for the things that people are actually trying to do in our product. And this is, again, where tools like Langsmith are really helpful to understand why or what you're optimizing for. One additional note on quality is that quality is also a product problem. So we built an agent builder in Clay where people can actually test and iterate on their agents before they run it at market scale. And by giving users these kinds of iteration tools, They actually have way more confidence to be able to run their agents at market scale. And yeah, we found a lot of success with this. Okay, these are the four challenges that we've run into on infrastructure, maximizing throughput, cost, and quality for running these agents at production scale. But what's next for Clay and what's next for Clay's agents? Piggybacking off what I was talking about on quality and context, we really think that agents need great context to do great work. And so we spent the last... six months or so, building a product that we call Audiences. Audiences lets you aggregate all of your go-to-market data into one place. So from tools like Snowflake, Salesforce, Gong, other call recordings, you're able to aggregate all of your data in one place, layer on third-party signals like fundraising announcements, news articles, and more, and give that to Clay agents to be able to run outbound campaigns. Audiences is also... the foundation for our agent memory. And we're using this to build what we call go-to-market intelligence, where agents are actually able to recommend plays based on the things that they've tried before and the context that they have. And so they're able to actually complete this flywheel of improving over time based on the things that have actually worked before. This comes with all sorts of additional infrastructure challenges that we have. Things like... virtual file systems that are able to actually reason over the context that we have in audiences, things like sandboxes. And if you're interested about these challenges, I would love to talk to you after this. To recap, a couple of things that I've talked about. One, go-to-market is fundamentally an engineering challenge. You really want to optimize your agents for infrastructure, reliability, throughput, costs, and quality to be able to get meaningful results. You need to actually run your agents across an entire market to get great go-to-market alpha. And finally, observability is the feedback loop that actually makes these agents better over time. And with that, thank you. Have a great rest of your day.
⚙️ Pipeline jobs
| Stage | Status | Att. | Updated | Error |
|---|---|---|---|---|
| download | done | 2/3 | 2026-07-20 13:40:25 | |
| transcribe | done | 1/3 | 2026-07-20 13:40:34 | |
| summarize | done | 1/3 | 2026-07-20 13:40:46 | |
| embed | done | 1/3 | 2026-07-20 13:40:47 |
📄 Описание YouTube
Показать
Jeff Barg, Head of AI at Clay, breaks down what it actually takes to run go-to-market agents at production scale — not just one agent, but 350 million a month across an entire addressable market. He covers the four hard problems Clay solved: infrastructure reliability, throughput under spiky workloads, cost (including a 70% reduction via caching), and agent quality. He also introduces Audiences, Clay's new product for giving agents the context they need to recommend plays autonomously. Chapters: 0:00 What Clay does and why GTM is an agent problem 0:55 350 million agents a month: Clay's scale 1:10 Why no creative advantage lasts forever 1:44 How to actually win: the fastest to iterate wins 2:00 Go-to-market alpha: the three levels 3:16 Why most teams stay stuck at level one 3:47 The loop Clay's best customers run 4:18 Why this looks like an engineering challenge 4:47 Four challenges at production scale 5:25 Challenge 1: infrastructure and durable workflow execution 6:21 Challenge 2: rate limits and the TCP/IP approach to throughput 7:30 Challenge 3: cost and caching strategies 8:36 Challenge 4: quality, context, and evals 9:39 What's next: Audiences and agent memory 11:12 Recap Extra resources: • Everything we shipped at Interrupt: https://www.langchain.com/blog/interrupt-2026-overview • Meet LangSmith Engine: https://www.langchain.com/blog/introducing-langsmith-engine • About LangChain: https://www.langchain.com/