State of Agentic Coding #8 with Mario, Armin, and Ben
Armin Ronacher · 2026-07-13 · 1ч 20м · 20 928 просмотров · YouTube ↗
Топики: ai-loop-engineering
🎧 Аудио
📝 Summary
model=deepseek-v4-flash · prompt=summary-v7 · 18 902→3 124 tokens · 2026-07-20 11:50:43
🎯 Главная суть
Armin, Mario и Ben обсуждают текущее состояние agentic coding, фокусируясь на недавних моделях (Fable, Sonnet 5), проблемах совместимости из-за RL‑обучения Anthropic на собственном Claude Code, регрессиях в tool calling и растущей стоимости токенов. Спикеры скептичны к hype вокруг «loops» и «dark factory», предупреждают о потере контроля и росте зависимости от провайдеров моделей.
Fable: первые впечатления
Mario попробовал Fable на своих типовых задачах (реализация интерфейсов по готовым типам) и не заметил большого скачка — с ними справляется и более дешёвая модель. Для дизайна системы он не успел протестировать, а в задаче личных финансов Fable полностью провалился (перепутал базовые проценты). Armin дал Fable рекреацию страницы Arendelle.com — в течение 5‑6 часов модель работала, раскрашивала экран красным для детекции теней, но при повторном запуске через Pi оказалась «скучной». Главная претензия — скорость ответа: в Claude Code даже простой вопрос занимал 2 минуты, что ломает рабочий ритм.
Ben использовал Fable через Claude Desktop и заметил новые визуальные возможности (inline preview HTML). Однако он не смог понять, насколько улучшения связаны с моделью, а насколько — с новым harness. После исчезновения Fable он переключился на Opus и ощутил, что Opus тоже хорош, опять же из‑за обвязки.
Регрессии моделей и RL‑обучение на собственном harness
Armin пояснил, что модели теперь тренируются через reinforcement learning на реальном поведении Claude Code. Claude Code очень снисходителен к tool calls — принимает невалидные JSON, ошибочные имена параметров (oldster/newster вместо old_string/new_string). Отсутствие отрицательного сигнала в обучении приводит к тому, что модель начинает систематически генерировать некорректные вызовы. В Pi (который строгий) около 20% edit-запросов от новых моделей заканчиваются ошибкой из‑за grammar constraint sampling: модель «случайно» делает лишнюю запятую, после чего вынуждена выдумывать несуществующие поля, что отравляет контекст.
Это создаёт проблему для всех сторонних разработчиков, использующих Anthropic API со своими инструментами. Если единственным потребителем модели становится Claude Code, качество работы с не‑coding агентами и MCP‑инструментами неизбежно ухудшится.
«Stochastic terrorism attack» на экосистему
Код Claude Code настолько «слоппи», что формирует де‑факто спецификацию, которая не документирована. Пример: навыки (skills) в YAML — Claude Code принимает некорректные переносы строк в полях, из‑за чего пользователи Pi не могут импортировать свои навыки. Разработчику Pi приходится либо нарушать YAML‑гамматику, либо терять пользователей. Armin сравнивает это с legacy‑программами Windows, где баги становились стандартом. Разница в том, что раньше люди обходили детерминированные ошибки, а теперь недетерминированные модели «решают», как работать, и все вынуждены подстраиваться.
Рост стоимости токенов: «peak model»?
Mario отмечает, что с весны 2025 он не чувствует серьёзного скачка возможностей (последний значимый был в октябре 2024). Новые модели становятся дороже, а не дешевле: стоимость решения задач на SWE‑bench растёт. Ben напоминает, что одновременно исчезают дешёвые «checkpoint» версии — провайдеры толкают к использованию только последних моделей. Armin указывает: Sonnet 5 дорогой, но не обязательно лучше в его задачах. Spинкеры сходятся — эпоха «дешёвого интеллекта» заканчивается, наступает AI‑инфляция.
Open‑weight модели: GLM 5.2 и DeepSeek Pro
Mario говорит, что GLM 5.2 работает с его задачами так же хорошо, как и «трёх‑поколенческие» модели, никакого прорыва. Ben добавил, что GLM любит «try‑hard» режим (force push, удаление файлов) — это может быть хорошо для некоторых сценариев, но для его workflow модель «просто норм». DeepSeek Pro показался Ben очень хорошим, но GLM в 3–4 раза дороже. Armin подводит: индустрии нужна «новая история каждую неделю», GLM стал такой историей, но он «не второе пришествие». При этом open‑weight ≠ open source: модели дорого воспроизводить, и зависимость от лабораторий сохраняется.
Loops и orchestration: модно, но не для всех
Peter от Factory AI и другие выступали на AI Engineer с концепцией «loops» — оркестратор распределяет задачи по суб‑агентам. Armin напоминает, что паттерн RALF loop (поставить цель, дать агенту итеративно её достигать с верификацией) существует с июня 2024. Сейчас его оборачивают в системы, управляемые внешними событиями (issue, PR). Ben пробовал авторецензирование через dumb model (Haiku) и находил баги, но в целом не доверяет полному циклу, так как loop‑режим генерирует «самый сложный и уродливый код». Mario критикует: «loops» выгодны продавцам токенов — чем больше итераций, тем больше расход. Призыв «review less» не решает проблему, а лишь снижает порог.
Dark factory и токеновый гиперинфляция
Ben снова поднимает тему: все эти подходы (review, loop, dark factory) поощряют тратить больше токенов. Компании, продающие токены, заинтересованы в таком повествовании. Mario добавляет: Factory AI в маркетинге обещает «ай, все автоматизируется», но на практике требует сложной настройки. Окупаемость для обычной компании сомнительна. Armin замечает: на AI Engineer Peter сам признал, что решил проблему токенов, присоединившись к OpenAI — это роскошь, не доступная обычным разработчикам.
Вертикальная интеграция и потеря контроля
OpenAI недавно отключила возможность кастомного SFT — теперь нельзя дообучать их модели. Способность выбирать конкретную версию модели (checkpoint) постепенно исчезает. Провайдеры толкают к единому «latest» – «будет нормально». Armin говорит, что регрессии при этом неизбежны, а пользователь не может их обойти. Все три собеседника согласны: конечные пользователи вынуждены будут проводить собственные evals, что повышает порог входа.
Стоимость compute и вторичные эффекты
Mario поднимает глобальную тему: гигантские капитальные затраты на дата‑центры уже привели к росту цен на DRAM, GPU, консоли (Switch 2). Его сын узнал из новостей, что AI сделал Switch дороже — молодёжь формирует негативное отношение к AI. Опросы подтверждают, что AI больше нравится старшему поколению (boomer‑менеджеры используют для слайдов). Armin шутит про киберпанковый сценарий, когда люди будут убивать друг друга за железо. Ben надеется на цикличность: сначала рост цен, потом перенасыщение и коррекция.
Thin clients возвращаются?
Mario предполагает, что мечта гигантов – вернуться к модели мейнфреймов: пользователь платит за подписку на облачный compute, а локальное устройство — только тонкий клиент (например, телефон). Ben приводит пример X (Twitter), где теперь встроили функции стриминга, избавляя пользователя от локального OBS. Armin успокаивает: на самом деле вокруг много дешёвых GPU (например, Tesla на дорогах), которые можно кластеризовать для инференса. Эксперименты с распределённым DeepSeek Flash через Lightning показали, что передача данных — мегабайты, а не гигабайты. Возможно, децентрализация станет противовесом централизации.
Прогнозы: когда лопнет пузырь?
Вопрос от чата: когда Bubble burst? Ben называет вторник (шутка). Armin: вероятно, после IPO Anthropic или OpenAI. Mario добавляет: SpaceX тоже имеет завышенную оценку, но держится. Все согласны, что после первичного размещения акций реальность может скорректироваться.
How to handle FOMO?
Mario: он сам «создатель FOMO» (каламбур: создал Flask → FOMO). Практический совет — оглянуться на полгода назад и спросить себя, что реально изменилось в повседневной работе. Если ничего, то нет смысла читать новости каждый день; достаточно проверять раз в месяц. Armin добавляет: большинство хайпа — «притворство, что знаешь будущее». На конференциях чувствуешь дезориентацию, но через день понимаешь — «это просто игра». Ben: ключевой паттерн «while loop с LLM‑запросом и выполнением инструмента» не изменился за год; всё остальное — детали.
📜 Transcript
en · 14 222 слов · 169 сегментов · clean
Показать текст транскрипта
The sloppy behavior of Cloud Code becomes a stochastic terrorism attack on all other software products. Ben, when is the bubble? Tuesday. Good. I have no plans. Do you experience FOMO? No, not anymore. I'm just the creator of FOMO. I'm sorry, I'm just the creator of Glass. He created the FOMO. Hey, Mario. Hey, Ben. How's it going? I'm pretty okay-ish, I guess. I'm surprisingly on a podcast. thought would be just the two of you until the end of time because it works so well but apparently you decided to invite me oh that's very nice i hear you're a fan i actually am a fan like your podcast and the one the cloudflare people just set up recently are the only ones i listen to now sponsored by erendel yeah i'm coming to you live except it's not live um from vienna austria i'm here with armin who's normally here i'm normally here and mario zechner Did I pronounce that girl okay? Yeah, it's pretty good. Close enough. Someone on HECA News today said Marco Zecna. I think as long as it's not Luigi, it's good. I happened to have made the trip over here and we thought it'd be fun to do an episode here together and invite Mario who came in from Graz and to talk about agentic coding, whatever that means. Let's go in order. Armin, who are you? What do you do? this is a ever evolving and more complicated question at this point and the creator of pi creator of pi sitting next to it oh my god it was actually annoying that people now actually seem to think this um no i'm still uh working at erindale uh company started with colin and mar is now as a partner there and we are working on all things such and decoding seemingly at this point Special guess, Mario, creator pie, or I don't know, give us your backstory. How did you wind up here? In the mid 2000s, I started working in the industry and applied machine learning. At the end of the 2010s, I went to San Francisco to do a gaming startup, mobile gaming startup, came back disillusioned with everything San Francisco. Went into management, started another startup with two friends in Sweden. And then I was fun employed for the last 10 years. My machine learning stuff was before deep learning. So in 2022, JGPT came onto the scene that was like a pullback into my old mode of thinking. And I was immediately interested in seeing that because that kind of did all the things we worked on in the 2000s just better. So that was interesting. And over the past four years now, I guess, or three and a half years. It turned out that those things could also be useful for programming or software engineering. And then after suffering through the churn of cloud code for a couple of months, starting in 2025, I decided to be stupid and build my own coding agent because how hard can it be? And that coding agent was then subsequently invented by Armin and called Pi. For the people not familiar with this, I don't know how it started, but we have been trying to trick Grok into pretending. This guy made flask and I made pie. Every once in a while it does actually reply that way. Python is like, I love it. It's just such a good idea to have meaningful inundation. It's great. It seems like a lot of great stuff gets made from stupid ideas. Depends how long you want to maintain it. Well, I guess I should add, this is weird for me because you guys have done a lot of open source stuff. I have this mildly popular open source project that I never planned, I never planned for this to happen, which is HUNC, which is like a, which is a diff tool. And I was just commenting this week that the origin of that was, I thought it'd be interesting to plug in diffs, diffs.com, which had just come out. And I was interested in Open2E, which powered open code, and I thought that was neat. And I was like, does it even, is it even rational? Does it even make sense to plug in, um... diffs has a React exporter, a React library rather, and that OpenTui can consume basically React and put that on the Tui, and would that even function? I just, you know, it seemed dumb to me because we were living in this era, maybe we'll come back to this, where it's like, you know, JavaScript on the terminal, that's stupid. You're familiar with that? Maybe there's a recurring theme there, because this is the era where everybody's just writing everything in Rust, because we can now. Right and a lot of other terminal div tools actually coming out at the same time. We're doing that we're pursuing like building this and rust Anyways, I was just thinking like it's funny how dumb things Lead to interesting outcomes. Is that fair? Well, let me put it that way. I'm actually more of a C kind of person So writing something that a lot of people use in TypeScript or no chairs it kind of is a world model destroying task or effort to the other two key because that's not where I wanted to end up like I've never touched TypeScript that much before apart from some WebGL gainy stuff but it's not the worst ecosystem and today the tools don't matter that as much anymore I think yeah I think like without the agents I probably would not enjoy maintaining a TypeScript code base to the degree that although maybe I would actually enjoy it more without the agent Well, I have to say I actually like TypeScript the language. It just sounds like the ecosystem. So last episode, we didn't do monthly predictions. We did, I don't know, annual predictions turned into rambling, which is usually pretty difficult for us. But we did talk about AI sovereignty. We talked about open weight models. We talked about how nation states are involved. in subsidizing or at least protecting Frontier Model Labs. I'm just paving over things a lot. We recorded that and then Fable came out two days later. And I think that we should talk about that because it seems like... And it was really interesting because I remember when we released it, I was like, oh! we just missed Fable, but then Fable didn't survive from one of my weekends. I think that's right. I think the episode came out by the time... Fable was already blocked. Yeah, pretty close. I tried it for two days for my usual tasks. I didn't experience a big jump in capabilities, but that might just be because my stuff is so small in scope. Every task I throw at a model. can probably be done by any kind of model in a way where I'm okay with the quality. So I don't really feel a big step change with Fable. Could you give an example of what a prompt looks like in that world? Is it build a million dollar SaaS, no mistakes? Here's a bunch of types, interfaces I designed in another session, and that session was collaborative. Now fill in the gaps, implement the interface basically. So that usually works even with smaller models. And using Fable for that is probably not the smartest idea economically. What I didn't get it to use for was design work. I would have liked to see how it behaves in terms of system design work. If it's a better sparring partner for that kind of stuff. didn't get the chance to do that. I actually, I did for some personal finance stuff and it completely failed. Like it was an absolute fail whale. It's like basic percentages. It just got totally wrong. So I had basically Fable in two different experiences. When I released initially, generally kind of impressed by how much it does within the harness. So in particular, I, we have this, if we go to Arendelle.com that are waves, which were mostly sort of hand clanked with the clanker a couple of months ago. I was like, okay, here's a picture. I recreate this. And it kept working for my entire five hour budget until it hit the limit. Over six hours at night. I picked it up in the morning and then I looked through the session. Like initially I still saw the session as it was happening. On the side I had Pi built me a session viewer that I can see the screenshot. Because I noticed that it spawns Chrome all the time to take screenshots of it. what i was really impressed by is that it's it recolored i think every couple of steps it recolored the entire screen red so it can detect the shadows better so they must have like no model has done that before but then i tried it now like three days ago when it came out again in pi without all the bells and whistles and it was very boring but it was actually much more enjoyable for me in pi because when i did it in cloud code it I feel like even the simplest question would take like two minutes to come back. So it was like as a sparing partner, at least with my prompting style, it was very slow. Like just the speed of it affects how you're going to work with something. Totally. Mark, a quick question about whether you were using Pi, I presume, when you were working with it. Yeah, I was actually, yeah, I was using Pi. But I was also using it for the personal finance stuff. I was using it through the web. Cloud AI and cloud.ai web interface. I think boring in Pi makes sense. That's kind of what it felt like. I think I had an experience similar to you, Armin, where, you know, when Fable was announced, I was like, oh, where is it? You know, and I'm trying to restart all the agents to update it or to get into however, like to have it just show up as an option. And the first client for which that happened for me was actually called Desktop, which I hadn't used in a while. And so I use Cloud Code within the desktop experience. And I think I had a similar experience. It was like, oh, you know, Cloud Desktop is pretty good. You know, like there are some new capabilities in here that maybe I hadn't noticed. Maybe because of the way that I work, which is a lot of, you know, I do UI stuff. I'll be like, you know, give me five options. And I noticed that Cloud Code had this tool where it did these like little inline preview visualizations. and Fable was doing a good job of like rendering sort of like HTML mock-ups that were not like a pure HTML document. They're doing, they have like a preview tool, and I really enjoyed that experience, and that's what turned into my open source project, which is called Sideshow, where I was like, wow, that's neat, but now I want this everywhere. So in some ways, I'm not sure how much of my experience I thought it was good, and I gave it some tasks, and I felt like it was vaguely stronger. and maybe, you know, went on further. I did observe that. But I couldn't tell how much of that was the harness. I couldn't tell, you know, gee, maybe this is the, yeah, the harness experience. And when Fable went away and then it was Opus, I just started using Opus on Claude Desktop. And I was like, oh, this is, again, like, maybe Opus is good. And I don't know how much is the harness. Yeah. I mean, that's a theme that's been repeating for the past six months, I guess. In October, there was a noticeable step change. and since then i personally at least for my tasks and your mileage may vary i i haven't felt a big step change like the one from say april 2025 to october again for me it's mostly progressive enhancements and some regressions and some other tasks um i know there's a step changed uh step change into cost so that's one part but i also think like one thing people put up on socials is that the latest models are basically built for loops and token maxing so i guess the most difference you will probably not feel in this kind of collaborative workflow or small scope issue workflows but more in the you have an orchestrator and that orchestrator does ultra code and writes a bunch of workflows with sub agents blah blah blah and my expectation would be that that is where they've focused a lot of their rl this time around for fable Because I think that's also how Mythos is actually supposed to work, right? Like it's not a single session that does it. So like even the smallest task within Cloud Code spawns a gazillion of salvations now. It's nice if you sell tokens, right? Not insinuating anything here. But I think it is sort of interesting in some sense because I think we talked about this last time, but there's basically like bifurcation happening, community between people that are like hands-off ralphing. and then people who are still trying to read the code. And what they need in terms of tooling is sort of diverging a little bit here. Yep. Another observation you've just reminded me on the topic of tokens is by using Pi, one of the things that I think Pi did a lot for my headspace was being very keenly aware of the context window. Sort of just, you know, it's front and center. You pay attention to it, dumb zone, etc. in Cloud Code Desktop and in Cloud Code CLI, it doesn't show the context window. You have to go out of your way to ask. And more than a few times with Fable, I would be like, gee, this has been going for a while. And it'd be like, you know, 500,000 token context window. So I was mainly compacting quite a bit, which may have also impacted results. That was interesting for me. I think it was the first... model where it seemed not quite to get so stupid if it got into a higher token range but it also... I feel like it hides compactions now, does it? I don't even know that I saw it, maybe because I would always get ahead of it, so I'm not sure. In my longer design sessions I usually use a model with 1 million token context window size. And I think GPT 5.4 still offers that. They still offer it for 5.5 and Opus 4.8 obviously offers this as well. I think down to 4.6. And I think for this kind of designy kind of collaborative things where you might have a bunch of markdown files or scratch code files that you're working on to iron out the design, the dumb zone thing doesn't matter that much. Yeah. And then you have the implementation and there it does matter quite a bit even. It stops being good at tool calling. Yeah, exactly. I think it's still reasonably good at recording things. Yeah, but they all still fall on the nose. Like after 250k tokens, it used to be 64k and then it was 128k. Now you can push them up to like with GPT-5.5, I go until the end of the context window and it's usually fine. Yeah, I personally felt it was stronger. Again, I think it also depends what you're doing. I also try to use it for writing and I found that it was quite garbage at that. Really? Because the cloud models are usually the ones I go to when I want to have something at least checked in terms of... I feel like it's writing worse. The new model I think is writing worse. I did one blog post recently where I felt like I want to help it give me an idea of what the structure could be. So I give it letting the LLM write your blog posts. I definitely let the LLM help me structure my thoughts. Sometimes they sort of like put me a draft out and I found it so offensive that I threw it away almost immediately. Was this Fable? That was Sonnet 5. Okay. But it was, but previously I felt like Sonnet was a really good basic writer. Better than Opus from my experience. Interesting. This is, I love how anecdotal this is, by the way. It vibes. It's so vibes. Welcome to the vibe show. Yeah. But I also wrote like a research draft blog post with Fable. And I was certainly, you know, and this is like a multi-prompt affair where I'm like, okay, structure it this way. Hey, let's explore this idea. I also was like, I think I have to throw everything away and write this just by hand. Whereas I felt maybe... I didn't feel like I'd have to completely throw it away before. And this is so vines, it could matter like what is the topic? What are we writing about? I think it's like it's two things. One is like I think I'm much more attentive now to anything sort of generated. I think it's a pretty big part. See and for me it comes back down to my workflows are so fucking stupid caveman like that I don't notice a difference because I personally And first of all, flabbergasted and offended that you guys let LLMs draft your blog posts and help you structure them. That's the whole point of writing this thing, that you structure something in your head yourself. You don't use an LLM for that. So here's my workflow. I usually just dictate paragraph after paragraph, and I'm the guy who lays out the words and lays out the structure. I mean, I guess that's what I'm saying. That's what it looks like to me. I mean, we're just using a different language. I think so. There's a difference. If I already know what I want to write, I just write it down. But for instance, for the blog post where I did this was the... where i was writing about the loop thing and there i wasn't sure how i would like want to fret the story so like i was asking the llm to just give me ideas of like how i want to convey this yeah and it's just this like i feel like this works at one point where basically like hey these are the points that i want to convey like give me some ideas of how this would go and i would at least read through it i was like like i don't feel offended reading it and now i feel like offended reading it for me, when I do this with the dictation paragraph after paragraph, once I'm done with a section, I rework the section and see if it fits in the rest of the sections. My point being, I don't feel a difference because the only thing I ask the LLM to do for me is take the dictation and put it in a blog post in the it's been a marketing file. And then once I'm done and have re-triggered everything and then let the LLM move stuff around for me, that's the second task it has. And I use the LLM to say something like, can we shorten this? And then I ask it for suggestions on how I can shorten a paragraph. And that's why I never feel a difference using a different model. It's the same with code. I used to make one pass at the end, like, hey, fix obvious grammar mistakes. And I don't do this anymore because I feel like it's way too adventurous in changing even words around. Like all of that that I feel like I would... I felt like I treated it like a spell checker. Now I don't trust it anymore. So I don't even treat it as a spell checker. But you're treating it as a whole document spell checker or paragraph by paragraph spell checker? I feel like in the past it didn't matter because it's like, hey, I wrote this entire thing and now please fix bad cases, put commas, that sort of stuff. And now with the same prompt, it's like, well, I really felt like you should be writing this instead. And it rearranges entire sentences. And it puts words in that I wouldn't use. Mario, I think that we actually worked very similarly with this to the degree that on this post, and I want to make a comment here that I'm not some prolific writer. Like, what vlog posts are we talking about here? I actually just started editing it with Kimi K2-6 because I just rationalized to myself that my editing is almost so micro that I actually prefer a fast mechanical model to just, you know. Yeah, exactly. Right? So... I guess maybe that's from going where it's like I didn't see an advantage frankly and maybe that's... Oh yeah, definitely. A few episodes ago I made a comment that was like I have a theory that we're actually at peak models right now. I'm on your side, man. And that was an argument of just sort of like will we even know what better looks like? This conversation is almost a good example of that where we go and you go online and you're like, it was better for me here or was it or I don't know what I'm talking about my my my recent obsession with reinforcement learning here because I think like this is a data point on like peak peak model So at AI engineer I went with a bunch of conversations that were roughly related to like Some labs might be over RL-ing their models we sometimes throw terms out really quickly. Help me understand, because even I'm a dumb dumb. What are we talking about by reinforcement learning? Like, what does that mean? Yeah, so when you start it with a model, it might be able to answer you some questions, but it doesn't necessarily understand what other than giving you in terms of text back it should do. And through post-training, the models learn what an appropriate response for a given prompt would be. And for a lot of programming setups, this looks like here's my prompt, call a tool, let the agent harness execute the tool, feed that back in until finally you reach your reward condition, which is usually you committed or something. So the reinforcement learning is the process by which the model basically learns tool calling, what sort of unit tests are, like the whole shebang until problem solved. And this is always a true also to some degree. just giving you like human outputs and not just tool calls, but like specifically right now in the context of an agent, a lot of it has to do with tool calls. Humor me for a moment to not go too deep, but just like maybe to help folks understand. And I think even for me to understand. So model companies have basically some massive suite of tests, integration tests, who knows that they're doing the reinforcement learning? I assume that they're not starting from the field. What they would basically do is someone somewhere solved the problem. So there's an ancient trace that started, that like a human eventually made to completion. So you have initially sort of the prompt that triggered off, and then you sort of have an idea of what the end result should be. And ideally you also can directly check out the GitHub repository with the state of the world when that whole thing started. And then they throw that entire thing into a reinforcement learning environment where they run thousands of these simulations simultaneously and feed sort of the GPUs as this whole thing executes. And there's a reward at the end when the problem was solved. So, I mean, this is, I think, roughly the way you should think about it. And the reward is energon. All you can do when training a model is basically modifying its weights and biases. So the reward function is a signal for what's good or bad and if it's good you want to emphasize that behavior and if it's bad you want to de-emphasize it. That means that reward function needs to translate into changes to those weights that the numbers inside the model that make up the information. So here's some useful signals of bad data points. It's not just good to like here you committed. There's a bunch of things which an LLN shouldn't do. So if I ask it to produce JSON, and it doesn't produce JSON, there should be a signal going back to the LLM, so you did a bad job. In theory, there's a whole bunch of stuff that the model should learn over time over how good and bad outcomes look like from basically sampling the whole thing. To go back to the story on what started my going into a rabbit hole was that people pointed out that the edit tool in Pi wasn't performing as well. And that was confusing to me because I did not have this experience of the PI edit tool failing in any entropic model I was using, but I was also not using it significantly with Sonnet 4 and Opus 4.8. And by edit tool, you mean the tool, the fundamental act of like editing... Editing a file. Okay. It seems pretty critical. Seems pretty critical. And one of the things that PI does, which I think is a good thing, is that if the add-on produces bad output so it produces a tool call that for some reason is not able to execute. For instance it wants to make a change to a file but the reference string is actually not in the file so it hallucinated something. Pi will be relatively strict with regards to what it accepts and if the model gets it wrong it will error back to the model and say like hey you did a poor job try it again or differently or whatever. It will tell LLM what went wrong and so they have this in context learning thing going on where hopefully it gets better at calling this tool as it's going forward. So it will to some degree fix up some things because models was a terrible at understanding white space characters. So it will allow different white space characters to appear within an edit match to replace it. So that's a thing that it so fixes up. Seemingly newer entropic models are worse at doing these tool calls and I didn't expect that, also couldn't reproduce it for a while until I found someone, gave me some data to reproduce it. And then I was kind of shocked how it works because what seemingly is happening with, and this is a little bit of this hypothesis, a little bit of sort of thing you can measure. So there's a concept in LLMs called, I guess, grammar constraint sampling, which is if you pull tokens from the GPU, there are certain probability tokens that you can get. And the highest probability token might not be a token that's valid for the position of the text you want to pull. So the simplest version of this is like, if you want to produce a JSON object and you know there has to be a JSON object, then the first character, first token you have to pull has to be an opening brace. Any other character would be rejected. And so GPT models, for instance, represent all the two calls as JSON objects. In entropic models, it's a little bit more complicated, like top level. tool calls or strings and then if you have arrays of like if a parameter is an array then it's a percent as a JSON object. It's a long story short they're using grammar constraint sampling for a lot of these parameters in function calls which are complex objects and so what is very interesting is that PI's edit tool is an array of possible edits that it can do. There's like an array of objects first parameter is usually old string second parameter is new string and then there can be a third parameter. I think in pi there is no third parameter, actually there's only two. But more importantly in a plot code there can be a third parameter which is optional which is replace all. And when you do grammar constraint encoding, decoding, it means that it samples one token so that advances it. So let's say you have old string column, the value of the old string and then you sample a comma. That means the only possible valid other string that you can do in this chasen dictionary is another string which is the next key okay does it make sense so far i think so so if i if i pulled the comma i have to produce another key because i cannot close because chasing objects do not support trading commas yes okay so that is that is what happened that is the deepest technical i think we've done on this show but now imagine it samples this comma by accident So now it has to make the option fit, which is no longer valid in PyZeditool. And so we have found out that if you bring the session in a certain state, 20% of all edits that continue fail, and it makes up completely random keys that are completely invalid in the tool call. So it says like, okay, your string is now old string, value of the old string, new string, value of the new string, and require unique, or tools or it just makes up complete random strings. It is forced to do that. The sembler from the output tokens must conform to the grammar of the output language, in this case JSON for the tool call. And then you're fucked. And the real interesting part is once it has emitted that wrong tool call, that stays within the context. So that explains why You can't easily reproduce it because you need a session where it is in context, most likely because that then poisons subsequent tool calls because the model then sees in context there was a tool call that looked like this. So I'm just going to do it again. Or I try it with a different kind of additional property that's not allowed. Now, pi can be lenient when it comes to checking validity of the outputs, the tool call with respect to the tool schema, the input schema. I'm not sure if you have implemented that. So it's just not turned on for this particular tool. So if you make this tool sort of accept additional arguments, then it would sort of not do anything. But I think the interesting part here is A, old models didn't do that. So this is a regression on your models. The second thing is what changed from the old models to the new ones? Like what actually changed? One is they're training on their own harness. So cloud code is the reference harness now for these models when the older ones are trained with some sort of made up harness that they had. And I don't know if it's exactly Claude Code, but at least it's sort of clearly trained on Claude Code's behavior. But the second thing is Claude Code at this point is incredibly lenient and I would dare to say a little bit sloppy. So there's no signal in the training process that you did a bad thing here. Because Claude Code, like I looked a little bit at the deco and deminified code just to see what it accepts. Did you do an illegal thing? No, it's not illegal. The fable did not fail. Is the US government coming after you now? But so like there's no negative signal in the training process anymore because like the cloud harness is so willing to accept a whole bunch of nonsense. Yeah. And the cloud code harness has an edit tool obviously as well. Yeah. But it only takes file path, old text, new text and replace all or something. But for instance, the cloud code harness. You have to have a multi-edit. But for instance, it clearly says like your parameters are called old string, new string. And yet it also accepts newster, oldster, newster, through oldster, newster. Like it is very lenient in what it accepts. So they know the deficiency of the model, or at least the agent who wrote that tool in the cloud code harness was told there's a problem. Right. And then it fixed it by being super lenient and accepting any old garbage the model outputs. I think it could have been you who highlighted this for me, but I'm not sure. Months ago, you know, cloud code will write a markdown. header metadata with extra parameters that technically might be like out of the spec. But because it does it now, people are modifying, you know, their own applications to support this sort of alternative cloud code spec. Yeah, that was a while ago. So skills, right? Who invented skills? Anthropical, fiercely invented skills. And there's even a spec. And that spec says the header in your skills file, which is Janel, should be Janel. So you as an engineer are like, okay, which YAML, which version? Not specified. So you go with whatever's the latest and implement that. And then you have users who bring in their cloud code skills to Pi and say, my skills don't work. Yes. Why don't they work? Oh, there's a new line in the description field, which by YAML's grammar isn't allowed, but cloud code happily swallows that garbage. So now we have a spec by the company that created that spec, and that spec is wrong. And the software implements it leniently and doesn't adhere to the spec. So now everybody else has to adhere to the same slop. I'm thinking of a parallel to almost like old Windows applications. Infinite compatibility, backwards compatibility. You know, and maybe young people don't know this, but at some point, you know, you're a games developer and you, you know, the Microsoft DirectX or whatever says it's doing one, you know, one thing, but... Everybody in the industry knows that you basically have to code around it in a certain way. Then you have to send the signal via the keyboard controller to get more than one megabyte of my RAM. Many, many... Well, we are all very old. We've seen some shit. A little bit. I guess there is a history, obviously, of programmers, you know, writing around these problems. And so then I know when we were doing Century SDK development, lots of that. I guess what's interesting or different about this time is it's not that Anthropic or anybody has designed a spec and, you know, we're working around deterministic code, which is what we were doing in the past. It's that the non-deterministic solutions we have built have just sort of decided to go and do it this way. And as a result, we are all... yeah we are all the sloppy behavior of cloud code becomes a stochastic terrorism attack on all other software products that want to ingest skills this is by this i think it reaches a new level to me is that first of all we have a little bit of evidence now that this is actually sort of getting a little bit worse because like if an if an agent does not fully align with like cloud cause view of the world then you're going to be punished by that yeah but also there is no documentation on it like this There is no spec for it to begin with, but I also don't even know if the people at Anthropic necessarily know what their thing is doing. Nobody ever probably thought about whether the metadata header in a skill file parsed by cloud code is valid YAML or not. Nobody gives a fuck. So it's just live code. But I want to say something to that effect too. What does it mean for API consumers that send requests to Anthropic with their own custom tools? that are nothing like cloud code because they're building agents that are not coding agents but do something else. What does it mean for the error rates of these agents and their tool calls if Anthropic is now basically going all in on cloud code as the RL harness? That's bad and I think you alluded to earlier that MCP might also have a problem with that approach. If the only consumer of that model would be cloud code, even in that world cloud code doesn't control all the tools. it's because it can load MCP stuff in there which registers as tools. And so the quality of CloudCo to be able to invoke MCP tools is also, I mean I don't have a particularly good statistics on that but I remember at the conference we were talking that people exposed challenges with modern models ability to actually also reliably invoke MCP tools because if your training data does not incorporate your MCP tools but it has crazy amounts of go to the internet and do web searches or launch Chrome or do something. The likelihood of the model doing that is going to be higher than the model picking your MCP tool. So there's a within, the more you do this reinforcement learning, the harder you're going to make the life for everybody who is trying to use it as it's sort of as advertised, which is like it extends, it calls anything you want, and you can build your own future on this model, but simultaneously plot. makes or entropic makes most money with plot code at this point so the incentives within the company are like a little bit I don't know intention at the very least so I have a I have a related question that I want to ask which seems highly relevant to this which is to go back like Mario said when we were he was trying fable it's like boy it's really bad at this financial stuff right or budgeting One of the first things I did on Opus 48 was someone on Twitter said like you asked that how many which days of the week have the letter D in it and I said like yes three Monday, Tuesday, Wednesday. And I was like are you sure? I was like oh yeah first D also the D or like a variation thereof. It's because it doesn't invoke. It does not invoke for that specific kind of prompt like a trivial prompt but for the prompts I gave it and especially the data files I gave it a bunch of CSVs. And it obviously wrote code in the background that it executed in a small little container somewhere. So that wasn't the problem. The problem was just that it was a pretty long planning session. I can't say how many tokens because they don't tell you. I regularly ask it for a summary of what's been discussed so far. So it keeps things fresh, so to speak. And eventually it just started deviating from the percentages we calculated previously. And it just couldn't remember them anymore. That might be an artifact of the web harness, if you want. But I don't think it was because I then also took that data and put it into PI with... That wasn't Fable then, it was Opus, I think. Yeah, anyways, I wasn't impressed. At least if I think about a non-technical user using Fable through the web interface or a desktop app, I'm not sure about Cowork, but the desktop app is basically the web app. Yeah, so my question was, did you see that as a regression from previous models? No. Okay, I expect there to be regressions. Okay, well then maybe this question won't make sense anymore, but I'll ask it anyways, which was, you know, let's say that you are anthropic or open AI and you've identified, you know, some valuable set of operations, i.e. writing code right now seems to be one of them, you know, but there's only so much intelligence, let's say, that you can pack into this. Do you, you know, almost like you let some parts go? Right so that you can reinforce harder on the more valuable parts and so when you brought up the financing I was wondering like, you know, maybe there's a humorously an accountant at a topic who is like, you know, the math just doesn't work on finance So we're just not gonna focus on that, you know, just put all the GPUs on code. Is that a thing that happens? That's like my genuine question. I mean we would like we said we had any in insights there because we don't know about the training regimes we don't know about the mixture between what goes into pre-training what goes into supervised fine tuning what goes into RL we have no idea but we can make some guesses some educated guesses the first guess is model size there's probably still room training bigger and bigger models but eventually you need to serve that model right and then economics of scale hit you. So that means even if you have a humongous model with I don't know how many trillion gazillion parameters, you probably have to distill it down for it to become economical to serve it. Even if you manage to scram more information, more abilities into the big fat model, once you serve it you've got to distill it down to something that you can actually serve at economical prices. So I think for that reason, it points back to your earlier point, we might be at the top of the S curve. in terms of capabilities. Yeah, I guess I was wondering, let's say there's only so much you can pack in there, but over the course of deploying this model and talking with customers and doing your forward-deployed experience thing. I don't think it's... knowing how big organizations are, I don't think anyone actually has the power to sort of steer something in one way or another. I think what might actually overwhelm is like, what do I have access to in terms of data? Where I'm going with this is like, let's say you have a SaaS product, okay? Old-school deterministic SaaS product. Lots of products, sunset features all the time. Not economically viable, we couldn't get enough customers on this, right? But when that happens, hopefully and usually, we're a good company, sends you an email, lets you know, hey, you know what? Google product 4000 is shutting down, right? And you're like, nuts. But at least, you know, you can make an educated decision about what to do next. Are we possibly entering a world where you can, you know, you can basically, you know, these models serve different purposes, I guess is what I'm trying to say, right? But you, you know, you could sort of adopt a model for a purpose like finance and a new version of that model could come out. And then we just sort of, you know, vibes-y identify that. It's sort of ambiguous, you know, like we're not getting a release notes that says, by the way, we just made this aspect of its intelligence 20% lower. So I think we sort of have the opposite problem, but for exactly the same outcome, which is imagine for a second you have this model. So I think coding is sort of like, it's fine, we've understood this. But imagine in December, and I think we were talking about this in around December, was like, hey, people start figuring out they're really good at security research. It was long before Memphis. Yes, yeah. And then it got so good at security research. that if you are now going to get to the latest models you cannot use it at all because it basically says like screw you this is getting too risky i'm not going to serve you you need to now be one of the hundred companies that the u.s government gives you yeah you're right yeah but i just want to come back to um the abilities of the model and regressions right like how can they evaluate that basically how can i ensure that once I train the new line of models that they don't regress in specific verticals, right? And obviously they have a metric ton of ablations running there and ablation studies and blah, blah, blah. But ultimately whatever they have can probably not cover all the real world examples or all the real world usage that's out there, even though they have those traces probably as well. And I think it's going to become even harder in the future if they add more capabilities to ensure that old capabilities that we started relying on in our workflows are still there. And I think that's also why they are pushing for you have to use our harness because that's the only controllable thing that kind of can keep this in check. And that's a rabbit hole I don't want to go down to. Yeah, but I think it is quite likely that if the models are... This is why I think this reinforcement thing is like it's really on my mind because like... if they're doing more and more reinforcement learning on the harness, they are going to increase the capabilities of that. But at the cost of seemingly deviating from this path. And so if you have a less capable model that's not reinforced learned to death, then you can actually maybe solve some sort of adjacent use cases with a quality that sort of stays within more or less the bounds of what the expected path was. Whereas now, seemingly, if you're getting a little bit too close to where they're really trained on but not quite your experience is going to be worse and my suspicion is that this is actually totally okay with the model providers because they don't have to sell you that model because you want to buy it for some other reason like obviously we have all these um You know humanities the last exam and these benchmarks that exist right and I think you were alluding to this which is there, you know There's a ton of benchmarks However, if you think about the multitude of tasks that people are being pushed upon for every random bespoke way that they're using LLMs they could have this real deviation and You know from model to model and they may not understand it and that and I was just chuckling to myself because the answer is like most other things Everyone's going to have to run evals, even end users. And you know what that's going to cause? That's great. Yeah, I don't see a future where this changes, to be honest. Obviously, all the emails we have available at the moment are basically just proxies for capabilities. They will not necessarily translate to your own use cases, but they're a good kind of, like, ignoring all the benchmarking that's going on. And most of the coding benchmarks are basically useless now as a kind of... I think that the coding benchmarks are even worse now in my head because they basically just really care, like, did you solve the problem? And so, like, I think the Sonic 5 release was an interesting one because Sonic 5 goes really well on some of those benchmarks. But if you look at how much it costs you to actually complete the benchmark, it's more expensive than Fable. Because the other thing is, like, the cost of solving problems seems to be going up rather than down, right? Funny how it works. That's a good point. we could get into that, which is the benchmarks are only capturing one dimension. Not at all. I think artificial analysis also takes into account cost or at least number of turns or something like that. There's definitely benchmarks that try to be more truthy, let's say. But like terminal bench or something like that, that doesn't take anything into account. Not cost, not runtime, not anything. We should address the fact that Fable went away. We mentioned it did. It went away. It's back. It could be gone again. Jesus. And it will be gone from the subscription. Don't forget that. Yeah. So it goes away. Even that's an interesting topic, which is we have been so, you know, like model provider subscriptions have been one of the biggest ways that people have experienced these models. And this feels like for the first time that. You might not have access to the sort of like Frontier model on the subscriptions. Seems that way to me. I think they're price testing. We'll find out. But regardless, when Fable went away, you know, everyone was like, this is concerning. Concerning, period. Can people just take my models away? What about open weight models? And then GLM 5.2 came out. A lot of people got really excited about that. What I saw for this window of time was a range of opinions on GLM 5.2. Obviously, there's the benchmarks that looked pretty good. But people saying, hey, anywhere from this is Opus to this is Fable quality. Let's talk about that. Is it that good? Is it, you know, should everybody... You give your lips first. I'm kind of tired of that. Can it make 3DS scenes? Yes, it can make very nice 3DS scenes. The question is, does it... Does it work with my tasks? And just like the models of the past three point generations, yeah, it works. I don't feel a lot of difference, but again, that might be down to my workflow, my code basis, and so on and so forth. Your mileage may vary if you do more complex things than I'm doing. Ultimately, I think that Armin's always saying open source wins, always wins. Models aren't really open source, they're open weights. It's quite a bit different difference. So it's really hard to reproduce them. It is also very costly, even if you had all the training process and data in hand, you could probably not afford reproducing that. So it's far cry from open source. Ultimately, we are still dependent on a bunch of labs and their kindness or their, let's put it that way, genanigans with respect to markets by throwing those out for free, basically. But again, I'm a proponent of open-weight models and I think a lot of people not only in Europe but also the US have started to wake up to the fact that it's not happening. Alex Karp now goes and see NVC to complain about closed-weight models. I have not a lot of opinion on GLM. I think there are two things that I noticed. One is like it likes to... I think it has maybe sort of similar kind of vibes to some other models with regards to... it just likes to succeed no matter the cost. So it will force push, it will delete file, it will do a whole bunch of stuff that's maybe a little bit out of my comfort zone of what I want a model to be. And so it's actually probably a really good programmer, but I also found Deep Seek to be pretty good. Like Deep Seek Pro is actually, to me, felt really good. And from that experience, GLM, I didn't see a huge difference, but I also don't loop, like I don't really do any of those things. It's fine. Yeah, I think for the workflows that at least we have established all of the models are fine now, which is kind of good. I think an interesting thing is that GLM is significantly more expensive than DeepSeq Pro. This might be that DeepSeq Pro is also sort of subsidized a little bit right now on the providers, but GLM is seemingly like between three and four times as expensive in practice than DeepSeq Pro. So is it four times as good? I think I got an OpenCodeGo um subscription just to use it because for a brief period it was only on go and then i eclipsed that limit very fast and then conveniently by the next day it was on the pay per use on zen and so then i was just paying for token and i felt like it was going pretty well um i guess it's just more vibe stuff i didn't think it was the second coming of christ i this industry needs a story every week you know and Conveniently GLM had I guess like a moment where this story was really good for them and not to take anything away from the model, but it is not the second coming It's a good model, right? In a way not much has happened the entire last year other than the models got a little bit better. We figured out that we can put the agents into like different setups like you can put them into Discord, Slack and they can do a bunch of different things but like On a very fundamental level, even this discussion about looping and orchestration, whatever, was actually there pretty much exactly one year ago. Sorry. Just not to the same amount of people being involved, not to the same amount of quality of what they can do with the model. We learned a little bit in the process, but actually somewhat, somewhere we have been there. Nuance. What Geoff, is that the correct pronunciation of his name? Geoff? Geoff? Geoff? Geoffrey. Geoffrey. What he did with Ralph Loop was just basically create a pattern that has its usage. Like all the research, the thing Karparsi and then later Toby from Shopify codified. That kind of came from the Ralph Loop, I would say. Basically set up a goal for the thing and let it do with the thing for as long as it needs to with some verification at the end that it reached the goal. And that's a super useful pattern. Like I've used this quite a bunch now for performance optimizations for... I don't even remember anymore. That's one style of loop, right? And that to me is known good. I would use this pattern. I have actually used for that pattern. Interesting enough, I think the Ralph loop came out when? In June last year? And then it had a renaissance in November because... I think because they put it into the marketplace. Yeah, something like that. Also with the original name. Right. And then now you have Boris and you have Peter talking about loops. And now everybody's going crazy about loops. And if I understood correctly, AI engineer in San Francisco last week was basically all about loops, right? Well, a lot of loops. So was there ever a good definition of what a loop is? I mean, I think they're basically two loops. There's the in-agent loop that sort of naturally ends. And then there's the harness level loop that... keeps the damn thing going until some externally defined condition arises. Which is basically still what the raw focus is. I don't think it has that dramatic... I think the main thing now is like it's actually cues in a way. People put a bunch of tasks into some sort of system where it picks it up and then starts looping. Right. So what Peter was presenting at AI Engineer, I think it was like a 15-minute sequence where he kind of, kind of, sort of went into how his loops work. And it's basically an orchestrator that then distributes work to other sub-agents. possibly based on an external trigger like an issue coming in and PR coming in and stuff like that and that to me is the new style loop you have external events that trigger a predefined thing basically. But I will say did you ever look at Trophy's programming language that he built in August last year? I think I looked at the website was disgusted and went away. But as far as I remember it was basically one Ralph loop hurting issues over and over and another Ralph loop busy picking up issues. Which is I think in a way sort of maybe a crap and like the name of beats which is also basically people are talking about the seriously now and i think like even peter says like if you didn't expect to do it like um sort of say like steve jaggi was right but that's literally what he said and i think it's like it is obviously performative art to some degree but like the underlying principles are not completely wrong it's just i am not a looper i don't know let me spend i'm not a looper i yeah i am a insane ADHD, you know, multiple thread conductor. I think objectively I do produce a lot. I also take, you know, people use the software. I think that's worth something. But, and the way that I achieve it is mostly like, you know, it's almost like spinning plates, right? Like, how you doing over here? How you doing over here? Okay. All right. All right. Oh shit. What's going on? Let's fix you. Right. Okay. It's not healthy. There's certainly some world where, you know, if I could relax more and use loops, I was interested in that. But I just don't know how to achieve it other than for the problem that you highlighted, which is with some deterministic outcome, right? So I do have, for almost all my open source projects and closed source projects, performance benchmarks, unit tests, right? Hunk, for example, has like 13 different tests. tracking memory growth and you know startup time and scroll time, right? And part of the reason why it's even mildly fast is that there was a whole bunch of auto research loops in Pi that was like improve this. You're goal, you're right. Totally works. What I have yet to see or what I don't understand is like I understand how to construct loops. I just don't know for what kind of tasks or for what kind of outcomes, other than like a deterministic numeric one, you know, am I missing out on? Maybe I'm misparaphrasing Peter here, but the way I understood, for example, is something like you have your issue tracker, you might go into the issue tracker or you might tell your agent, go into the issue tracker, identify a bunch of issues that are valid, open them back up and distribute that work to some other agent to then, process that issue and come up with a PR for the issue so I, the human, at the end can just go through the PRs, see if they're good and say, okay, merge this. I have tried this a long time ago and the models might have gotten better, but I don't trust that this leads to any kind of maintainable software at all. I understand this better though, because what he's doing is he's saying, I don't trust what you're creating, but I put some, you know, some value in that you may have identified something. So I'm just going to regenerate everything in the tools that I do trust. Yeah, but I don't trust the tools that he trusts. And I think that there's, I think, a little bit of the delta here because, so one thing finds that I think, I don't know, we already discussed this, but like this, the simplest loop that I can imagine that they can conjure up is that task comes in, generate. Now, start from scratch, review it, feed all the review findings back in the generator until eventually the thing achieves some stable equilibrium where the other language is satisfying. Because it always finds something that needs improvement. And that produces the most complex and ungodly code possibly imaginable from my experience so far. And yet that is a loop that a lot of people run in. Yeah. So there's probably a version of this loop that's not quite as crazy and I think one of their... suggestion that it has like you should be running the review but without thinking because if it has lower reasoning budget then it doesn't hang itself as much. But like it's all kind of like I don't really yet understand which tool I can actually use that I trust so much that because he said earlier like it's a I don't know he didn't say anxiety inducing but it sounded a little bit the description of like doing orchestration on that level is a little bit like you wish he had more free time and maybe Luke's give it a free time. or like sometimes we were thinking last year is this like the more people started doing all of this the more i felt like all this great free time that i had a little bit is like completely gone and at this point it's just pure anxiety and so the loops are not in any way a future it feels like fills me with the idea that it's going to get back to um like being in control i don't know it's just yeah i think there's another another um side of this in the discos The more you have running in the background, the less control over you spend, basically, you have. And the question is, people promoting this kind of workflow, they obviously don't have any limitations in terms of tokens. Maybe OpenAI has... Peter's talk was literally like, the first was Tokuswap. I solved this by joining OpenAI. And then he went to the next thing, which was like, now my only problem is orchestration. And then the question becomes, like... Peter, obviously his system works for him because OpenClaw is still running as it should. And it's super big in China still. I just recently had a talk with two native Chinese people and they were like, oh my God, it just changed the entire landscape in China. So I'm not disputing that it works for Peter. I just wonder, does it scale down to us? Does it scale to your run-of-the-mill enterprise company that may have a software department or which is mostly software? And given today's pricing, I would say it probably doesn't. Specifically when you then compare it to the actual return on investment you get. Like, does this produce things that increase my revenue stream? Or reduce churn or whatever, right? And I question whether that's the case. I mean, for an open source project like OpenClaw, awesome. If OpenAI basically sponsors the tokens to use to keep OpenClaw running and becoming better and more secure. specifically, which is what they've been working on. That's great. Love it. I just wonder, OpenClaw itself doesn't need ROI per se. Like they don't need to make revenue and they don't have any liabilities either. Which doesn't mean that OpenClaw is garbage or anything because they've put a lot of work into security over the past couple of months and it paid off. It's just the incentive structure is different here. And I don't know, I'm kind of feel like we're being sold a version of how things should work that are not to our benefit but to the benefit of people who sell us tokens. Let's put it that way. A common theme of this show. But so far it hasn't been completely wrong, I think. Yeah. And I think that this one of the ways in which this would hold on the token factory right now, the dark factory, is that well... We have reached a point now where the the the we have reached a point a couple of months ago But the general point is like now we are sort of piling up on reviews. It's like and and so the answer is like review less Right it seems to be sort of the mode right now. We haven't solved this problem with the reviews so now The recommendation is just like not everything that you're shipping needs to be reviewed and I don't think that's actually a solution. That just happens to be personally do think though that a lot of reviews in the past before agents were purely performative and weren't actually used. Yeah, I also agree that you don't need to review every single line of fucking code an agent generates. I think that's not, that cannot work, right? I usually give the example of Pi as an HTML export. I have not read a single line of how that export works. I just look at the export and it looks correct or it doesn't. That's it. I don't care. It's not mission critical, right? So and some companies are actually now going full in on the dark factory pattern like Factory AI. It's in the name, obviously. And you will have their marketing material. And then on a Twitter thread, you will have Eno, who I really like. He's a great guy. I had a call with him a while ago. Where he's like, yeah, obviously it needs a little bit of setup. First, it doesn't work out of the box. And then I wonder, what does this entail? And what are my results after the setup? I don't know. I'm not ruling out this is going to be a part of the future or the future. I wouldn't want it to be the future for various reasons, but I don't see it at the moment. Just had a random thought on reviews. And maybe it was ambiguous because sometimes we're talking human review, but also agentic review. Randomly thinking, I've actually had quite a bit of results using dumb models to do reviews. And my argument... I don't know. Sometimes I like to think about what would this look like with humans? And, you know, sometimes having junior employees look at stuff, is half of it wrong? Half of their suggestions make no sense? Sure. Do they occasionally, maybe because they're just, you know, addled on something else, do they identify something that maybe you wouldn't notice? Also true. It's also a way for them to learn. Yes. And this is also something I experimented with this with Fable. I would explicitly say, like, run a haiku. sub-agent review. Boy did that find a lot of bugs, which also made me think what the heck is going on in Fable, right? I don't know what the point of this was. I guess I do want to come back just to punctuate last episode we talked a lot about this isn't 7D chess. We have a company that benefits a lot from you spending more tokens. Isn't it convenient that a lot of these products seem to come out that consume more tokens? and that the solution for all of your ills is to spend more tokens. And we have two things that came out actually that do this, which is Fable, of course, but also Sonnet 5. You touched on this earlier, but we didn't connect it in this way, which is it's better purportedly. I mean, it probably is. Sonnet 4.7, at least on some benchmark, but also, surprise, it also seems to, you know, use more turns. It also costs as much as GPT 5.5, I think. And it's definitely worth the GPT 5.5. Who's going to... Yeah, I mean, but it's not the only thing that has to cost more tokens now. So, the SWE bench is a good example. There's websites where you can sort of see people running them independently and also give you like token cost. So basically, cheap models are falling away because they're being replaced by slightly more expensive cheap models. So, like the cost points that they got at one point for... Heiko 3.5 has not, of course there's like DeepSeek Flash, so there's a cheap model appearing, but the flagship cheap models are, the cost of solving a problem actually getting more expensive. And so, yeah, you've seen this with Gemini Flash as well. Yeah, or just SaaS products. Like, you can't go back in time and use the version of Netflix that was totally great that cost half as much. Yeah, I guess we have a little bit of an inflation problem. Just, you know, yeah, tying it back that these are not even some grand conspiracy. This is just, you know, traditional marketplace dynamics. I remember in the early beginnings of LLMs available to us, perhaps you had things like model checkpoints. You could select a specific version of a specific model from a specific date and use that in your production software or backend system, whatever. And I think they're still kind of doing this, but it kind of slowly goes away and you kind of push towards just use the latest whatever, right? It's going to be fine. And we know it's not. We see regressions. And another interesting bit was 20, I think around the end of 2023 and beginning of 24, there was a large push towards do your own supervised fine tuning on our models. And then just a couple of weeks ago, OpenAI actually, turned it off. So you can also not customize them anymore. Probably because SFT is really hard to do well and the outputs are usually garbage anyways. But still it just points in the direction where they're trying to become vertically integrated. They want to own the entire stack probably because they know that just selling tokens is not enough. Our friend David Kramer highlighted the fact that if you're working with agents you benefit a lot from determinism. We talked about, you know, hey loops with a deterministic output you can have actually really great results because you're just smashing tokens till you get the answer. What that also means, so that can be unit tests, that can be benchmarks I alluded to earlier. However, to drive all those things you also need compute. They're not tokens and we've historically thought about, at least in the last few years, as compute as being increasingly cheap. However, compute... is also going more. We talked about we shouldn't mention the price point of like how much is left of this. Oh how much is left of this now? Yeah, yeah, we've talked so much about computers being expensive. I don't even think we were prepared for literally devices in your hand. It's almost like used cars when used cars went. through the roof. I don't know if they did here, but... I don't know if we discussed this on this podcast, but I did talk with someone else about this recently. We just... I mean, this sort of... So, in our bubble, clearly this tech is amazing. Yeah, how much? Well, this is 10 grand now. Can you imagine? But I mean, like, the tech of AI is amazing. There's no doubt, right? We have no doubt about it. You know what my son sort of found out via listening to news on the internet? That AI makes the switch to more expensive. And I think this is sort of the fact that you have a child who doesn't really care about AI in any sort of meaningful capacity. But he learns that his friend is not going to buy a switch to now because it's more expensive because of AI. There's a sort of negative association with this thing now. I think that's great. I mean, if you look in the future, man versus machine, it's great that our youth now builds up this internal hate towards the things that we've done. They also had this, I don't know, was it Forbes or is it like Financial Times, but it had a bunch of surveys of age groups by how much they like AI. It's actually the older ones that like it and the younger generation. I think that makes intrinsic sense because if you think about it, the older generation... white collar workers they love this shit because now they can can let the ai build all their useless slide decks and chop materials that makes total sense to me and i i actually talking to people in our neighborhood that that's what they're telling me basically it helps them be lazy at the job which is great for them that makes total sense for me i just think it's interesting that this Like we are obviously a bubble, but our little bubble sort of destroys the world economy right now. Not in a way that's like, oh, AI makes your jobs useless. But like we're building out so much CapEx for data centers that all of a sudden like your computer costs more, educational computers cost more, consoles cost more, the electricity costs more. Then the stock market crashes and your pension's gone. So I mean, this is weird. I may have bought a Switch 2 as a response to increasing prices. And when I was looking at getting an SD card to install on it, they are also more expensive. So here's a question for you. And I guess, you know, now we're turning into predictions because we've been talking for a while, but is this kind of the beginning or is this the peak or is this sort of on the way up for, you know, just compute as a commodity, right? We've talked about sovereign AI. And I'm going to bring it back to what Mario brought up, which is, what does it matter that you have an open-weight model if you actually don't, you can't afford the hardware to run it? Doesn't matter, right? Does that trend continue? Because the value of what now people are realizing with computers is going so high up that it is just going to, like, never mind sovereign GPUs. Frankly, people should be, you know, should I be hoarding a bunch of AMD Ryzen's? in my computer, you know, in my basement because I don't know where that's going to be or is that just contributing to the problem? I mean, hoarding definitely contributes to the forward, right? The moment I said the word hoarding, I think that that answered the question. Big product out of the market prices increase, so that makes sense. But I'm convinced that the wet dream of all of these people at the moment is that we go back to the 70s where you had just a terminal and compute was basically not something you owe. but something you rent. And I have a feeling it's gonna go that way. The only problem with that is that everybody's addicted to those little rectangles we have in our pants, and they are also affected by the increase in DRAM prices, for example. And that's not gonna play well. Like, you're not gonna be able to take away phones from people. But I have a feeling that everything else is gonna be... That's a great insight. I wanna say they won't because the value... derived from this compute depends on the fundamental availability of the rectangle and this machine being distributed in people's homes. So there's some fundamental... But here's the thing, like, this is a production machine, right? I can do stuff with this. My phone, anytime I have to do any actual work on it, I fucking hate it. But I think what we're gonna get, and this is now conspiracy tinfoil hat territory, but my guess is the wishes of the grandmasters, the wishes... You're only going to use your phone. We're going to make your phone real fucking cheap. But you're going to sign up for our AI services that can then do your work for you by just giving them prompts on your phone. And you will never need this again. We're going to turn this thing that you can buy for a one-time payment. We're going to turn it into a subscription service because you now need a computer in the cloud where your agent can run and do all your work. And I think that's where we're going to go. Thin clients feel like they're coming back. I want to give one quick example, which is X. the everything app introduced, they've had a live streaming component for a little while, but the way that you did that is you ran sovereign software like OBS and you streamed that to their platform. Now this is actually being the actual sort of like recording capabilities are now part of that platform and you're going to go on the web on a thin client and you're going to use it that way. Yeah, I'm not too worried about that. So I had a conversation with Yudan Gazit, who works at GitHub Next, I think he's head of GitHub Next, and he brought up an interesting thought. This was maybe a little bit as an anti-piece as to the whole thing. So he was reminiscent on a time when computer game developers basically paid for IncrediBuild. So IncrediBuild is a thing I also used, which is... it took us too long to compile Unreal Engine. So every machine in our studio had IncrediBuild installed and then when someone wanted to compile it distributed the build across everybody else. And so if you do the math it actually turns out that in some office there are enough probably sufficiently large Mac chip is sitting around that you could sort of cluster up and actually have a local weight model do pretty decently. if you get it to run decently well. And so in terms of like random compute sitting around in different places, take all the Tesla cars that are standing around full packed with GPUs for self-driving, there's actually enough deployed infrastructure outside of data centers that you could actually run a bunch of models. And eventually that like, even in your version of this, when the computer goes away, that's a problem. But in a world where there's still going to be a bunch of fewer computers. there will be a surplus of deployed GPUs and maybe eventually someone's going to build a distributed inference engine that will harness all of this stuff. Like the limits of physics will throw a wrench in the other thing. I know. I think like maybe we're not all going to plug them in, but there's not that much data. Like the anti-resilist thing where he distributed DeepSeq flash, I think. All with lightning connectors. No, no, no. Because the actual on the way he probed it the actual data that he had exchanged between the sort of the checkpoints as he did was megabytes not gigabytes Interesting. Um, I mean more power to that kind of effort I I mean this there still it can still see Well, that's not completely making us dependent on this. I haven't thinking this during this conversation that like cyberpunk 2077 where we're all just gonna start killing each other for like the you know, the newest You know hardware upgrade or it's So there's like on this sort of thing is like there's like everybody like prices go up because like obviously There's a demand and then hopefully production goes up and then there's a glut and then the prices collapse So presumably is going to happen here, too, right? It doesn't seem like we're actually going to need as much stuff As we're probably in the process of building out there could be a correction. Okay, so we've just been kind of rambling about random topics and maybe coming to some of our favorite things like tokens are expensive um we have one question here from chewy distraction q how to handle fomo do you experience fomo no not anymore i got my eye psychosis out of my way last year while being fun employed where everybody else was ignoring agents I think the easiest way is to just look back at the last six months or so and ask yourself what really changed in your day-to-day and then ask yourself if it makes sense to be kind of like this and getting all news every minute into your system and whether it has any impact on your day-to-day if you do this and my prediction would be it does not so I think checking in with the state of things every month is probably enough That's all well and good for you Mario because you're literally at the at the You know, you are the creator of FOMO. You are leaving a drain. I'm sorry. I'm just the creator of flask. He created the phone now I think I mean ultimately it's still a while loop with an LLM request and then the tool execution nothing changed about that Yeah, I don't know how to handle FOMO. I think it is the problem I think that to some degree is like it's a question like how do you deal with your own emotions? I think the suggestion is the correct one which is I also have given wishes just don't read the news all the time but if something is interesting it's going to be there in two weeks and then if it's two weeks late it's not going to be a massive problem um but despite me saying that how how good am i actually managing my own emotions is a completely different question right um i think when i went to ai engineer like on the first day i was like i i felt like i don't know what i wrote i think i wrote something into this question like oh yeah I don't want to read the history. But it was like, because you're showing up in a place where everybody sort of lives the future and it's like crazy disorienting. And then like a day and a half and I was like, okay, it's just pretending to live in the future. But there is, I think that actually, I think that's the guy from open codes, not the one from human layer. I actually say this recently, which is like, we all live in a world of like... uncertainty and a good way to feel better about it is like you're pretending to know the future and telling it to everybody else so that you build up yourself there. Maybe that's one way to deal with FOMO but also just recognizing that probably a lot of people are doing that. It's like nobody really knows what's going on. Some hunches maybe. But I think we'll find out at the end of the year. I think by that. So you think by the end of the year we're right? That's my prediction by the way for the end of the thing. My prediction for the end of the year is we will have found out what the end state of all of this carnage is. We should predict when the bubble bursts. I think that would be good. As soon as Entropico open AI IPO. I think the IPO is basically like, the good thing is like, I think it's a relatively safe point to say like a little bit after the IPO is going to be when this all plays out. But on the other hand, we have SpaceX who still haven't gone down to the actual dollar amounts they should go down to. But they only issued 5% of the stocks. Ben, when is the bubble? Not the first. Tuesday. Good. I have no plans. Well, should I start down here? Yeah. All right. Thanks for being here. Thanks for having me. Okay. Bye.
⚙️ Pipeline jobs
| Stage | Status | Att. | Updated | Error |
|---|---|---|---|---|
| download | done | 1/3 | 2026-07-20 11:49:05 | |
| transcribe | done | 1/3 | 2026-07-20 11:49:59 | |
| summarize | done | 1/3 | 2026-07-20 11:50:43 | |
| embed | done | 1/3 | 2026-07-20 11:50:44 |
📄 Описание YouTube
Показать
00:00 Cold Open 00:50 Live-ish from Vienna: Meet Mario Zechner 06:43 Vibe-Checking Fable 21:35 RL 101 & How Harnesses Normalize Model Jank 36:56 When Models Get Worse, Will You Know? 47:21 GLM-5.2 and Open Weights 51:18 Agent Loops, Trust, and ROI 01:09:00 AI Taxed Your Nintendo Switch 01:17:23 Handling AI FOMO & Closing Pi: https://pi.dev hunk (terminal diffs): https://hunk.dev Sideshow: https://sideshow.sh The original Ralph loop — https://ghuntley.com/ralph/ Mario Zechner on X: https://x.com/badlogicgames Ben Vinegar on X: https://x.com/bentlegen Armin Ronacher on X: https://x.com/mitsuhiko Modem: https://modem.dev/ Earendil: https://earendil.com/