The Saga Pattern: Implementing State Persistence
The Dev World - by Sergio Lema · 2026-04-23 · 12м 8с · 277 просмотров · YouTube ↗
Топики: durable-execution
🎧 Аудио
📝 Summary
model=deepseek-v4-flash · prompt=summary-v7 · 3 505→2 460 tokens · 2026-07-20 15:09:58
🎯 Главная суть
Saga Pattern — способ управления долгоживущими распределёнными транзакциями без ACID-изоляции. Вместо единой транзакции используется последовательность локальных действий с возможностью отката (компенсации) для тех шагов, которые ещё не дошли до «точки невозврата». После точки невозврата (например, списание денег) все последующие шаги должны быть повторяемыми до успеха. Это единственный способ не просыпаться в 2 часа ночи и не править три базы данных вручную.
Проблема распределённых транзакций: пример с кофемашиной
В пятницу в 16:00 вы развернули микросервисный флоу заказа: Service A (order) → Service B (payment) → Service C (inventory). REST-вызовы выглядят чисто, сервисы декомпозированы. Вы уходите домой с чувством гордости. В 2 часа ночи — инцидент: клиент купил кофемашину за $2000. Service A создал запись заказа, Service B успешно списал деньги, а Service C (инвентарь) упал из-за ночного бекапа. Инвентарь не обновился. В результате у клиента нет машины, у вас — $2000 и база данных в Service A, которая считает, что всё в порядке. Следующие 4 часа вы вручную правите SQL в трёх продакшен-базах, сверяясь с логами. Это и есть «кошмар распределённых транзакций»: убив монолит, мы вместе с ним сожгли ACID-транзакции, получив тысячу маленьких невидимых пожаров.
Три ложных убеждения, которые приводят к хаосу
Разработчики часто обманывают себя, чтобы не внедрять Saga:
- «Сеть надёжна» — на самом деле это «трубы, скреплённые изолентой и молитвами».
- «Я использую двухфазный коммит (2PC)» — это не 1998 год, ваша NoSQL база его не поддерживает. Даже если бы поддерживала, латентность похоронила бы производительность.
- «Просто добавлю
@Transactional» — аннотация работает только на одной базе данных, а у вас три AWS-региона и легаси-сервер в чьём-то подвале. - «Исправлю неконсистентность крон-джобом в понедельник» — аналог «начну диету завтра». Крон падает, и вы получаете две проблемы.
Два типа разработчиков, усугубляющих ситуацию
- Оптимист — верит в доброту людей и стабильность интернета. Пишет код в расчёте, что всё сработает с первой попытки. Например, если вызов инвентаря выбрасывает 500, деньги уже списаны, заказ сохранён, клиент звонит в поддержку. Такой код — не сервис, а обязательство.
- Архитектор-абстрактор — на прошлой неделе узнал о паттернах и решил применить их все. Добавляет 5 слоёв абстракции, кастомную шину событий, которую понимает только он, и столько интерфейсов, что найти логику — квест. Итог: построенный до решения бизнес-задачи generic Saga-фреймворк, который невозможно оттестировать и отладить. При первом сетевом таймауте state machine зависает в состоянии
pending_compensation_final_retry_v2навсегда.
Как работает Saga: State Machine + Compensating Transactions
Saga — это длительный разговор между сервисами, который нужно записывать. Ключевые элементы:
- Compensatable (откатываемые) действия — то, что можно отменить, например, резервирование инвентаря.
- Pivot (точка невозврата) — обычно списание денег. Если пэймент прошёл, мы обязаны завершить всю сагу.
- Retriable (повторяемые) действия после pivot — например, отправка товара. Они не могут упасть навсегда: если сбой, мы повторяем до успеха.
Логика: если pivot упал, отменяем всё до него. Если упало что-то после pivot — не отменяем, а ретраим.
Saga теряет изоляцию (ACD, не ACID). Пока сага выполняется, другой процесс может увидеть «грязные» данные. Для защиты используется семантический блок (semantic lock): резервируя товар, помечаем его статусом pending_saga, чтобы другие транзакции не считали его доступным.
Оркестрация против хореографии
В комментариях часто спрашивают: почему не event-driven choreography (каждый сервис сам знает, что делать дальше)? Причина — наблюдаемость. В хореографии, когда цепочка из шести сервисов, у вас нет центрального места, чтобы увидеть состояние. Это игра в «испорченный телефон» с $1000. Для 90% бизнес-кейсов центральный оркестратор (state machine) лучше: вы можете открыть дашборд и точно сказать, на каком шаге умер заказ.
Практические советы по внедрению Saga
- Найдите pivot — обычно это действие, которое сложнее всего отменить (деньги).
- Определите компенсации — если действие нельзя отменить, оно должно быть после pivot и обязано быть retriable.
- Не пишите свой фреймворк — используйте AWS Step Functions, Temporal или Spring State Machine. Они решают проблему «а что, если сам оркестратор упадёт».
- Обеспечьте идемпотентность — каждый сервис должен уметь принять одну и ту же команду 5 раз, но выполнить её один раз. Если эндпоинт возврата денег не идемпотентен, будет очень плохо.
Распределённые системы сложны не потому, что они «монолит с сетевыми вызовами», а потому, что это хаотичная среда, где всё может и будет ломаться, пока вы спите. Saga не делает систему идеальной — она делает отказы управляемыми, давая план на случай, когда мир рушится. Удалите вложенные try-catch и спроектируйте путь отказа.
📜 Transcript
en · 1 254 слов · 21 сегментов · clean
Показать текст транскрипта
Imagine it's Friday 4 p.m. You've just deployed the new order to delivery microservice flow. You used a nice clean set of REST calls. Service A calls service B, B calls C. It's elegant. It's decoupled. You go home, open a beer and feel like a cloud native goat. Then at 2 a.m. the page goes off. A customer bought a 2000 express machine. Service A. the orders, created the record, service B the payments, successfully charged the credit card, but service C the inventory was having a moment because someone decided to run a backup during peak hours. The inventory update failed. The result, the customer has no expiration machine, you have their 2000 years and your database in service A thinks everything is fine, you spend the next 4 hours manually running SQL updates across the 3 different production databases, squinting at logs like they are the matrix, trying to figure out who to refund and what to delete. This is the distributed transaction nightmare. We killed the monolith and in our excitement we burned AC transaction in the same grave. We traded one big manageable fire for a thousand tiny invisible ones. Welcome to the saga pattern. It's the industry's way of admitting that distributed systems are a mess and the eventual consistency is just a fancy term for it will be great eventually, hopefully, maybe. My name is Sergio and I spent way too much time thinking about why we built systems that are destined to fail. This is the second video in our microservices architecture patterns playlist. If you missed the first one, you're already behind on your technical depth. If you ever looked at a distributed stack trace and felt your soul will leave your body, hit subscribe, we're digging through the architectural trenches together. Before we look at the code, let's address the psychological denial we live in as developers. We tell ourselves these lies to avoid the complexity of sagas. The network is reliable. It's not. It's a series of tubes held together with duct tape and prayers. I will just use a two-phase comment. No, you won't. This isn't 1998 and your more than no SQL database doesn't support it. Even if it did, the latency wouldn't turn your high performance API into a carrier pigeon service. I will just add a transactional annotation. That works great. On one database. Too bad, your database is scattered across three AWS regions and the legacy servers in someone's basement. I will fix data inconsistencies with a cron job on Monday. This is the developer equivalent of I will start my diet tomorrow. Monday cons. The Chrome job fails and now we have two problems. This developer believes in the goodness of humanity and the stability of the internet. They write code that assumes everything works on the first try. The quality check. If inventory client reserves throws a 500, the payment is already gone. The order record is saved. The customer is calling support. This code isn't a service. It's a liability. This developer discovered design patterns last week and decided to use all of them. They've added 5 layers of abstraction, a custom event bus that only they understand and so many interfaces that finding the actual logic is like a game of Waze World. The reality check, they've built a generic saga framework before they even solved the business problem. It's hard to test, impossible to debug and the first time a network timeout happens, the state machine gets stuck in a pending compensation final retry v2 state forever. We're going to build a state machine based saga. We aren't going to use magic, we're going to use compensating transactions and a clear pivot point. First we need to track where we are. A Saga is just a long running conversation. If you don't write down what was said, you're going to forget. In Saga, we have Compensate table, Pivot and Retriable transactions. Compensate table can be undone, like unreserved inventory. Pivot the point of no return usually the payment and the payment is settled we must finish Retriable transactions after the pivot that cannot fail permanently like shipping if they fail we retry until they succeed If the pivot the payment fails we have to undo everything before it if something fails after the pivot we don't undo we retry Sagas are ACD not 80 we lost the isolation If another process looks at our inventory while the saga is running they might see dirty data. To fix this we use a semantic lock. We don't just reserve the item, we mark it as pending saga. Why do we struggle with this? Because of the regime driven development. Implementing a proper saga with AWS step functions or SPRIC state machine feels heavy. It's not agile. We choose the easy way, the nested REST goals. We do this because we want to ship features today and let the maintenance version of us deal with the data corruption six months from now. We also suffer from the cargo called programming. We see Google or Netflix using microservices. So we do it too. But Google has a thousand engineers to build their infrastructure for SEGAs. You have a Jira ticket. And the deadline for Tuesday. When you move to microservices without a SEGA strategy, you aren't building a distributed system. You're building a distributed mode. I can already hear the comments section. Well, actually Sergio, why not just use an event-driven choreography saga instead of an orchestrator? It's more decoupled. Sure, if you enjoyed not knowing what the hell is happening in your system, in choreography, each service knows what to do next. It sounds great until you have to debug flow that spans six services and realize you have no central place to see the state. You're essentially playing telephone with 100 euro bills. For 90% of business cases, a central orchestrator, like a state machine, is better because it's observable. You can actually point at a dashboard and see where the order died. But Sergio, what about the overhead of persisting the saga state? What's more expensive? A few milliseconds of DB write for the state machine or your lead developer spending 10 hours on a Saturday morning manually reconciling bank statements with database records? Do the math. If you are going to implement Cygas tomorrow, do this. Identify your pivot point. Usually this is the action that is hardest to undo. The money. Define your compensations. If you can't undo an action, it must happen after the pivot and it must be retryable. Don't rule your own framework. Use AWS step functions, temporal or spring state machine. These tools handle the what if the orchestration itself crashes problem, which trust me, you don't want to solve yourself. Embrace idempotency. Every service in a SAIGA must be able to receive the same command five times and only execute it once. If your refund endpoint isn't idempotent, you're going to have a very bad time. Distributed systems are hard because we pretend they are just monolith with network calls. They aren't. They are chaotic environments where everything can go wrong. will go wrong, usually while you're trying to sleep. The saga pattern isn't about making things perfect, it's about making the failures manageable. It's about having a plan for when the world burns down. Now go delete the nested try catch block and actually design your failure path. And see you in the next video. Bye!
⚙️ Pipeline jobs
| Stage | Status | Att. | Updated | Error |
|---|---|---|---|---|
| download | done | 1/3 | 2026-07-20 15:09:25 | |
| transcribe | done | 1/3 | 2026-07-20 15:09:34 | |
| summarize | done | 1/3 | 2026-07-20 15:09:58 | |
| embed | done | 1/3 | 2026-07-20 15:09:59 |
📄 Описание YouTube
Показать
Stop manually fixing database inconsistencies at 2 AM. Learn how to implement the Saga Pattern to manage state across microservices, handle isolation challenges, and design reliable compensating transactions. This deep dive into Microservices explores the Saga Pattern as an alternative to Two-Phase Commits (2PC). We focus on State Persistence, Idempotency, and Eventual Consistency using Spring Boot and Java. By addressing Semantic Locking and Compensating Transactions, we solve the Lack of Isolation inherent in Distributed Systems. This video belongs to the Advanced Microservices Patterns playlist: https://www.youtube.com/watch?v=PzzH-1y5yzE&list=PLab_if3UBk99JzCygewVzzk1hhrHMda-A Chapters: 0:00 The 2 AM Ghost in the Machine 1:27 Signature Intro 1:52 The Lies We Tell Ourselves 2:47 The "Junior vs. Senior" Evolution 6:00 The "Real World" Live-Coding Breakdown 9:09 The Psychology of the Error 9:53 Addressing the "Well, Actually" Crowd 10:52 Actionable Solution 11:36 The "Go Code" Outro Github repository: https://github.com/serlesen/microservices-architectures/tree/video_2 My NEW eBook: https://sergiolema.dev/best-practices-to-create-a-backend-with-spring-boot-3/ Blog: https://bit.ly/47ornJL LinkedIn: https://bit.ly/41Nn61q