The Compensating Transaction Pattern
Pastel Sketchbook · 2026-04-30 · 14м 5с · 19 просмотров · YouTube ↗
Топики: durable-execution
🎧 Аудио
📝 Summary
model=deepseek-v4-flash · prompt=summary-v7 · 4 352→2 276 tokens · 2026-07-20 15:08:02
🎯 Главная суть
Компенсирующая транзакция (compensating transaction) — это шаблон для обработки сбоев в распределённых системах с конечной согласованностью. Вместо традиционного отката к исходному состоянию (который в распределённой среде опасен) система на каждом успешном шаге сохраняет метаданные для обратного действия и при ошибке выполняет бизнес-логику, возвращающую систему в логически согласованное состояние, учитывающее параллельные изменения других экземпляров.
Проблема традиционного rollback в распределённых системах
Простой откат на уровне БД в распределённом окружении не работает по трём причинам. Во-первых, параллельные изменения: пока шаг 3 терпит неудачу, другие экземпляры уже могли легально изменить те же записи — жёсткий откат уничтожит эти данные. Во-вторых, сложность отмены: логика, обращающая бизнес-операцию, часто не менее сложна, чем сама операция. В-третьих, внешнее состояние: если действие затронуло внешний сервис, локальный откат БД его не отменит — нужно явно вызвать этот сервис с компенсирующим вызовом. В результате фокус смещается с восстановления исходного состояния («state restoration») на управляемый откат с явной бизнес-логикой («logic-driven reversal»).
Механика компенсирующей транзакции
Компенсирующая транзакция — это workflow, который «отматывает» выполненные шаги неудавшейся операции, применяя бизнес-правила для каждого обратного действия. Ключевые элементы: запись подробной информации о каждом шаге по мере его выполнения (чтобы потом знать, как именно откатывать) и учёт параллельной работы других экземпляров. Пример с бронированием: шаг 1 (билет) и шаг 2 (ещё один билет) успешны, шаг 3 (отель) падает. Система запускает компенсацию: отменяет шаг 2, затем шаг 1. При этом:
- Компенсация не всегда строго последовательна — некоторые шаги можно выполнять параллельно.
- Логика домен-специфична: отмена авиабилета — это не удаление строки из БД, а вызов бизнес-процесса с частичным возвратом средств.
- Вместо немедленной отмены всего можно попытаться устранить проблему (например, предложить альтернативный отель), сохранив уже сделанный прогресс.
Иерархия обработки сбоев: retry → компенсация → человек
Устойчивость к сбоям строится по пирамиде. Основание — автоматические повторные попытки для транзиентных ошибок (сетевые глитчи). Если повторные попытки исчерпаны или ошибка нетранзиентна (например, товар закончился), включается следующий уровень — компенсирующая транзакция: прямой прогресс останавливается, запускается workflow для безопасной отмены или корректировки. На вершине — человеческое вмешательство для необратимых внешних эффектов или решений высокой ценности. Система приостанавливает workflow и генерирует детальные оповещения.
Пример реализации на Azure: Container Apps, Service Bus, Cosmos DB
В архитектуре на Azure orchestrator размещается в Azure Container Apps и управляет последовательностью шагов. Клиент отправляет запрос, orchestrator через Azure Service Bus связывается с микросервисами (Service A, Service B). После успешного выполнения каждого шага orchestrator логирует состояние выполнения и метаданные для компенсации в Azure Cosmos DB. Если на любом этапе возникает ошибка, orchestrator извлекает сохранённые метаданные и запускает точные обратные действия. Service Bus дополнительно отправляет повторно упавшие сообщения в dead-letter queue.
Практические правила внедрения
- Идемпотентность команд: компенсирующие действия сами могут сбоить, поэтому их нужно делать повторяемыми без побочных эффектов.
- Точки невозврата: необратимые шаги (юридически обязывающие действия) должны выполняться только после прохождения всех критичных валидаций.
- Краткосрочные блокировки: временно блокировать необходимые ресурсы заранее, финализировать операции до истечения блокировки — это повышает успешность многошаговых транзакций.
- Сквозная наблюдаемость: структурированная телеметрия с correlation ID обязательна, чтобы отслеживать как прямой поток, так и его отмену.
Сравнение: традиционный rollback vs компенсирующая транзакция
| Аспект | Традиционный rollback | Компенсирующая транзакция |
|---|---|---|
| Механизм | Атомарный откат на уровне хранилища | Пошаговая бизнес-логика отмены |
| Блокировки | Блокирует записи или всю БД до завершения | Не использует долгие блокировки, разрешает параллельные изменения |
| Обработка ошибок | Автоматически через БД | Приложение-специфичная логика, требуется orchestrator |
| Конечное состояние | Точное исходное состояние | Новое скомпенсированное состояние с учётом промежуточной работы |
Когда применять, а когда не стоит
Паттерн подходит, когда workflow затрагивает несколько сервисов или хранилищ и традиционные ACID-транзакции невозможны. Особенно эффективен, если восстановление требует сложной доменной логики (отмена брони отеля, частичный возврат). Идеален для систем с конечной согласованностью, которые не могут полагаться на атомарные транзакции между компонентами.
Избегать паттерна стоит в трёх случаях: если ошибки можно решить простыми повторными попытками (большинство сетевых глитчей); если система не может терпеть даже временную несогласованность; если компенсирующие действия не могут надёжно восстановить валидное состояние — тогда нужно использовать строгую согласованность.
Связь с другими паттернами и Well-Architected Framework
Компенсирующая транзакция — ядро самоисцеляющейся архитектуры. Её окружают:
- Retry pattern — первая линия обороны, сокращает частоту вызова компенсации.
- Saga pattern — управляет согласованностью данных между микросервисами, используя компенсацию как основной механизм восстановления.
- Transactional outbox — гарантирует атомарную запись команд компенсации вместе с бизнес-данными в той же транзакции, исключая потерю сообщений.
- Pipes and filters — разбивает сложную задачу на независимые шаги, каждый из которых можно откатить индивидуально.
Паттерн прямо отвечает за надёжность (reliability pillar) Well-Architected Framework: предотвращает каскадные сбои (критические пути защищены компенсациями) и обеспечивает корректное восстановление без «зависших» данных. Соответствует рекомендациям RE02 (защита критических потоков) и RE09 (катастрофоустойчивость).
📜 Transcript
en · 2 029 слов · 33 сегментов · clean
Показать текст транскрипта
Welcome everyone. Today we are exploring the compensating transaction pattern. In the realm of eventually consistent distributed systems, handling failures requires more than just a simple rollback. This pattern provides a structured approach for orchestrating graceful failure recovery, ensuring that even when a step in a long-running process fails, the system can systematically undo previous actions to maintain a consistent state. Let's delve into how this strategy enables resilience and reliability in complex modern architectures. Cloud scale necessitates a fundamental shift from strong consistency to eventual consistency. Historically, legacy applications have relied on strong transactional consistency, which maintains a unified state by locking databases. While effective for smaller systems, this approach creates significant resource contention and performance bottlenecks as we attempt to scale. To maximize performance and eliminate data contention across disparate geographic locations, modern cloud applications adopt eventual consistency. This model breaks business operations into separate, manageable steps. Although the system state may temporarily disagree across different loads, it is designed to converge toward a consistent state once all operations are complete. The distributed dilemma highlights why traditional rollbacks often fail in modern distributed workflows. Simply put, restoring a system to its exact original state is not just difficult, it is dangerous. The first major challenge involves concurrent changes. In a distributed environment, multiple application instances are constantly modifying data. While a specific operation is running and potentially failing at step 3, other instances may have already committed valid changes to the same records. As illustrated by the brick wall in the diagram, a hard database rollback would blindly overwrite all that valid work, leading to data loss and inconsistency. Furthermore, we must navigate complex reversals. Undoing a multi-step business process is rarely as simple as reverting a single database row. The logic required to unwind an operation is often just as complex as the operation itself. Finally, we have to account for external state. In a service-oriented architecture, different services often hold their own internal state. To undo an action that involved an external service, we cannot rely on a local database rollback. Instead, we must explicitly invoke that service again to perform a compensating action that reverses the initial effect. This reality shifts our focus from simple state restoration to the more complex requirement of explicit, logic-driven reversals. The compensating transaction pattern is a workflow that intelligently rewinds through the completed steps of a failed operation. By applying business-specific rules to reverse effects rather than forcing a blind system restore, it ensures a more precise and context-aware recovery. Key to this process is recording detailed information about each step as it runs, which allows the system to know exactly how to reverse those actions later. Furthermore, this approach accounts for concurrent work by other active application instances, maintaining system-wide integrity throughout the failure resolution process. This diagram illustrates the concept of workflow topology, specifically focusing on the balance between forward progress and intelligent reversal. In a distributed transaction model, we move through a series of forward steps. In this example, steps one and two, booking flights, succeed, but step three, booking the hotel, results in a failure. When a failure occurs, the system initiates compensating actions. These are not simple rollbacks, but specific logic designed to return the system to a consistent state. As indicated by the transition from the failed hotel booking, the system moves to undo step two and then step one. However, effective reversal logic is more nuanced than a simple mirror image of the forward path. First, compensation is not always strictly sequential. Some undo steps can be executed in parallel to increase efficiency. Second, the logic is domain-specific. Canceling a flight rarely involves a simple database deletion. Instead, it triggers business-defined processes such as partial refunds. Finally, we prioritize selective reversal. Rather than immediately canceling the entire transaction, an intelligent workflow might first attempt to resolve the failure, for example by offering an alternative hotel, to preserve the progress already made. The escalation of failure recovery follows a structured approach designed to maintain system integrity while minimizing manual workload. At the Foundation, we handle transient failures through automated retry logic. These involve brief interruptions, such as network blips, which are managed by retrying identipotent commands to preserve forward progress. When failures are non-transient, meaning retries have been exhausted or specific business rules have failed, such as an out-of-stock condition, we move to the next level of escalation. Here, we utilize compensating transactions. Forward progress is halted, and a compensation workflow is triggered to safely undo or adjust previous actions. Finally, at the top of the pyramid, we address high-impact or ambiguous failures that necessitate human intervention. For actions with irreversible external side effects or those requiring high-value decision-making, the system pauses the workflow and generates detailed alerts for manual review. This ensures that the most complex and critical issues receive the direct human oversight they require. This architectural blueprint illustrates a robust approach for orchestrating compensation within Azure. When dealing with long-running workflows, managing state and ensuring reliability across distributed services is paramount. In this implementation, an orchestrator hosted within Azure Container Apps serves as the central brain for the process. The workflow begins when a client initiates a request. The orchestrator then manages the sequence of operations by communicating through Azure Service Bus to various downstream components, represented here by Service A and Service B. As each forward step is successfully completed, the orchestrator logs critical execution state and compensation metadata into Azure Prosmos DB. This database serves as a durable record of both progress and the necessary steps to undo an operation if a failure occurs. If an error is encountered at any point in the chain, the orchestrator retrieves the stored metadata to trigger the precise reversal actions needed to maintain system consistency. This slide outlines the critical component roles within the compensation ecosystem. Starting on the left, Azure Container Apps serves as the orchestrator. It is responsible for coordinating each step of the workflow, applying retry logic to handle transient faults, and determining if an alternative path is viable before triggering a full compensation process. In the center, Azure Service Bus manages the messaging layer. It routes both forward and compensation commands between microservices and automatically moves repeatedly failing messages to a dead letter queue to maintain system integrity. To the right, Azure Cosmos DB handles state and audit functions. It records the current execution state and stores the specific compensating actions required to resume, correlate, and audit the workflow. Together, these components provide a robust framework for managing complex distributed transactions and error recovery. Implementing effective compensation strategies in distributed architectures requires robust guardrails to ensure system consistency and reliability. First, we must design item potent commands. Because compensating transactions themselves are susceptible to failure, it is essential that undo steps can be safely repeated during retries. By ensuring item potency, we prevent unintended side effects when a system attempts to recover from a partial failure. Second, it is crucial to clearly define points of no return. We must identify any irreversible steps, such as external, legally binding actions, and ensure these occur only after all critical validations have successfully completed. This minimizes the risk of reaching a state that cannot be reverted if an error occurs later in the process. Third, we should utilize short-term locks. By placing timed locks on required resources in advance, we ensure that all necessary components are available before performing the actual work. Finalizing operations before these locks expire significantly improves the success rates of complex multi-step transactions. Finally, we must maintain end-to-end observability. Since compensations run after original operations have already committed, we rely on structured telemetry and correlation IDs to maintain a clear audit trail. This visibility is vital for monitoring both the original transaction flow and its subsequent reversal, ensuring we can verify the final state of the system. When managing data recovery, it is crucial to understand the functional differences between strong and eventual consistency models, specifically through the lens of traditional rollback versus compensating transactions. Starting with the primary mechanism, a traditional rollback relies on atomic data reversion, essentially undoing changes at the storage layer. In contrast, A compensating transaction uses step-based business logic reversal, which is better suited for distributed environments where a single database rollback is not possible. In terms of resource locking, traditional rollbacks often lock affected records or even the entire database until the process completes, which can lead to high contention in busy systems. Compensating transactions avoid long-term locks, embracing highly concurrent modifications and allowing the system to remain responsive even during recovery. Failure handling also differs significantly. In a traditional model, the system automatically handles recovery via the database engine itself. With compensating transactions, failure handling is application-specific and requires custom orchestrator logic to navigate the various states of a distributed process. Finally, we look at the end state. A traditional rollback returns the system to its exact original state, as if the failed operation never occurred. A compensating transaction, however, reaches a new compensated state that accounts for any intermediate work performed by concurrent processes, ensuring the system remains logically consistent even if it does not return to its original bit-for-bit configuration. Now let's look at when to deploy this pattern, and equally importantly, when it is best to avoid it. You should consider this pattern when your workflows span multiple services or distinct data stores where traditional ACID transactions are not feasible. It is particularly effective when failure recovery requires complex, domain-specific logic, such as cancelling a hotel reservation or issuing a partial refund, rather than a simple database rollback. Furthermore, it is ideal for systems designed around an eventual consistency model that cannot rely on atomic transactions across distributed components. On the flip side, you should avoid this pattern if operations can be safely handled with simple retries, which is often the case when the vast majority of failures are strictly transient network blips. You should also steer clear if your system requirements dictate that you strictly cannot tolerate even temporary inconsistency. Finally, if a compensating action cannot reliably restore the system to a valid state, this pattern is not the right choice. In those instances, you should utilize strong consistency models instead. This architectural pattern is designed to align directly with the well-architected framework's reliability pillar. By implementing this approach, we significantly enhance the system's resiliency to malfunction. Specifically, it utilizes compensation actions to address failures within critical paths, which effectively prevents localized issues from escalating into cascading system crashes. Beyond mere prevention, the pattern ensures graceful recovery. Rather than leaving data in a corrupted or orphaned state following a failure, the mechanism guarantees that the system returns to a fully functioning state. This is achieved through automated processes, such as rolling back data, breaking resource locks, or executing native reverse behaviors. Ultimately, this pattern provides concrete support for WAF best practices, RE02, which focuses on protecting critical flows, and RE09, which governs robust disaster recovery strategies. The compensating ecosystem represents a strategic approach to building self-healing architecture. At its core is the compensating transaction, supported by a network of design patterns that ensures system resilience and data consistency. The retry pattern serves as our first line of defense. By automatically addressing transient faults, it significantly reduces the frequency with which we must resort to full compensation. When immediate recovery isn't possible, we turn to the saga pattern. This is essential for managing data consistency across microservices, using compensation as a primary recovery mechanism to maintain a consistent state when long-running distributed processes encounter errors. To guarantee the reliability of these recovery actions, we utilize the transactional outbox. This pattern ensures that compensating commands and state changes are recorded atomically within the same transaction as the business data, effectively eliminating the risk of message loss. Furthermore, the pipes and filters pattern allows us to break down complex tasks into reusable, independent steps. This modularity enables us to individually compensate specific filters if a failure occurs downstream, providing a high degree of control over the recovery process. Together, these components form a robust ecosystem that allows our systems to gracefully handle failures and maintain operational integrity.
⚙️ Pipeline jobs
| Stage | Status | Att. | Updated | Error |
|---|---|---|---|---|
| download | done | 2/3 | 2026-07-20 15:07:29 | |
| transcribe | done | 1/3 | 2026-07-20 15:07:40 | |
| summarize | done | 1/3 | 2026-07-20 15:08:02 | |
| embed | done | 1/3 | 2026-07-20 15:08:03 |
📄 Описание YouTube
Показать
Orchestrating graceful failure recovery in eventually consistent distributed systems. https://learn.microsoft.com/en-us/azure/architecture/patterns/compensating-transaction