# Disaster recovery

## Operating model

Regional recovery is active-passive for ordered event processing, replay administration, credential issuance, mutable partner configuration, provider execution and reconciliation ingestion. Stateless documentation, schema retrieval, authentication and health monitoring may be distributed.

## Recovery objectives

- Tier 0 partner ingress and canonical event acceptance: RTO 60 minutes, RPO 5 minutes.
- Tier 1 provider delivery and webhook egress: RTO 120 minutes, RPO 15 minutes.
- Tier 1 reconciliation ingestion: RTO 240 minutes, RPO 60 minutes.
- Immutable assurance evidence: RTO 24 hours, RPO 15 minutes.

These objectives remain provisional until an approved exercise demonstrates them.

## Preconditions

1. Multi-AZ operation is healthy in the primary region.
2. Aurora backups and point-in-time recovery are current.
3. Evidence objects replicate to an Object-Lock-enabled recovery bucket.
4. Images are available by immutable digest in the recovery region.
5. Secrets and certificate trust material are replicated through approved lifecycle processes.
6. MSK topic configuration, retention and ACL definitions are reproducible.
7. Provider endpoints permit the recovery-region egress identities.

## Controlled recovery sequence

1. Declare the incident and freeze partner configuration, replay and certificate issuance.
2. Fence primary provider-call workers and record the fencing token.
3. Confirm the last signed audit checkpoint and database recovery point.
4. Restore Aurora into isolated recovery subnets and run schema integrity checks.
5. Create MSK topics and restore only replay-safe event classes.
6. Start tasks with outbound provider execution disabled.
7. Verify duplicate suppression, certificate registry, OAuth binding and canonical schema hashes.
8. Transfer event-consumer ownership using a new fencing token.
9. Enable ingress, then non-consequential egress, then approved provider capabilities.
10. Reconcile every ambiguous provider instruction before normal operation.
11. Produce and immutably retain the recovery evidence pack.

## Mandatory exercises

- quarterly backup restoration into an isolated account or VPC;
- six-monthly task and Availability Zone failure;
- annual regional recovery;
- annual certificate-authority and secret-rotation recovery;
- replay-after-failover and reconciliation-after-recovery.

Production failover is never inferred solely from a health check. An authorised incident commander must approve ownership transfer for consequential processing.
