AWS Deep Dives

AWS / Architecture Guide

AWS Backup Deep Dive

// How Vaults, Plans, Jobs, Copy Jobs and Restore Jobs are orchestrated

Concepts Architecture Backup Plans Jobs Flow Copy Jobs Restore Jobs Coordination Full Example Supported Services Security Lifecycle Monitoring

New to AWS Backup? Start Here

AWS Backup has a lot of moving parts. Here's the mental model that makes the rest of this page easy to follow:

AWS Backup

Think of it like a Bank Vault

A Vault is your secure safe-deposit box. Your backups (called recovery points) sit inside it. You can have multiple vaults — one for prod, one for DR (disaster recovery), one for compliance.

EventBridge Schedule

Think of Plans like a Calendar Subscription

A Backup Plan is like a recurring calendar event: "every night at 2 AM, back up everything tagged Backup=true and keep it for 30 days." You write the plan once; AWS does the work automatically.

RDS Snapshot

Recovery Points are Snapshots in Time

Each time a backup runs, it creates a Recovery Point — a frozen copy of your resource at that exact moment. You can restore from any of these points later. Think of them like iPhone backups: you can roll back to yesterday's or last week's.

ARN

ARN = Amazon Resource Name (a unique ID)

You'll see ARNs everywhere in AWS. They're just unique identifiers for any resource, like arn:aws:rds:us-east-1:123456789:db:my-database. When the docs say "Recovery Point ARN", they mean the unique ID of a specific backup snapshot.

AWS Region

Region = Physical AWS Data Center Location

AWS has data centers worldwide (us-east-1 = N. Virginia, eu-west-1 = Ireland, etc.). "Cross-region copy" means sending a backup to a different geography so that if an entire region fails, you still have your data elsewhere.

IAM

IAM Role = Permission Pass for AWS Backup

AWS Backup needs permission to access your databases, EC2 instances, etc. An IAM Role grants those permissions. AWS provides a default one called AWSBackupDefaultServiceRole that works for most cases — just use that to start.

The 30-second summary: You tag your AWS resources (databases, servers, file systems) with Backup=true. You create a Backup Plan that says "back these up nightly". AWS Backup then runs on schedule and stores each backup in a Vault. It can copy that backup to a second region for safety. When you need the data back, you restore a backup into a brand-new resource.

The 5 Building Blocks

AWS Backup is a fully managed service that centralizes and automates data protection across AWS services. Before diving into flows, understand each primitive.

AWS Backup vault
VAULT

Backup Vault

A container that holds recovery points (backups). Each vault has a KMS key, an access policy and an optional vault lock (WORM). You can have many vaults per account and Region.

Scheduled backup plan
PLAN

Backup Plan

A policy document made of rules. A rule sets when to back up (schedule), how long to keep the backup (lifecycle) and where to copy it (copy actions). A separate selection attaches the plan to resources by tag or ARN.

Backup job execution
JOB

Backup Job

One run against one resource. It takes a snapshot, or a continuous backup, and stores the recovery point in the target vault. A plan rule starts it, or you start it on demand.

Cross-region copy job
COPY JOB

Copy Job

Copies an existing recovery point into another vault. The target can be in the same Region, another Region, or another AWS account. This is how you build DR and compliance isolation.

Restore job
RESTORE JOB

Restore Job

Rebuilds a resource from a recovery point. You pass the target configuration as restore metadata, and AWS Backup provisions the new resource.

How Everything Connects

The diagram below shows the top-level relationships between AWS Backup components and the protected resources.

graph TD subgraph ACCOUNT["AWS Account (us-east-1)"] direction TB PLAN["Backup Plan ───────────── Schedule: cron(0 2 * * ? *) Retention: 30 days warm / 365 delete Copy rule to DR vault"] SEL["Resource Selection ───────────── Tag: Backup=true or specific ARNs"] subgraph RESOURCES["Protected Resources"] EC2["EC2 Instance"] RDS["RDS Database"] EFS["EFS File System"] DDB["DynamoDB Table"] end subgraph PRIMARY_VAULT["Primary Vault (us-east-1)"] RP1["Recovery Point 1 2024-01-15 02:00"] RP2["Recovery Point 2 2024-01-16 02:00"] RP3["Recovery Point 3 2024-01-17 02:00"] end end subgraph DR_ACCOUNT["DR Account / DR Region (eu-west-1)"] DR_VAULT["DR Vault Cross-region copy"] RP_DR["Recovery Points copied from primary"] end PLAN --> SEL SEL --> RESOURCES RESOURCES -->|"Backup Job"| PRIMARY_VAULT PRIMARY_VAULT -->|"Copy Job"| DR_VAULT DR_VAULT --> RP_DR style ACCOUNT fill:#111827,stroke:#1e2d45,color:#e2e8f0 style DR_ACCOUNT fill:#0f1e35,stroke:#1e2d45,color:#e2e8f0 style RESOURCES fill:#0a0e1a,stroke:#1e2d45,color:#e2e8f0 style PRIMARY_VAULT fill:#1a1a2e,stroke:#3b82f6,color:#e2e8f0 style DR_VAULT fill:#1a1a2e,stroke:#8b5cf6,color:#e2e8f0 style PLAN fill:#1e2d45,stroke:#3b82f6,color:#93c5fd style SEL fill:#1e2d45,stroke:#f59e0b,color:#fbbf24
A single Backup Plan can protect hundreds of resources simultaneously — AWS Backup runs one Backup Job per resource per rule execution.

Anatomy of a Backup Plan

A Backup Plan contains one or more rules. A selection decides which resources the plan covers. Here's an example:

  // Example: Production Backup Plan
  {
    "BackupPlanName": "prod-daily-backup-plan",
    "Rules": [
      {
        "RuleName":             "DailyToUsEast1",
        "TargetBackupVaultName": "prod-primary-vault",
        "ScheduleExpression":   "cron(0 2 * * ? *)",  // 2 AM UTC daily
        "StartWindowMinutes":   60,
        "CompletionWindowMinutes": 180,
        "Lifecycle": {
          "MoveToColdStorageAfterDays": 30,
          "DeleteAfterDays": 365
        },
        "CopyActions": [         // triggers a Copy Job after backup
          {
            "DestinationBackupVaultArn": "arn:aws:backup:eu-west-1:DR_ACCOUNT_ID:backup-vault:dr-vault",
            "Lifecycle": {
              "DeleteAfterDays": 90
            }
          }
        ]
      },
      {
        "RuleName":             "WeeklyToUsEast1",
        "TargetBackupVaultName": "prod-primary-vault",
        "ScheduleExpression":   "cron(0 3 ? * SUN *)",  // Sunday 3 AM
        "Lifecycle": { "DeleteAfterDays": 1825 } // 5 years
      }
    ],

    "Selections": [  // sent separately via create-backup-selection
      {
        "SelectionName": "all-tagged-resources",
        "IamRoleArn": "arn:aws:iam::ACCOUNT:role/service-role/AWSBackupDefaultServiceRole",
        "ListOfTags": [
          { "ConditionType": "STRINGEQUALS",
            "ConditionKey":  "Backup",
            "ConditionValue":"true" }
        ]
      }
    ]
  }
    

KEY FIELDS EXPLAINED

Field Purpose Example
ScheduleExpression Cron expression for when jobs fire cron(0 2 * * ? *) = 2 AM UTC daily
StartWindowMinutes How long the job may wait in CREATED before it starts. If it never starts it goes EXPIRED. Minimum 60; console default 8 hours 60 = job must start within 1 hour
CompletionWindowMinutes Minutes after the job actually starts before AWS Backup cancels it. Console default 7 days 180 = 3 hours max runtime
MoveToColdStorageAfterDays Days in warm storage before the recovery point moves to cold. Ignored for resource types without cold storage 30 = after 30 days → cold tier
DeleteAfterDays Days before AWS Backup deletes the recovery point. With a cold transition it must be at least MoveToColdStorageAfterDays + 90 365 = deleted after 1 year
CopyActions Starts a Copy Job to another vault once the backup completes Copy to EU DR vault

Backup Job Lifecycle

When a plan rule fires, AWS Backup creates a Job for each matching resource. Each job goes through these states:

stateDiagram-v2 [*] --> CREATED : Schedule triggers CREATED --> PENDING : Resource being prepared PENDING --> RUNNING : Snapshot started RUNNING --> COMPLETED : Backup written to vault RUNNING --> PARTIAL : Completed with partial results RUNNING --> FAILED : Error occurred RUNNING --> ABORTING : Cancellation requested ABORTING --> ABORTED : Cancelled CREATED --> EXPIRED : Not started within StartWindow COMPLETED --> [*] PARTIAL --> [*] FAILED --> [*] ABORTED --> [*] EXPIRED --> [*] note right of COMPLETED Recovery Point now in Vault. Copy Job triggers if CopyActions defined. end note

STEP-BY-STEP FLOW

1

Schedule Fires

The rule's cron expression fires. AWS Backup works out which resources the selection matches, by tag or by ARN.

2

Job Created → PENDING

AWS Backup creates one Backup Job per resource. The job waits in CREATED until it can begin, then moves to PENDING while the source service prepares the snapshot.

3

Job RUNNING — data transfer

The backup data is written to the vault. For EBS this is an EBS snapshot. For EFS and S3, AWS Backup moves the data itself. Track progress with aws backup describe-backup-job.

4

Recovery Point Created

On COMPLETED the vault holds a new recovery point with its own ARN. Creation time, resource type and encryption details are stored with it.

5

Copy Job Triggered (if configured)

If the rule has CopyActions, AWS Backup starts one Copy Job per target vault.

Copy Job — Cross-Region & Cross-Account

Copy Jobs replicate recovery points between vaults. They are the backbone of multi-region DR strategies and compliance isolation.

flowchart LR subgraph SOURCE["Source — us-east-1, Account A"] VAULT_SRC["prod-primary-vault Recovery Point: arn:...rp/abc123"] end subgraph DEST1["Same-Region Vault"] VAULT_SAME["prod-compliance-vault Copied Recovery Point locked / WORM"] end subgraph DEST2["DR Region — eu-west-1, Account B"] VAULT_DR["dr-vault Copied Recovery Point 90-day retention"] end VAULT_SRC -->|"Copy Job 1 — same-region"| VAULT_SAME VAULT_SRC -->|"Copy Job 2 — cross-region / cross-account"| VAULT_DR style SOURCE fill:#111827,stroke:#3b82f6 style DEST1 fill:#111827,stroke:#10b981 style DEST2 fill:#111827,stroke:#8b5cf6

COPY JOB REQUIREMENTS

Scenario Requirement Notes
Same-region copy A second vault in the same Region and account Good for isolating a locked compliance vault from the working vault
Cross-region copy A destination vault in the target Region, plus an IAM role AWS Backup can assume The copy is re-encrypted with the destination vault's KMS key. Recovery points already in cold storage cannot be copied
Cross-account copy Destination vault policy must allow backup:CopyIntoBackupVault for the source account Both accounts must be in the same AWS Organization, and the management account must switch on cross-account backup
Cross-account + cross-region All of the above. The destination cannot be the account's default vault Resource types AWS Backup does not fully manage need a customer managed KMS key shared with the destination account
  // Destination vault access policy (allows Account A to copy in)
  {
    "Version": "2012-10-17",
    "Statement": [{
      "Effect":    "Allow",
      "Principal": { "AWS": "arn:aws:iam::SOURCE_ACCOUNT_ID:root" },
      "Action": [
        "backup:CopyIntoBackupVault"
      ],
      "Resource":  "*"
    }]
  }
    

Restore Job — Recovering Resources

A Restore Job rebuilds an AWS resource from a recovery point. Think of it like loading a saved game. For most services, it creates a brand-new resource and never touches the original. (Exceptions: S3 can restore objects into an existing bucket; EFS can restore items into a new directory inside an existing file system.)

Do I need to spin up a new database?

Yes — always. AWS Backup cannot restore "in place". Restoring an RDS database gives you:
• A new DB instance with a brand-new hostname (e.g. prod-db-restored.abc123.rds.amazonaws.com)
• A new ARN and resource ID
• Your original database, still running and untouched

When the restore finishes you redirect your app yourself: update the connection string, the Secrets Manager secret, or the environment variables. That is deliberate — it stops a restore from overwriting a healthy database.
flowchart TD A["Operator / Automation"] --> B["1. List Recovery Points aws backup list-recovery-points-by-backup-vault"] B --> C["2. Choose a Recovery Point ARN e.g. from last night 02:30 UTC"] C --> D["3. Get required restore metadata aws backup get-recovery-point-restore-metadata"] D --> E["4. Start Restore Job aws backup start-restore-job"] E --> F{Job Status} F -->|"RUNNING 15-30 min"| G["AWS provisions new RDS instance..."] G --> F F -->|"COMPLETED"| H["New DB available new-prod-db.xyz.rds.amazonaws.com"] F -->|"FAILED"| I["Check CloudWatch Logs + DescribeRestoreJob"] H --> J["5. Validate data run queries / smoke tests"] J --> K{OK?} K -->|Yes| L["6. Cut over traffic Update Secrets Manager or Route 53 CNAME"] K -->|No| M["Try earlier recovery point"] L --> N["7. Delete old instance or keep for rollback"] style A fill:#1e2d45,stroke:#3b82f6,color:#93c5fd style H fill:#1a2e1a,stroke:#10b981,color:#6ee7b7 style I fill:#2e1a1a,stroke:#ef4444,color:#fca5a5 style L fill:#1a2e1a,stroke:#10b981,color:#6ee7b7 style N fill:#1e1a2e,stroke:#8b5cf6,color:#c4b5fd
Service What's restored New resource? Cutover needed?
RDS / Aurora New DB instance / cluster from snapshot New endpoint + ARN Yes — update connection string
EC2 (EBS) New EC2 instance from AMI created from snapshot New Instance ID + Volume IDs Update target groups / DNS
EFS New EFS file system; files land in a recovery directory, not their original paths New FS ID + DNS Yes — remount or update mount target
DynamoDB New table built from the backup you pick New table name Yes — update app table reference
S3 Objects to same or different bucket Same or new bucket Only if new bucket name
Aurora New Aurora cluster New cluster ARN + endpoint Yes — update connection string

CODE SAMPLES

Click a tab to see the restore code for each service.

What happens when you run this?
You get a new RDS instance on a new hostname, such as prod-mysql-restored-20240117.abc.us-east-1.rds.amazonaws.com. The old database keeps serving traffic until you change the connection string.
Prerequisites: AWS CLI installed & configured (aws configure), and your IAM user must have backup:* and rds:* permissions.
# ─────────────────────────────────────────────────────────────────────
  # STEP 1 — Find available recovery points (backups) in your vault
  #   This lists all RDS backups. Look at "Created" to find the one
  #   from the date/time you want to restore from.
  # ─────────────────────────────────────────────────────────────────────
  aws backup list-recovery-points-by-backup-vault \
    --backup-vault-name "prod-primary-vault" \
    --by-resource-type "RDS" \
    --query 'RecoveryPoints[*].{ARN:RecoveryPointArn,Created:CreationDate,Status:Status}' \
    --output table

  # Example output:
  # -----------------------------------------------------------------------
  # |            ListRecoveryPointsByBackupVault                          |
  # +------------------------------+-----------+---------------------------+
  # | ARN                          | Created   | Status                    |
  # +------------------------------+-----------+---------------------------+
  # | arn:aws:rds:...:awsbackup-.. | 2024-01-17| COMPLETED                 |
  # | arn:aws:rds:...:awsbackup-.. | 2024-01-16| COMPLETED                 |
  # +------------------------------+-----------+---------------------------+
  #   ↑ Copy the ARN of the backup you want to restore from

  # ─────────────────────────────────────────────────────────────────────
  # STEP 2 — Ask AWS what parameters are needed to restore this backup.
  #   AWS Backup returns a JSON object with all the config of the
  #   original DB (instance class, engine, subnet group, etc.)
  #   You'll use this in Step 3 — just change the DB name.
  # ─────────────────────────────────────────────────────────────────────
  aws backup get-recovery-point-restore-metadata \
    --backup-vault-name "prod-primary-vault" \
    --recovery-point-arn "arn:aws:rds:us-east-1:123456789:snapshot:awsbackup-2024-01-17-02-30"

  # Returns something like:
  # {
  #   "DBInstanceIdentifier": "prod-mysql",      ← original DB name
  #   "DBInstanceClass":      "db.t3.medium",    ← instance size
  #   "Engine":               "mysql",           ← database engine
  #   "MultiAZ":              "false",           ← high-availability setting
  #   "DBSubnetGroupName":    "prod-subnet-group",
  #   "VpcSecurityGroupIds":  "sg-0abc123"
  # }
  #   ↑ Copy this output. You'll paste it into Step 3, changing only
  #     DBInstanceIdentifier to a new unique name.

  # ─────────────────────────────────────────────────────────────────────
  # STEP 3 — Start the Restore Job.
  #   IMPORTANT: Change "DBInstanceIdentifier" to a NEW name.
  #   If you use the same name as the original, it will FAIL because
  #   a DB with that name already exists.
  # ─────────────────────────────────────────────────────────────────────
  aws backup start-restore-job \
    --recovery-point-arn "arn:aws:rds:us-east-1:123456789:snapshot:awsbackup-2024-01-17-02-30" \
    --iam-role-arn "arn:aws:iam::123456789:role/service-role/AWSBackupDefaultServiceRole" \
    --resource-type "RDS" \
    --metadata '{
      "DBInstanceIdentifier": "prod-mysql-restored-20240117",
      "DBInstanceClass":      "db.t3.medium",
      "Engine":               "mysql",
      "MultiAZ":              "false",
      "DBSubnetGroupName":    "prod-subnet-group",
      "VpcSecurityGroupIds":  "sg-0abc123"
    }'

  # Returns: { "RestoreJobId": "ABCDEF123456" }
  #   ↑ Save this ID — you need it to check progress in Step 4

  # ─────────────────────────────────────────────────────────────────────
  # STEP 4 — Monitor restore progress (takes 15-30 min for most DBs)
  #   Run this every few minutes. Status goes:
  #   PENDING → RUNNING → COMPLETED (or FAILED)
  # ─────────────────────────────────────────────────────────────────────
  aws backup describe-restore-job \
    --restore-job-id "ABCDEF123456"

  # When COMPLETED, you'll see:
  # { "Status": "COMPLETED", "CreatedResourceArn": "arn:aws:rds:...:db:prod-mysql-restored-20240117" }

  # ─────────────────────────────────────────────────────────────────────
  # STEP 5 — Get the hostname of the new database
  #   This is the address your app needs to connect to.
  # ─────────────────────────────────────────────────────────────────────
  aws rds describe-db-instances \
    --db-instance-identifier "prod-mysql-restored-20240117" \
    --query 'DBInstances[0].Endpoint.Address'

  # Output: "prod-mysql-restored-20240117.abc123.us-east-1.rds.amazonaws.com"
  #   ↑ This is your new database hostname

  # ─────────────────────────────────────────────────────────────────────
  # STEP 6 — Update Secrets Manager so your app picks up the new host
  #   (If you store DB credentials in Secrets Manager — recommended)
  #   After this, restart your app containers/servers to reconnect.
  # ─────────────────────────────────────────────────────────────────────
  aws secretsmanager update-secret \
    --secret-id "prod/db/connection" \
    --secret-string '{"host":"prod-mysql-restored-20240117.abc123.us-east-1.rds.amazonaws.com","port":3306,"username":"admin","password":"your-password"}'
  
After step 6: Restart your app servers so they re-read the secret. Delete the old database only once you have confirmed the new one works.

End-to-End Coordination

Here is how all components interact in a complete backup + DR + restore scenario, from schedule fire to successful restore.

sequenceDiagram participant SCHED as EventBridge Scheduler participant BACKUP as AWS Backup Service participant RESOURCE as Resource (RDS) participant VAULT_P as Primary Vault participant VAULT_DR as DR Vault (eu-west-1) participant OPS as Operator Note over SCHED,BACKUP: Backup Plan triggers at cron(0 2 * * ? *) SCHED->>BACKUP: Rule fires — create Backup Jobs BACKUP->>RESOURCE: Request snapshot / backup RESOURCE-->>VAULT_P: Stream backup data BACKUP-->>BACKUP: Backup Job: RUNNING RESOURCE-->>BACKUP: Snapshot complete BACKUP->>VAULT_P: Store Recovery Point (RP-001) BACKUP-->>BACKUP: Backup Job: COMPLETED Note over BACKUP,VAULT_DR: CopyAction defined in rule — auto-spawn Copy Job BACKUP->>VAULT_P: Read RP-001 BACKUP->>VAULT_DR: Write copy of RP-001 (cross-region) BACKUP-->>BACKUP: Copy Job: COMPLETED Note over OPS,VAULT_DR: Incident: Production DB corrupted OPS->>VAULT_DR: Browse recovery points VAULT_DR-->>OPS: List: [RP-001, RP-002, RP-003...] OPS->>BACKUP: Start Restore Job from RP-001 BACKUP->>VAULT_DR: Retrieve recovery point data BACKUP->>RESOURCE: Provision new RDS instance BACKUP-->>OPS: Restore Job: COMPLETED — new-rds-arn OPS->>OPS: Validate, then update DNS / app config

Complete Example: 3-Tier Web App

A production web app with EC2, RDS, and EFS — backed up daily with cross-region DR copies. Here's the setup and what happens each night. The clock times below are an illustration, not a guarantee: real job durations depend on data size, resource type and region.

flowchart TB subgraph APP["Production App Stack (us-east-1)"] EC2["EC2 Auto Scaling Tag: Backup=true"] RDS["RDS MySQL Tag: Backup=true"] EFS["EFS Tag: Backup=true"] end subgraph PLAN_BOX["prod-backup-plan"] RULE1["Rule: DailyBackup cron 02:00 UTC Retention: 30 days CopyTo eu-west-1"] RULE2["Rule: WeeklyBackup cron SUN 03:00 Retention: 5 years No copy"] end subgraph JOBS["Nightly Jobs (02:00 UTC)"] JOB1["Backup Job EC2 — EBS snapshot"] JOB2["Backup Job RDS snapshot"] JOB3["Backup Job EFS backup"] end subgraph VAULT1["Primary Vault (us-east-1)"] RP_EC2["RP: EC2 daily"] RP_RDS["RP: RDS daily"] RP_EFS["RP: EFS daily"] end subgraph COPY_JOBS["Copy Jobs (auto-triggered)"] CJ1["Copy: EC2 RP to eu-west-1"] CJ2["Copy: RDS RP to eu-west-1"] CJ3["Copy: EFS RP to eu-west-1"] end subgraph VAULT_DR["DR Vault (eu-west-1)"] DR_EC2["Copy: EC2 RP"] DR_RDS["Copy: RDS RP"] DR_EFS["Copy: EFS RP"] end PLAN_BOX --> APP APP --> JOBS JOB1 --> RP_EC2 JOB2 --> RP_RDS JOB3 --> RP_EFS RP_EC2 --> CJ1 RP_RDS --> CJ2 RP_EFS --> CJ3 CJ1 --> DR_EC2 CJ2 --> DR_RDS CJ3 --> DR_EFS style APP fill:#111827,stroke:#3b82f6 style PLAN_BOX fill:#111827,stroke:#f59e0b style JOBS fill:#111827,stroke:#10b981 style VAULT1 fill:#111827,stroke:#3b82f6 style COPY_JOBS fill:#111827,stroke:#8b5cf6 style VAULT_DR fill:#111827,stroke:#ef4444

NIGHTLY TIMELINE — WORKED EXAMPLE

02:00 UTC — Schedule fires CREATED

The DailyBackup rule's schedule fires. AWS Backup evaluates all resources tagged Backup=true — finds EC2, RDS, EFS. Creates 3 Backup Jobs.

02:01 — Jobs start running RUNNING

Each job begins. RDS creates a native snapshot; EBS snapshot taken for EC2 volumes; EFS backup streamed to vault. These run in parallel.

02:30 — Backup jobs complete COMPLETED

3 recovery points now stored in prod-primary-vault in us-east-1. Retention lifecycle: warm for 30 days, then auto-deleted.

02:31 — Copy Jobs auto-spawn COPY RUNNING

Because CopyActions is defined, 3 Copy Jobs are automatically created. They stream the recovery points to dr-vault in eu-west-1 under the DR account.

03:15 — Copy Jobs complete COMPLETED

All 3 recovery points are now replicated to EU. DR account has 90-day retention. Primary backups remain independent in us-east-1.

09:00 (next day) — Incident: RDS corruption detected

Ops team decides to restore RDS from the 02:30 recovery point in the DR vault in eu-west-1.

09:05 — Restore Job started RESTORE RUNNING

Restore Job provisions a new RDS instance in eu-west-1 from the copied recovery point. Parameters: new DB identifier, same instance class, target VPC.

09:25 — Restore complete COMPLETED

New RDS instance available at new endpoint. Ops validates data integrity, then updates application config / Route 53 to point to new DB. In this example that is an RPO of about 7 hours (last backup 02:30, incident 09:00) and an RTO of about 20 minutes.

What Can AWS Backup Protect?

AWS Backup supports a wide range of services. Not all features are available for every service — check the matrix below.

Category Service Continuous / PITR Cold Storage Cross-Region Copy Cross-Account Copy
Compute Amazon EC2 (incl. VSS-enabled Windows) No No Yes Yes
Block Storage Amazon EBS No Yes Yes Yes
File Storage Amazon EFS No Yes Yes Yes
File Storage Amazon FSx (Windows, Lustre, ONTAP, OpenZFS) No No Yes Yes
Object Storage Amazon S3 Yes No Yes Yes
Relational DB Amazon RDS (all engines) Yes No Yes Yes
Relational DB Amazon Aurora Yes No Yes Yes
NoSQL DB Amazon DynamoDB (advanced features required) No Yes Yes Yes
Document DB Amazon DocumentDB No No Yes Yes
Graph DB Amazon Neptune No No Yes Yes
Data Warehouse Amazon Redshift (manual snapshots only) No No No No
Time Series Amazon Timestream No Yes Yes Yes
Containers Amazon EKS No No Yes Yes
Hybrid AWS Storage Gateway (Volume) No No Yes Yes
Hybrid VMware VMs (via Backup Gateway) No Yes Yes Yes
SAP SAP HANA on EC2 Yes Yes Yes Yes
IaC AWS CloudFormation (stacks) No Yes Yes No
Continuous Backup / PITR restores to any second in the last 35 days — that 35-day cap is also why a continuous backup can never move to cold storage. AWS Backup supports it for RDS (not Multi-AZ clusters), Aurora, S3 and SAP HANA on EC2. DynamoDB has its own PITR, separate from AWS Backup.

Vault Lock, Encryption & Access Control

AWS Backup provides multiple layers of protection for your recovery points — from encryption and access policies to immutable WORM locks and legal holds.

ENCRYPTION

KMS Encryption

Every vault has an AWS KMS key. Resource types AWS Backup fully manages (S3, EFS, DynamoDB with advanced features, Timestream, CloudFormation, SAP HANA, VMware) are encrypted with that key. The rest (EBS, EC2, RDS, Aurora, FSx, DocumentDB, Neptune) keep the source resource's encryption — back up an unencrypted resource and the backup is unencrypted too.

ACCESS POLICY

Vault Access Policies

A resource-based policy on each vault controls who can create, copy or delete recovery points. Use it to deny deletion, or to let one specific account copy backups in.

CLOUDTRAIL

Audit Trail

All AWS Backup API calls are logged to CloudTrail. Every backup, copy, restore, and deletion is recorded with who did it, when, and from where — critical for compliance audits.

VAULT LOCK (WORM PROTECTION)

Vault Lock enforces a Write-Once, Read-Many (WORM) model on a vault. Once locked, recovery points cannot be deleted before their retention period expires — not even by the root user. The lock's minimum and maximum retention apply to new backup and copy jobs; recovery points already in the vault keep the lifecycle they were created with.

flowchart LR subgraph GOV["Governance Mode"] G1["Lock applied"] G2["Can be removed by privileged IAM users"] G3["Good for testing before compliance"] G1 --> G2 --> G3 end subgraph COMP["Compliance Mode"] C1["Lock applied"] C2["Minimum 72-hour cooling-off period"] C3["After cooling-off: IMMUTABLE forever"] C4["Cannot be removed by anyone — incl. AWS"] C1 --> C2 --> C3 --> C4 end style GOV fill:#111827,stroke:#f59e0b style COMP fill:#111827,stroke:#ef4444
Feature Governance Mode Compliance Mode
Removable? Yes, by a user with sufficient IAM permissions Only during the grace time you set (minimum 3 days / 72 hours). After that no user and not AWS can change or delete it
Delete recovery points early? No, unless someone removes the lock first No — AWS Backup denies the delete, including for the root user
Change the lock's min/max retention? Yes, by re-running put-backup-vault-lock-configuration Only during the grace time. After that the settings are frozen
Regulatory compliance Good for internal policy, but the lock can still be removed Assessed by Cohasset Associates for SEC 17a-4, CFTC and FINRA environments
Use case Test before committing to compliance Production compliance vaults

LEGAL HOLD

A Legal Hold stops specific recovery points from being deleted, whatever their lifecycle says. Vault Lock protects a whole vault; a Legal Hold covers only the recovery points you select, by vault, resource type, resource ID or creation date. The hold never expires — it lasts until someone with the right permissions cancels it with CancelLegalHold. Two limits to know: a hold does not cover continuous (PITR) backups, and it does not follow a recovery point that is copied to another Region or account. Each account can have 50 active holds.

LOGICALLY AIR-GAPPED VAULTS

A logically air-gapped vault is a second vault type built for ransomware recovery. AWS Backup keeps its contents in a service-owned account and always locks it in compliance mode, so recovery points cannot be deleted before their retention expires. It is encrypted with an AWS owned key by default, or a customer managed key if you supply one, and its minimum retention period is 7 days. You can share the vault with other accounts through AWS Resource Access Manager (RAM) so they can restore from it, and you can add Multi-party approval (MPA) so the backups stay recoverable even if the owning account is lost.

  # ── Governance mode: leaving out --changeable-for-days is what picks it ──
  aws backup put-backup-vault-lock-configuration \
    --backup-vault-name "prod-compliance-vault" \
    --min-retention-days 30 \
    --max-retention-days 365

  # ── Compliance mode: adding --changeable-for-days is what picks it ────
  aws backup put-backup-vault-lock-configuration \
    --backup-vault-name "prod-compliance-vault" \
    --min-retention-days 30 \
    --max-retention-days 365 \
    --changeable-for-days 3

  # --changeable-for-days is the grace time: minimum 3, maximum 36500.
  # Once it elapses, neither you, the root user nor AWS can alter this lock.
      

Lifecycle Management — Warm & Cold Storage

AWS Backup can move recovery points between storage tiers for you. Old backups you rarely touch drop to cold storage, which costs much less per GB-month.

flowchart LR A["Backup Job completes"] --> B["Recovery Point created in WARM storage"] B -->|"After N days (MoveToColdStorageAfterDays)"| C["Recovery Point moved to COLD storage"] C -->|"After N days (DeleteAfterDays)"| D["Recovery Point DELETED"] style A fill:#1e2d45,stroke:#3b82f6,color:#93c5fd style B fill:#1a2235,stroke:#f59e0b,color:#fbbf24 style C fill:#1a2235,stroke:#3b82f6,color:#60a5fa style D fill:#1a2235,stroke:#ef4444,color:#f87171
Tier Cost Retrieval Minimum Duration
Warm Storage Standard pricing per GB/month Immediate — restore anytime None
Cold Storage Much lower per GB-month — see the AWS Backup pricing page Slower retrieval, higher restore cost 90 days, billed in full even if you delete sooner
The 90-day rule: AWS Backup rejects a lifecycle where DeleteAfterDays is less than MoveToColdStorageAfterDays plus 90. Cold storage is billed for a 90-day minimum, and once a recovery point has moved to cold you can no longer change its transition day.

SERVICES SUPPORTING COLD STORAGE

Cold storage is a per-resource-type feature. AWS Backup supports it today for EBS (through EBS Snapshot Archive, so set OptInToArchiveForSupportedResources), EFS, DynamoDB with advanced features, Timestream, SAP HANA on EC2, VMware VMs and CloudFormation. EC2 instances, RDS, Aurora, S3 and Storage Gateway do not support it — those backups stay warm for the whole retention period, and AWS Backup simply ignores the cold setting for them. Cold recovery points also cannot be copied to another Region or account.

  // Example lifecycle in a backup rule
  "Lifecycle": {
    "MoveToColdStorageAfterDays": 30,   // warm for 30 days
    "DeleteAfterDays":             365   // deleted after 1 year
  }
  // 30 warm + 335 cold = 365 days total. Valid because 365 >= 30 + 90;
  // AWS Backup rejects the rule otherwise.
      

Monitoring & Alerts

Backups are only useful if they actually succeed. AWS Backup integrates with EventBridge, CloudWatch, and SNS so you always know when something goes wrong.

EVENTBRIDGE

Amazon EventBridge

Events for backup, copy and restore job state changes, plus changes to vaults and plans. AWS Backup emits them on a best-effort basis roughly every 5 minutes. Route them to Lambda, SNS, SQS or any EventBridge target.

CLOUDWATCH

CloudWatch Metrics & Alarms

AWS Backup publishes metrics to the AWS/Backup namespace every 5 minutes, with resource type and vault name as dimensions. Alarm on NumberOfBackupJobsFailed or NumberOfCopyJobsFailed to catch trouble early.

SNS

SNS Notifications

put-backup-vault-notifications subscribes one vault to events such as BACKUP_JOB_COMPLETED, BACKUP_JOB_FAILED and BACKUP_JOB_EXPIRED. The topic must be in the same account. From there, fan out to email, Slack via Lambda, or PagerDuty.

EVENTBRIDGE RULE — ALERT ON BACKUP FAILURE

The most common monitoring setup: an EventBridge rule that triggers an SNS notification whenever a backup job fails.

  # ── Create SNS topic for backup alerts ────────────────────────────────
  aws sns create-topic --name "backup-failure-alerts"

  aws sns subscribe \
    --topic-arn "arn:aws:sns:us-east-1:123456789:backup-failure-alerts" \
    --protocol "email" \
    --notification-endpoint "ops-team@company.com"

  # ── Create EventBridge rule to catch backup job failures ──────────────
  aws events put-rule \
    --name "backup-job-failed" \
    --event-pattern '{
      "source": ["aws.backup"],
      "detail-type": ["Backup Job State Change"],
      "detail": {
        "state": ["FAILED", "ABORTED", "EXPIRED"]
      }
    }'

  # ── Connect the rule to the SNS topic ─────────────────────────────────
  aws events put-targets \
    --rule "backup-job-failed" \
    --targets '[{
      "Id": "sns-target",
      "Arn": "arn:aws:sns:us-east-1:123456789:backup-failure-alerts"
    }]'
      

BACKUP AUDIT MANAGER

For compliance-heavy environments, Backup Audit Manager checks your backup activity against controls you choose — "every resource is in a backup plan", "recovery points are encrypted", "backups run at least daily". You group controls into a framework, and AWS Backup publishes a fresh report to S3 every 24 hours, plus on-demand reports. It needs AWS Config resource tracking switched on, which costs extra. You can also feed the results into AWS Audit Manager.

RESTORE TESTING

Backups you never test are backups you can't trust. A restore testing plan restores recovery points on a schedule and records how long each restore took, which is your evidence for an RTO target. Validation is your own code: keep the restored resource for 1 to 168 hours, run a check (usually an EventBridge rule firing a Lambda), then report the outcome with PutRestoreValidationResult. AWS Backup deletes the restored resource once validation finishes or the window closes.