<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Disaster-Recovery on EXPLAIN ANALYZE</title><link>https://explainanalyze.com/tags/disaster-recovery/</link><description>Recent content in Disaster-Recovery on EXPLAIN ANALYZE</description><generator>Hugo -- gohugo.io</generator><language>en-us</language><lastBuildDate>Sat, 10 Oct 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://explainanalyze.com/tags/disaster-recovery/index.xml" rel="self" type="application/rss+xml"/><item><title>Postgres, Kafka, Elasticsearch on EC2: Keep the Data on EBS, Not the Instance</title><link>https://explainanalyze.com/p/postgres-kafka-elasticsearch-on-ec2-keep-the-data-on-ebs-not-the-instance/</link><pubDate>Sat, 10 Oct 2026 00:00:00 +0000</pubDate><guid>https://explainanalyze.com/p/postgres-kafka-elasticsearch-on-ec2-keep-the-data-on-ebs-not-the-instance/</guid><description>&lt;img src="https://explainanalyze.com/" alt="Featured image of post Postgres, Kafka, Elasticsearch on EC2: Keep the Data on EBS, Not the Instance" /&gt;&lt;div class="tldr-box"&gt;
 &lt;strong&gt;TL;DR&lt;/strong&gt;
 &lt;div&gt;For a self-managed storage system on EC2 (a relational database, a Kafka broker, an Elasticsearch node, a Redis replica), moving EBS volumes is far faster and cheaper than draining a node and refilling its replacement over the network: a resize or a replacement reattaches the data, and an OS upgrade swaps only the root volume while the data volumes stay attached. The price is a network hop on every fsync, one EBS ceiling per instance that every volume shares, and snapshot hydration inside the DR plan&amp;rsquo;s RTO.&lt;/div&gt;
&lt;/div&gt;

&lt;p&gt;This is the first of two posts on running self-managed stateful systems on EC2 (PostgreSQL and MySQL, Kafka, Elasticsearch and OpenSearch, Redis with persistence on) with storage and compute separated. This one covers why; Part 2 covers how the DR side gets built on AWS.&lt;/p&gt;
&lt;p&gt;The two storage options behave very differently. Instance-store NVMe is physically attached to the host, fast, and ephemeral: per the &lt;a class="link" href="https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/instance-store-lifetime.html" target="_blank" rel="noopener"
 &gt;AWS docs&lt;/a&gt;, its data does not survive a stop, a hibernation, a host retirement, or an automatic recovery, and the blocks are cryptographically erased. EBS is network-attached block storage with its own lifecycle. A volume outlives the instance it&amp;rsquo;s attached to, detaches and reattaches, snapshots incrementally to S3, and changes size and performance while mounted. When the data is on the instance, every instance event becomes a data event. When it&amp;rsquo;s on EBS, most of them become a reattach.&lt;/p&gt;
&lt;h2 id="why-not-the-managed-service"&gt;Why not the managed service
&lt;/h2&gt;&lt;p&gt;The managed services are the obvious alternative: RDS or Aurora, Amazon MSK, Amazon OpenSearch Service, ElastiCache. Aurora is the extreme form of this post&amp;rsquo;s argument (the &lt;a class="link" href="https://www.amazon.science/publications/amazon-aurora-design-considerations-for-high-throughput-cloud-native-relational-databases" target="_blank" rel="noopener"
 &gt;SIGMOD 2017 paper&lt;/a&gt; pushes redo processing into a storage service so compute holds almost no state), and for plenty of teams it&amp;rsquo;s the correct call. The reasons to stay on EC2 are concrete. I/O-heavy workloads cost far more on a managed service, and AWS&amp;rsquo;s own pricing admits it: &lt;a class="link" href="https://aws.amazon.com/about-aws/whats-new/2023/05/amazon-aurora-i-o-optimized/" target="_blank" rel="noopener"
 &gt;Aurora I/O-Optimized&lt;/a&gt; (May 2023) drops per-I/O charges and is pitched at clusters where I/O exceeds 25 percent of the Aurora bill. A multi-cloud estate, or a team that won&amp;rsquo;t be locked to one vendor, keeps the engine, its tooling, and the runbooks portable; EBS is AWS-specific, but the same split works on GCP Persistent Disk or Azure Managed Disks. Then there are builds the service doesn&amp;rsquo;t offer (Percona Server, a patched PostgreSQL, Elasticsearch past 7.10, the last version &lt;a class="link" href="https://aws.amazon.com/opensearch-service/faqs/" target="_blank" rel="noopener"
 &gt;Amazon OpenSearch Service runs&lt;/a&gt;), plugins outside the allow-list, kernel tuning, and constraints a managed service can&amp;rsquo;t accommodate. Without one of those, the managed service usually wins.&lt;/p&gt;
&lt;h2 id="reattach-the-volume-instead-of-draining-the-node"&gt;Reattach the volume instead of draining the node
&lt;/h2&gt;&lt;p&gt;The alternative to moving a volume is draining a node, and every system has its own version. Elasticsearch excludes the node from allocation and &lt;a class="link" href="https://www.elastic.co/docs/reference/elasticsearch/configuration-reference/cluster-level-shard-allocation-routing-settings" target="_blank" rel="noopener"
 &gt;relocates its shards&lt;/a&gt; to the rest of the cluster. Kafka reassigns the broker&amp;rsquo;s partitions with a plan the &lt;a class="link" href="https://kafka.apache.org/43/operations/basic-kafka-operations/" target="_blank" rel="noopener"
 &gt;operations docs&lt;/a&gt; say has to be written by hand for a decommission, usually under a replication throttle (the docs&amp;rsquo; example is 50 MB/s between brokers) so the move doesn&amp;rsquo;t starve producers. Zalando&amp;rsquo;s Kafka team &lt;a class="link" href="https://engineering.zalando.com/posts/2017/10/reattaching-kafka-ebs-in-aws.html" target="_blank" rel="noopener"
 &gt;put the cost at&lt;/a&gt; around 6–7 hours of rebalancing for one broker lost with its disk. A relational database builds a new replica and fails over to it. Redis does a &lt;a class="link" href="https://redis.io/docs/latest/operate/oss_and_stack/management/replication/" target="_blank" rel="noopener"
 &gt;full sync&lt;/a&gt;, snapshotting the whole dataset on the primary and shipping it. Every one of them copies the node&amp;rsquo;s whole dataset across the network while the cluster carries the extra load, and a rolling replacement of Elasticsearch or Kafka nodes on local disk pays it twice per node: once to drain the old node, once to fill the new one.&lt;/p&gt;
&lt;p&gt;Moving the volume copies nothing. Stop the process cleanly, detach, attach to the new instance, start. PostgreSQL and MySQL come up without recovery after a clean shutdown, and an unclean stop costs only a log replay.&lt;/p&gt;
&lt;p&gt;OS patching doesn&amp;rsquo;t even need a new instance once the OS lives on its own volume. EC2&amp;rsquo;s &lt;a class="link" href="https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/replace-root.html" target="_blank" rel="noopener"
 &gt;root volume replacement&lt;/a&gt; swaps the root volume of a running instance for one built from a new AMI and reboots it. Per the docs, the data EBS volumes stay attached, instance-store data survives (temp directories included), and the instance stays on the same host with the same network interfaces, IP addresses, and DNS name. To the rest of the cluster the node was briefly away and came back on the same address with the same data. Two conditions apply. The data directory must not sit on the root volume, which the replacement detaches. And the AMI must match the instance&amp;rsquo;s architecture, so a move to Graviton is still a new instance and a volume move.&lt;/p&gt;
&lt;p&gt;Resizing is the same move: an &lt;code&gt;r7i.4xlarge&lt;/code&gt; that ran out of memory becomes an &lt;code&gt;r7i.8xlarge&lt;/code&gt; and the data stays put. Compute gets sized for CPU and memory, never for disk. For Redis, memory includes persistence headroom: the &lt;a class="link" href="https://redis.io/docs/latest/operate/oss_and_stack/management/admin/" target="_blank" rel="noopener"
 &gt;admin guide&lt;/a&gt; warns that a write-heavy instance can use up to twice its normal memory while the forked child writes an RDB or rewrites the AOF.&lt;/p&gt;
&lt;p&gt;A Kafka broker that comes back on its old volumes, under the same broker identity, holds every segment it had and fetches only what was produced while it was gone. Zalando&amp;rsquo;s &lt;a class="link" href="https://engineering.zalando.com/posts/2017/10/reattaching-kafka-ebs-in-aws.html" target="_blank" rel="noopener"
 &gt;result&lt;/a&gt; was a lost broker recovered &amp;ldquo;in a number of minutes without copying data,&amp;rdquo; and an upgrade of 15 brokers in two hours, 42 times faster than before. An Elasticsearch node that returns with its data inside the delayed-allocation window gets its shards back &amp;ldquo;with the minimum of network traffic,&amp;rdquo; per the &lt;a class="link" href="https://www.elastic.co/docs/deploy-manage/distributed-architecture/shard-allocation-relocation-recovery/delaying-allocation-when-node-leaves" target="_blank" rel="noopener"
 &gt;delayed allocation docs&lt;/a&gt;. The default is one minute, shorter than most instance launches, so a planned swap raises it first.&lt;/p&gt;
&lt;p&gt;Redis is the odd one: a replica shut down cleanly keeps its replication offset in its snapshot file and can resume where it left off, while one restarting from its append-only log reloads from the volume and takes a full sync anyway.&lt;/p&gt;
&lt;p&gt;Storage changes are online. &lt;a class="link" href="https://docs.aws.amazon.com/ebs/latest/userguide/ebs-modify-volume.html" target="_blank" rel="noopener"
 &gt;Elastic Volumes&lt;/a&gt; grows size, IOPS, and throughput on a volume while it serves traffic. Nothing shrinks, so oversizing is a one-way door.&lt;/p&gt;
&lt;p&gt;Engine upgrades are the same root swap, because the binaries live on the root volume and the data doesn&amp;rsquo;t. Bake the new PostgreSQL, MySQL, Kafka, Elasticsearch, or Redis version into the AMI, replace the root volume, and the process starts on the new version against the same data directory. For a minor version that&amp;rsquo;s the whole job. A major version also converts the data on disk, which is where the rollback below comes in. The one move that does need a copy is crossing Availability Zones, since an EBS volume lives in one: snapshot it, create a volume from the snapshot in the target AZ, attach.&lt;/p&gt;
&lt;p&gt;Major versions get the rollback neither relational engine provides. The &lt;a class="link" href="https://www.postgresql.org/docs/current/pgupgrade.html" target="_blank" rel="noopener"
 &gt;&lt;code&gt;pg_upgrade&lt;/code&gt; docs&lt;/a&gt; say that in link mode, once the new cluster has started, the old one is unsafe and must be restored from backup. MySQL 8.0 to 8.4 upgrades the data dictionary in place on first start, and the &lt;a class="link" href="https://dev.mysql.com/doc/refman/8.4/en/downgrading.html" target="_blank" rel="noopener"
 &gt;downgrade docs&lt;/a&gt; allow going back only by logical dump and load or replication. Root volume replacement can keep the old root volume instead of deleting it, so the rollback is swapping that volume back in and restoring the data volume from a snapshot taken after a clean shutdown. Upgrade a pre-warmed copy of the data volume, not the original, or the rollback waits on hydration.&lt;/p&gt;
&lt;p&gt;Each of these moves has one trap on the PostgreSQL side.&lt;/p&gt;
&lt;div class="warning-box"&gt;
 &lt;strong&gt;Warning&lt;/strong&gt;
 &lt;div&gt;Attaching a PostgreSQL data volume to an instance with a newer C library can silently corrupt every text index on it. PostgreSQL collates through glibc by default, and glibc 2.28 rewrote the sort order for most locales; per the &lt;a class="link" href="https://wiki.postgresql.org/wiki/Locale_data_changes" target="_blank" rel="noopener"
 &gt;PostgreSQL wiki&lt;/a&gt;, B-tree indexes on &lt;code&gt;text&lt;/code&gt;, &lt;code&gt;varchar&lt;/code&gt;, &lt;code&gt;char&lt;/code&gt;, and &lt;code&gt;citext&lt;/code&gt; built under the old rules must be reindexed before the instance takes production traffic. &lt;a class="link" href="https://docs.aws.amazon.com/linux/al2023/ug/glibc-gcc-and-binutils.html" target="_blank" rel="noopener"
 &gt;Amazon Linux 2 ships glibc 2.26 and AL2023 ships 2.34&lt;/a&gt;, and AL2 &lt;a class="link" href="https://aws.amazon.com/amazon-linux-2/faqs/" target="_blank" rel="noopener"
 &gt;reached end of support on June 30, 2026&lt;/a&gt;, so &amp;ldquo;new AMI, reattach the volume&amp;rdquo; this year crosses that line. The symptom is a &lt;code&gt;UNIQUE&lt;/code&gt; index admitting duplicates and lookups missing rows that exist. Pin the OS major version of database AMIs or put the &lt;code&gt;REINDEX&lt;/code&gt; in the cut-over plan, and consider ICU collations, which carry a version PostgreSQL can check. MySQL ships its own collations and doesn&amp;rsquo;t inherit this from the OS, though its collations do move between MySQL versions, &lt;a class="link" href="https://explainanalyze.com/p/collation-drift-the-silent-index-killer/" target="_blank" rel="noopener"
 &gt;a different trap&lt;/a&gt;.&lt;/div&gt;
&lt;/div&gt;

&lt;h2 id="split-the-io-then-hit-the-instance-ceiling"&gt;Split the I/O, then hit the instance ceiling
&lt;/h2&gt;&lt;p&gt;Once the data is on volumes, nothing says it has to be on one. Separate volumes give each write stream its own IOPS budget and its own queue.&lt;/p&gt;
&lt;p&gt;The relational engines have the finest controls. PostgreSQL can put WAL on one volume, the data directory on another, and give a hot table, a large index, or the current partitions of a time-series table a tablespace of their own. MySQL goes further: binlog, redo, undo, and doublewrite can each sit on their own volume, and &lt;a class="link" href="https://dev.mysql.com/doc/refman/8.4/en/innodb-create-table-external.html" target="_blank" rel="noopener"
 &gt;individual tables or partitions&lt;/a&gt; can live outside the data directory. The win is isolation as much as aggregate. A commit waits on a log fsync, and on its own volume that fsync doesn&amp;rsquo;t queue behind a checkpoint flushing gigabytes of dirty pages. A durable MySQL commit syncs both redo and binlog, and separate volumes keep the two from contending.&lt;/p&gt;
&lt;p&gt;Kafka and Elasticsearch moved in opposite directions on the same idea. Kafka can spread a broker&amp;rsquo;s log across several directories; per the &lt;a class="link" href="https://kafka.apache.org/43/operations/hardware-and-os/" target="_blank" rel="noopener"
 &gt;hardware docs&lt;/a&gt;, partitions are assigned round-robin across directories and each sits entirely in one, so an uneven partition mix means uneven disks. JBOD under KRaft left early access &lt;a class="link" href="https://kafka.apache.org/blog/2024/07/29/apache-kafka-3.8.0-release-announcement/" target="_blank" rel="noopener"
 &gt;in Kafka 3.8&lt;/a&gt; (July 2024). Elasticsearch &lt;a class="link" href="https://www.elastic.co/docs/reference/elasticsearch/configuration-reference/path" target="_blank" rel="noopener"
 &gt;deprecated multiple data paths in 7.13&lt;/a&gt; and recommends one filesystem spanning several disks or one node per data path, which on EC2 means one data volume per node and more nodes.&lt;/p&gt;
&lt;p&gt;The split also enables tiering. Cold partitions go on baseline gp3, and provisioned io2 pays only for hot data. PostgreSQL can move a partition to another tablespace, but the move rewrites it under an exclusive lock, so it suits partitions that no longer take writes. MySQL decides placement when &lt;a class="link" href="https://explainanalyze.com/p/designing-partitioning-you-dont-have-to-babysit/" &gt;the partition rotation&lt;/a&gt; creates the partition; moving it later means a stopped server. Kafka and OpenSearch take the tier off block storage entirely. Kafka&amp;rsquo;s &lt;a class="link" href="https://kafka.apache.org/43/operations/tiered-storage/" target="_blank" rel="noopener"
 &gt;tiered storage&lt;/a&gt;, production-ready &lt;a class="link" href="https://kafka.apache.org/blog/2024/11/06/apache-kafka-3.9.0-release-announcement/" target="_blank" rel="noopener"
 &gt;since 3.9&lt;/a&gt; (November 2024), copies completed log segments to S3 and keeps only the tail on the broker, though not for compacted topics. OpenSearch&amp;rsquo;s &lt;a class="link" href="https://docs.opensearch.org/latest/tuning-your-cluster/availability-and-recovery/remote-store/index/" target="_blank" rel="noopener"
 &gt;remote-backed storage&lt;/a&gt; (2.10) does the same for segments and translog.&lt;/p&gt;
&lt;p&gt;Temp needs no durability, which makes it the natural tenant of local NVMe on instance types that have it. Sort spills, hash joins, and the scratch files of online schema changes run at local-disk speed and stay off the EBS budget, and both PostgreSQL and MySQL can point their temporary files at the instance store. The PostgreSQL docs warn that &lt;a class="link" href="https://www.postgresql.org/docs/current/manage-ag-tablespaces.html" target="_blank" rel="noopener"
 &gt;a tablespace on temporary storage risks the cluster&lt;/a&gt;. One holding only temp files is the survivable exception, provided its directory is recreated on every boot, because instance store comes back empty after any stop.&lt;/p&gt;
&lt;div class="note-box"&gt;
 &lt;strong&gt;Note&lt;/strong&gt;
 &lt;div&gt;The engines fail in opposite directions when that boot step is missing, tested on PostgreSQL 17 and MySQL 8.4 with the temp directory wiped between restarts. MySQL refuses to start, which is loud. PostgreSQL starts normally, fails to create temporary tables, and quietly writes sort spills to the data volume instead, the exact I/O the split was meant to keep off it. Nothing in the startup log says so.&lt;/div&gt;
&lt;/div&gt;

&lt;p&gt;Then the ceiling. Every volume on an instance shares that instance&amp;rsquo;s EBS bandwidth, and it sits lower than the volumes&amp;rsquo; combined limits suggest. Since &lt;a class="link" href="https://aws.amazon.com/about-aws/whats-new/2025/09/amazon-ebs-size-provisioned-performance-gp3-volumes" target="_blank" rel="noopener"
 &gt;September 2025&lt;/a&gt; one &lt;a class="link" href="https://docs.aws.amazon.com/ebs/latest/userguide/general-purpose.html" target="_blank" rel="noopener"
 &gt;gp3 volume&lt;/a&gt; provisions up to 80,000 IOPS and 2,000 MiB/s, and one &lt;a class="link" href="https://docs.aws.amazon.com/ebs/latest/userguide/provisioned-iops.html" target="_blank" rel="noopener"
 &gt;io2 Block Express&lt;/a&gt; volume 256,000 IOPS and 4,000 MiB/s. Per the &lt;a class="link" href="https://docs.aws.amazon.com/ec2/latest/instancetypes/mo.html" target="_blank" rel="noopener"
 &gt;instance tables&lt;/a&gt;, an &lt;code&gt;r7i.8xlarge&lt;/code&gt; gets 40,000 IOPS and 1,250 MB/s in total: one maxed gp3 already exceeds the instance, and a second volume buys isolation, not headroom. Only from &lt;code&gt;r7i.24xlarge&lt;/code&gt; (120,000 IOPS, 3,750 MB/s) do several volumes add up to more than one. Sizes through &lt;code&gt;4xlarge&lt;/code&gt; advertise a burst they hold for 30 minutes a day before dropping to baseline, the same trap as &lt;a class="link" href="https://explainanalyze.com/p/its-almost-always-the-queries-part-v-disk-has-two-alarms-not-one/" &gt;gp2 burst balance&lt;/a&gt;. Splitting scales I/O and nothing else. A hot row, a replica applier that can&amp;rsquo;t parallelize the primary&amp;rsquo;s writes, PostgreSQL waiting on its WAL write lock, a hot Kafka partition with one leader: none improve with a third volume, and the headroom before sharding or adding nodes ends at a number in the instance-type table.&lt;/p&gt;
&lt;p&gt;Latency is the other price, and how much of it a system pays depends on whether it waits for the disk. AWS describes gp3 as &lt;a class="link" href="https://docs.aws.amazon.com/ebs/latest/userguide/general-purpose.html" target="_blank" rel="noopener"
 &gt;single-digit millisecond&lt;/a&gt; and io2 Block Express as under 500 microseconds on average for a 16 KiB I/O; local NVMe has no network hop in the path. The relational engines wait: every commit pays a WAL or redo fsync. Group commit amortizes it across concurrent transactions, so most OLTP never notices, but small serial commits will see their p99 move. Kafka mostly doesn&amp;rsquo;t wait. Its docs &lt;a class="link" href="https://kafka.apache.org/43/operations/hardware-and-os/" target="_blank" rel="noopener"
 &gt;recommend the default flush settings&lt;/a&gt;, which &amp;ldquo;disable application fsync entirely,&amp;rdquo; because &amp;ldquo;a failed node will always recover from its replicas,&amp;rdquo; so EBS latency stays out of the produce path and the volume&amp;rsquo;s job is throughput. Redis sits between them: its default once-a-second fsync runs in the background, but per the &lt;a class="link" href="https://redis.io/docs/latest/operate/oss_and_stack/management/optimization/latency/" target="_blank" rel="noopener"
 &gt;latency guide&lt;/a&gt; a slow one eventually blocks the main thread. Where p99 is set by fsync time and nothing can hide it, a primary on local NVMe with synchronous replicas on NVMe is the better design, and the re-seed window is the price of the latency.&lt;/p&gt;
&lt;h2 id="the-dr-plan-is-a-snapshot-schedule-and-a-hydration-rate"&gt;The DR plan is a snapshot schedule and a hydration rate
&lt;/h2&gt;&lt;p&gt;EBS snapshots are incremental and copy across regions, which supports two shapes of standby in the second region. Cold standby keeps only snapshots; failover creates volumes, launches instances, and attaches. Warm standby keeps volumes already restored from the latest snapshot, with no compute running. Both cost a fraction of a running replica, and both depend on things the cost estimate doesn&amp;rsquo;t show, starting with whether the snapshot is a usable copy at all.&lt;/p&gt;
&lt;p&gt;A snapshot is only a backup if it&amp;rsquo;s consistent, and splitting the log from the data breaks that by default. Snapshot the volumes one at a time and each captures a different instant. The &lt;a class="link" href="https://www.postgresql.org/docs/current/backup-file.html" target="_blank" rel="noopener"
 &gt;PostgreSQL docs&lt;/a&gt; require simultaneous snapshots when data files and WAL sit on different disks, and on MySQL a binlog captured ahead of the redo log produces replicas that diverge from their first minute. EBS can &lt;a class="link" href="https://docs.aws.amazon.com/ebs/latest/userguide/ebs-create-snapshots.html" target="_blank" rel="noopener"
 &gt;snapshot several volumes at the same instant&lt;/a&gt;, which gives crash consistency: what a power cut would leave, and what both engines recover from by design. Knowing where that instant sits in the log stream, so a restore can replay forward from it, takes the engine&amp;rsquo;s own backup hooks around the snapshot.&lt;/p&gt;
&lt;p&gt;Some systems don&amp;rsquo;t accept a volume snapshot as a backup at all. Elasticsearch&amp;rsquo;s &lt;a class="link" href="https://www.elastic.co/docs/deploy-manage/tools/snapshot-and-restore" target="_blank" rel="noopener"
 &gt;snapshot docs&lt;/a&gt; call its snapshot API &amp;ldquo;the only reliable and supported way to back up a cluster&amp;rdquo; and rule out restoring from filesystem-level copies. Kafka&amp;rsquo;s cross-region answer is cluster-to-cluster mirroring with &lt;a class="link" href="https://kafka.apache.org/43/operations/geo-replication-cross-cluster-data-mirroring/" target="_blank" rel="noopener"
 &gt;MirrorMaker 2&lt;/a&gt;, since one broker&amp;rsquo;s volume holds only that broker&amp;rsquo;s replicas. Redis treats its RDB file as the DR artifact. For these, EBS carries a node through instance events and the system&amp;rsquo;s own mechanism carries DR. Redis&amp;rsquo;s &lt;a class="link" href="https://redis.io/docs/latest/operate/oss_and_stack/management/replication/" target="_blank" rel="noopener"
 &gt;replication docs&lt;/a&gt; also hold the sharpest version of the instance-store risk: a primary that restarts empty is still a primary, and replicas that sync from it &amp;ldquo;will effectively destroy their copy of the data.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Warm standby isn&amp;rsquo;t warm, either. A volume restored from a snapshot is usable at once, but its blocks &lt;a class="link" href="https://docs.aws.amazon.com/ebs/latest/userguide/ebs-initialize.html" target="_blank" rel="noopener"
 &gt;stay in S3&lt;/a&gt; until first read, and every page arrives with the enthusiasm of a shared drive on hotel Wi-Fi. RTO includes that hydration. It gets paid in time, by reading every block before cutover, or in money, since AWS sells faster initialization. And it starts over on every refresh, so a standby is warm only if something reads it after each one.&lt;/p&gt;
&lt;p&gt;The standby doesn&amp;rsquo;t have to be provisioned for production while it waits. Kept at baseline performance it costs little beyond its gigabytes, and failover attaches compute and raises the volumes to production performance online. The raise isn&amp;rsquo;t instant: the &lt;a class="link" href="https://docs.aws.amazon.com/ebs/latest/userguide/monitoring-volume-modifications.html" target="_blank" rel="noopener"
 &gt;modification docs&lt;/a&gt; put performance changes at &amp;ldquo;a few minutes to a few hours,&amp;rdquo; longer on a volume that was never fully hydrated. Through the change the volume serves at no less than its old performance. That fits an RTO of hours. An RTO of minutes means paying for production performance all the time.&lt;/p&gt;
&lt;p&gt;Recovery point is the other number. Snapshots alone lose everything written since the last one, plus the copy time. Shipping the log continuously to the DR region (WAL archiving on PostgreSQL, binlog streaming on MySQL) shrinks that to the shipping lag, and a restore becomes snapshot plus replay. None of it helps with a write that was wrong when it landed, which every snapshot preserves faithfully; &lt;a class="link" href="https://explainanalyze.com/p/before-the-bad-write/" &gt;that needs a different layer&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Separation covers instance failure, host failure, and planned work, and not the storage layer. In the &lt;a class="link" href="https://aws.amazon.com/message/65648/" target="_blank" rel="noopener"
 &gt;April 2011 us-east-1 event&lt;/a&gt;, a misrouted network change set off a re-mirroring storm inside EBS that left about 13 percent of volumes in one AZ stuck for days, beyond reach of any instance replacement. Replacement also depends on the EC2 control plane: in the &lt;a class="link" href="https://aws.amazon.com/message/101925/" target="_blank" rel="noopener"
 &gt;October 2025 us-east-1 outage&lt;/a&gt;, new launches failed from 2:25 to 10:36 AM PDT while running instances kept running. A recovery that starts with a launch is only as available as the launch API. A replica already running elsewhere removes that dependency, but it&amp;rsquo;s compute paid for around the clock, and if the DR side serves no reads it does nothing until the day it&amp;rsquo;s needed, which is the cost a volume-based standby exists to avoid.&lt;/p&gt;
&lt;p&gt;Part 2 turns the DR half into an implementation outline: which AWS services schedule, copy, and restore the snapshots, how log shipping or the system&amp;rsquo;s own backup API fits alongside them, the hydration options and what they cost, and how a restore gets drilled against a stopwatch.&lt;/p&gt;</description></item></channel></rss>