AWS Database Blog
A decision framework for Amazon Neptune availability
When your Amazon Neptune workload requires resilience beyond what a single cluster provides, choosing the right patterns requires a decision framework. Such requirements include tolerance for Regional impairments or near-zero RPO during failover. This post helps you decide which patterns to layer on top of the Neptune baseline SLA, and in what order.
Many organizations building fraud detection systems, recommendation engines, or knowledge graphs encounter these tradeoffs. A Regional impairment during peak transactions can mean lost revenue, regulatory penalties, or eroded customer trust. Without a structured approach, teams either over-invest in resilience or discover gaps during an incident. This post maps each pattern to the workload characteristics that make it the right choice.
In this post, we cover multi-Region deployments with Neptune Global Database, cross-Region failover strategies, write durability tradeoffs during recovery, and single-Region resilience patterns. We provide a decision framework table (at the end) comparing all patterns side by side to help you match Recovery Time Objective (RTO), Recovery Point Objective (RPO), and complexity targets to your workload. The multi-Region patterns in this post follow a consistent principle: deploy an identical application stack in each Region with a regional endpoint, and fail over at the endpoint layer rather than migrating the full stack.
Understanding Neptune availability considerations
Neptune provides continuous read and write availability under normal conditions. The writer instance becomes temporarily unavailable during configuration changes, scaling operations, in-place engine version upgrades, and promotion events.
Events that cause writer unavailability
The following events cause writer unavailability, listed from most to least common:
The most common cause of writer unavailability is a static parameter change. Parameters like neptune_streams require an instance reboot to take effect, while dynamic parameters apply immediately with no interruption.
Scaling operations are the second most common event. Changing the writer instance type requires an in-place restart. However, since engine 1.2.0.0, reader instances remain available during writer restarts, so you can scale additional readers independently to maintain read availability throughout the process.
Failover promotion occurs when Neptune promotes a read replica to become the new writer. With one or more replicas present, service typically resumes in under 60 seconds. Without replicas, Neptune must create a new primary instance, which takes considerably longer. This is why maintaining at least two read replicas (three instances total) is a baseline best practice. After a failover promotes one replica to writer, the remaining replica continues serving read traffic without interruption.
In-place engine version upgrades cause all instances in a cluster to restart simultaneously. Minor versions apply in-place automatically during maintenance windows. Using the Neptune Blue/Green solution to perform blue-green updates reduces the impact window for major version upgrades.
In Global Database deployments prior to engine 1.4.0.0, a primary writer restart triggered secondary AWS Region restarts, disrupting reads across all Regions. Version 1.4.0.0, released November 2024, introduced global database survivable replica, allowing secondary readers to continue serving requests during primary writer restarts.
Neptune SLA commitments
The following SLA commitments provide context for the preceding unavailability events and help inform which availability patterns to adopt. In April 2025, Neptune increased its Multi-AZ SLA from 99.9 percent to 99.99 percent.
| Deployment Type | SLA |
| Multi-AZ DB Instance, Multi-AZ DB Cluster, Multi-AZ Graph | 99.99% |
| Single DB Instance, Single-AZ Graph | 99.5% |
| Global Database (each cluster covered by Multi-AZ SLA) | 99.99% per Region |
For current commitments, see the Amazon Neptune Service Level Agreement.
Multi-Region architecture with Neptune Global Database
Neptune Global Database, generally available since July 2022, replicates data across Regions at the storage layer. This approach operates independently of database instances, consuming no instance CPU or memory. Your clusters remain fully available for application workloads throughout replication.
The architecture follows a top-down, per-Region design. Each Region exposes the same regional endpoint and runs the same stack beneath it, including the application layer, write queue or write-ahead log (WAL), and Neptune cluster. During a cross-Region failover, you redirect traffic at the regional endpoint, typically through Amazon Route 53 DNS routing, rather than migrating application infrastructure between Regions. The secondary Region’s stack is already running and serving read traffic. Promotion shifts write responsibility and redirects the Regional endpoint.
Figure 1: Multi-Region architecture with Neptune Global Database. Each Region runs an identical stack with a regional endpoint at the top. Failover redirects traffic at the endpoint layer.
The primary Region’s storage volume streams redo log records to each secondary Region without consuming instance CPU or memory. Under normal write volumes, replication lag remains below one second, measurable through the GlobalDbProgressLag Amazon CloudWatch metric. You can deploy up to five secondary clusters across different Regions, each supporting up to 16 read replicas. Starting with engine version 1.4.0.0, secondary Region readers continue serving requests even during primary writer restarts through global database survivable replica.
Manual failover as the default approach
Neptune doesn’t offer automated cross-Region failover. This is a deliberate design choice. Cross-Region failover is a distributed consensus problem: a system must reliably determine that your workload in the primary Region is degraded, and that determination must itself be resilient to network partitions and split-brain scenarios.
Automated cross-Region failover introduces the following risks:
- False positives: A transient network issue between monitoring infrastructure and the primary Region could trigger an unnecessary promotion. This causes a write pause during the promotion process.
- Split-brain scenarios: If both Regions believe they are the primary writer, data divergence becomes possible. Resolving split-brain after the fact is significantly more complex than the initial service event.
- Cascading failures: An automated failover during a broader infrastructure event could add unnecessary load to surviving infrastructure, worsening the situation.
For these reasons, Neptune provides two cross-Region failover operations, switchover and failover, leaving the triggering decision to your operations team or your own automation with appropriate safeguards.
Cross-Region failover strategies
Neptune Global Database offers two failover mechanisms, each designed for different scenarios. Switchover provides zero data loss for proactive Region relocation. Failover handles unplanned events when the primary Region is experiencing an impairment that prevents write operations.
Switchover (RPO = 0)
Use switchover when you want to proactively relocate the primary cluster to a different Region. Common triggers include bringing the writer closer to your user base or preparing for a planned maintenance event. You can also use it to validate your disaster recovery procedures during a Game Day exercise.
During a switchover:
- Neptune pauses writes on the current primary and waits for all secondary clusters to converge, meaning replication lag reaches zero.
- The selected secondary cluster is promoted to become the new primary writer.
- The previous primary becomes a secondary cluster.
- Applications reconnect using the Global Database writer endpoint.
Failover
Use failover when the primary Region is experiencing a service event and is unreachable. Unlike switchover, failover doesn’t wait for replication convergence, so writes that had not yet replicated are stranded in the primary Region.
The RPO equals the replication lag at the time of failover, which is typically under one second but could be higher during periods of heavy write activity. After promotion, the new primary begins accepting writes immediately.
After a failover, remove the previous primary Region from the Global Database and re-add it as a secondary. Any in-flight writes that did not replicate before the event remain inaccessible until the original Region recovers. To avoid this, implement a WAL pattern as described in the “Write durability during cross-Region recovery” section.
Write durability during cross-Region recovery
The preceding failover mechanism restores write availability in a secondary Region, but writes buffered in the primary Region remain stranded until recovery. The following patterns address that gap.
In a multi-Region Neptune deployment, you can decouple write availability from writer availability. For details, see Building zero-downtime write architectures for Amazon Neptune. Your application sends writes to a regional queue. A consumer process reads from the queue and applies mutations to Neptune. This lets your application continue accepting writes even when the Neptune writer is unavailable.
During a cross-Region failover, writes buffered in the primary Region’s queue haven’t yet been processed. You have two options for addressing this. Option A uses Amazon Simple Queue Service (Amazon SQS) or Amazon Kinesis Data Streams and accepts delayed write visibility until the primary Region recovers. Option B adds an Amazon DynamoDB Global Table as a write-ahead log that replicates across Regions for near-zero RPO. The examples in this post use Amazon SQS, but Amazon Kinesis Data Streams is equally valid for high-throughput workloads.
Option A: Accept delayed visibility (simpler, lower cost)
Keep the dual-Region write queue architecture. Accept that the primary Region’s SQS queue strands writes until the Region recovers. When it does, the consumer drains the backlog and reconciles the data.
Consider this approach for workloads where:
- Your application can tolerate seconds to minutes of recent writes not visible in the secondary Region after failover.
- You want to minimize infrastructure cost and operational complexity.
- Your write volume is low enough that the queue backlog at any given moment is small.
- You have a defined recovery procedure for reconciling pending writes after the Region recovers.
Option B: Amazon DynamoDB Global Table as write-ahead log
In this pattern, your application writes each graph mutation as a row (or item) in a DynamoDB Global Table, using a timestamp or sequence number as the sort key. A consumer process in the active Region reads from the local DynamoDB replica and writes to Neptune. After a successful write, the consumer marks the row as processed.
DynamoDB Global Tables replicate items across Regions asynchronously, typically within one second. After a Regional failover, the secondary Region’s consumer activates and picks up from the last unprocessed row. To maintain idempotency, the consumer checks the processed flag before writing to Neptune. This prevents duplicate mutations if a row is replayed.
Important: This pattern depends on DynamoDB Global Tables availability. If DynamoDB experiences a Regional impairment simultaneously with Neptune, the WAL might not replicate before failover, reverting RPO to DynamoDB replication lag. Global Tables use last-writer-wins reconciliation based on timestamps, not strict ordering. Design mutation schemas to be conflict-free under eventual consistency. For workloads requiring strict ordering, make sure your consumer activation logic prevents both Regions from processing the same WAL entries simultaneously during the failover transition.
This pattern trades write latency for durability, adding a DynamoDB round-trip before each Neptune write. During a network partition between Regions, DynamoDB Global Tables prioritize availability over strict consistency, using last-writer-wins reconciliation. Design your mutation schema to be conflict-free if your workload requires strict ordering.
This option is most suitable when your workload has the following characteristics:
- Your application requires near-zero RPO during cross-Region recovery and can accept the additional latency and cost of a DynamoDB WAL layer.
- You need the secondary Region to have full visibility into all writes immediately after promotion.
- You are willing to accept the additional cost of DynamoDB Global Tables and the complexity of idempotent consumers.
Write durability: Decision summary
The following table summarizes the trade-offs between the two write durability approaches during cross-Region recovery. Choose based on your RPO requirements, cost tolerance, and operational complexity budget:
| Factor | Option A: SQS (Accept Delay) | Option B: DynamoDB Global Table WAL |
| Write loss during failover | Writes pending in primary queue until recovery | No write loss. Full log available in both Regions |
| Infrastructure cost | Lower (SQS Standard pricing) | Higher (DynamoDB Global Table replication charges) |
| Operational complexity | Lower (simple queue consumer) | Higher (idempotent consumer, TTL cleanup, ordering) |
| Recovery after failover | Requires reconciliation of pending writes | Consumer picks up from last unprocessed row |
| RPO | 0 (writes are pending, not lost. Visibility delayed until primary Region recovers) | Near-zero |
| Best for | Workloads tolerant of brief pending visibility | Workloads requiring continuous write durability |
Single-Region availability patterns
Even without a multi-Region deployment, several patterns improve Neptune availability within a single Region.
Blue/green deployments
Neptune natively supports blue/green deployments for engine upgrades. Clone the Neptune cluster before the maintenance window, serve traffic from the clone while the original undergoes patching, then switch back or keep using the clone. Blue/green deployments are temporary: you run two clusters only for the duration of the maintenance event, then discard the old one.
Cache-first architecture
Place Amazon ElastiCache in front of Neptune for frequently accessed read paths. Use Amazon ElastiCache and Neptune Streams to keep the cache synchronized. During brief Neptune transitions, the cache continues serving reads. This pattern suits workloads with hot-spot access patterns such as product recommendations or social feeds. Trade-offs include: cache invalidation logic tied to Neptune Streams, a staleness window between graph writes and cache updates, cache warming after evictions, and operational overhead of an additional distributed system. The RTO benefit applies only when Neptune is impaired but ElastiCache remains available.
Pre-computed graph results
For the highest read availability within a single Region, periodically compute graph query results and store them in Amazon DynamoDB or Amazon ElastiCache. Your application reads from the pre-computed store and falls back to Neptune only for fresh computation. Extending this to multiple Regions adds complexity: cross-Region store replication, computation scheduling, and staleness handling. The near-zero RTO assumes only Neptune is impaired. Evaluate application logic changes for cache hydration and fallback routing before adopting.
Feature decoupling (graceful degradation)
Architect your application so that graph-dependent features degrade gracefully when Neptune is temporarily unavailable. For example, a recommendation engine might serve cached recommendations during a brief transition period while real-time personalization pauses. The application remains functional. Only the graph-powered features experience reduced freshness temporarily.
These single-Region patterns complement multi-Region architectures. Blue/green deployments address planned maintenance but do not protect against unplanned failures. Cache-first and pre-computed patterns improve read availability but don’t address write continuity. Choose patterns based on your specific availability gaps rather than adopting all of them.
Monitoring and operational best practices
High availability isn’t only about architecture. It requires ongoing operational discipline across cluster configuration, monitoring, testing, and maintenance scheduling. The following practices help you maintain your target availability over time.
Cluster configuration
Start with at least two reader instances across different Availability Zones. This gives Neptune a promotion target during failover and distributes read traffic. Set failover priority tiers so Neptune promotes your largest reader first. For production workloads, use db.r6g.xlarge or larger so the promoted reader can handle the full write workload.
Amazon CloudWatch monitoring
Start by monitoring GlobalDBProgressLag, which measures replication lag between primary and secondary Regions in milliseconds. Configure an Amazon CloudWatch alarm when this metric exceeds your RPO threshold, as sustained lag indicates the secondary Region is falling behind.
For instance capacity, alarm if CPUUtilization sustains over 70 percent on the writer, as this increases replication lag. Alarm if FreeableMemory drops below 2 GB.
Configure Route 53 health checks from multiple Regions to avoid single-point-of-observation failures, and don’t depend on the impaired Region’s resources to trigger the recovery.
Game Day testing
Conduct quarterly failover exercises to validate your runbooks and automation. Trigger a switchover in a pre-production environment and measure actual RTO. Confirm that application-level validation, including read-after-write tests and queue drain monitoring, behaves as expected. Document the results and refine your procedures.
Maintenance window management
Schedule maintenance windows during your lowest-traffic periods. Coordinate across time zones if you serve a global user base. Combine maintenance events when possible to reduce the total number of transition windows per quarter.
Availability pattern decision framework
The following table summarizes each availability pattern with its expected RTO, RPO, operational complexity, and best-fit scenario. Use it to match patterns to your workload requirements and layer patterns together for defense in depth.
| Pattern | RTO | RPO | Complexity | Best For |
| Multi-AZ baseline, single Region | < 120 seconds | 0 (within Region) | Low | Most workloads. Default starting point |
| Global Database, multi-Region | < 1 second for reads | = replication lag at time of switchover | Medium | Global read latency, disaster recovery reads |
| Global Database + write queue, multi-Region | Minutes (queue drain) | = replication lag at time of failover | Medium-High | Write continuity during maintenance or failover |
| Global Database + DynamoDB WAL, multi-Region | Minutes (WAL consumer) | Near-zero | High | Near-zero RPO requirement across Regions |
| Blue/green deployment, single Region | < 30 seconds (DNS switch) | 0 | Low-Medium | Planned maintenance with minimal transition |
| Cache-first with ElastiCache, single Region | Near-zero (Neptune-only event) | Staleness of cache | Medium | Hot-spot reads, read-heavy workloads |
Start with the Multi-AZ baseline. It provides 99.99 percent availability with minimal configuration. Layer additional patterns as your requirements demand. Add Global Database for global read scalability and Regional disaster recovery. Add a write queue for write continuity during transitions. Add a DynamoDB WAL if you can’t tolerate any stranded writes not yet visible during cross-Region recovery.
Conclusion
Start by identifying your workload’s RTO and RPO targets. Use the decision framework table to find patterns that meet those targets. Then evaluate the complexity column against your team’s operational capacity. Layer patterns from simplest to most complex, adding each only when the preceding level no longer meets your requirements.
Neptune provides 99.99 percent availability through its Multi-AZ SLA. Some workloads require write continuity, faster recovery, or read availability during maintenance events. These patterns extend the baseline capabilities of Neptune so you can select and layer resilience strategies based on your specific RTO, RPO, and operational complexity requirements.
To get started:
- Evaluate your workload’s RTO and RPO requirements using the preceding decision framework table.
- Evaluate your current deployment against the Multi-AZ baseline: two or more readers across Availability Zones.
- Plan for Neptune Global Database if you need cross-Region read scalability or disaster recovery capability.
- Plan a write queue pattern if you need write continuity during Neptune transitions.
- Conduct a Game Day exercise to validate your failover procedures before relying on them in production.
For more information, refer to the Amazon Neptune documentation and the Amazon Neptune Samples repository on GitHub.
If you have questions or feedback about this post, leave a comment in the Comments section.