AWS Big Data Blog
How Autodesk migrated 2.3 billion documents to Amazon OpenSearch Service using Migration Assistant and intelligent routing
This post walks through how Autodesk re-architected a single-index Elasticsearch 7.1.1 domain on Amazon OpenSearch Service into four multi-index OpenSearch Service domains, using Migration Assistant for Amazon OpenSearch Service and a routing layer that directs each query to the shards that hold the data for that query.
Trace cascading decision failures with a blame graph on Amazon OpenSearch Service
When multiple AI agents collaborate on a decision and get it wrong, standard logs can’t tell you which agent caused it. This post shows how to build a blame graph on Amazon OpenSearch Service and Amazon Bedrock that measures influence between agents and walks backward from a failed decision to find the root cause.
How AppFolio transformed its data streaming architecture with Amazon MSK Express brokers
Learn how AppFolio transformed its data streaming architecture by adopting Amazon MSK Express brokers, replacing hours-long rebalances and manual storage planning with a platform that scales automatically across workload-isolated clusters.
How GPU acceleration builds billion-scale vector indexes on Amazon OpenSearch Service
GPU-accelerated vector (k-NN) indexing on Amazon OpenSearch Service and OpenSearch Serverless lets you build billion-scale vector indexes in hours instead of days. This post goes deep on the decoupled GPU architecture, the CAGRA-to-HNSW conversion, a one-billion-vector benchmark, and operational best practices for production.
Centralized CloudTrail monitoring across 100+ AWS accounts
Learn how to build a centralized AWS CloudTrail monitoring solution on Amazon OpenSearch Service, with Terraform managing the full stack. It handles 200 GB/day of logs across 100+ accounts, provides automated threat detection, delivers on-demand SOC 2, PCI DSS, and HIPAA compliance reporting, and gives four teams isolated access.
How a team at Epic Games tuned Amazon OpenSearch Service for Fortnite analytics
Learn how a team at Epic Games tuned their Amazon OpenSearch Service cluster for Fortnite analytics. This post details how Epic Games partnered with AWS to right-size instances, rebalance shards, optimize index mappings, and upgrade engine and Graviton versions, improving query latency and throughput while reducing costs.
AI-powered cost optimization agent for Amazon Kinesis Data Streams
Learn how to deploy an open-source, AI-powered agent built on Amazon Bedrock that automatically analyzes every Amazon Kinesis Data Streams stream in your account, compares costs across the three capacity modes, and recommends the optimal mode to help you save over 60% on streaming costs on a schedule you choose.
Amazon OpenSearch Service extends version lifecycle support timelines
In November 2024, we announced Standard and Extended Support dates for legacy Elasticsearch and OpenSearch versions on Amazon OpenSearch Service. We are extending security and operating system patch coverage for these versions by 12 months, through November 7, 2027, and announcing support dates for additional Elasticsearch and OpenSearch versions.
Scaling fine-grained access control for enterprise lakehouse using SageMaker Unified Studio and AWS Lake Formation
As enterprise lakehouses grow to thousands of tables across business domains and regions, fine-grained access control becomes a governance bottleneck. This post shows how to combine AWS IAM Identity Center, AWS Lake Formation tag-based access control, and trusted identity propagation in Amazon SageMaker Unified Studio for automated, auditable, least-privilege access.
Event-driven pipeline orchestration with Amazon MWAA and Airflow 3.0
Data engineering teams running Apache Airflow across multiple AWS accounts have no built-in way to coordinate workflows between separate Amazon MWAA environments. With Airflow 3.0 on Amazon MWAA, you can use asset-based scheduling and Asset Watchers with Amazon SQS to build event-driven, cross-account orchestration that replaces polling with near real-time triggers.









