AWS Big Data Blog

Category: Analytics

Amazon MSK simplifies configuring custom domain names

Amazon MSK simplifies configuring custom domain names

With Amazon MSK, you can now configure custom domain names for provisioned clusters using a single configuration property that works identically on ZooKeeper and KRaft. Define the domain once and Amazon MSK applies it across every broker, so custom domain names keep working as the cluster scales.

How Zepto powers sub-second search using OpenSearch Service OR2 instances

How Zepto powers sub-second search using OpenSearch Service OR2 instances

Learn how Zepto, India’s fast-growing quick-commerce platform, migrated Amazon OpenSearch Service to OpenSearch Optimized (OR2) instances to scale sub-second product search across hundreds of delivery hubs, achieving over 100% higher indexing throughput and 30% cost savings while serving the same workload on two-thirds the data nodes.

NaranjaX manages multiple Amazon MSK Serverless clusters in different accounts from their IDP using AWS RAM and Route 53

NaranjaX manages multiple Amazon MSK Serverless clusters in different accounts from their IDP using AWS RAM and Route 53

Learn how NaranjaX built a cross-account, many-to-many connectivity model for Amazon MSK Serverless using AWS Resource Access Manager and Amazon Route 53 Resolver, so teams across more than 40 AWS accounts can adopt event-driven architecture from a centralized internal developer platform.

How Autodesk migrated 2.3 billion documents to Amazon OpenSearch Service using Migration Assistant and intelligent routing

This post walks through how Autodesk re-architected a single-index Elasticsearch 7.1.1 domain on Amazon OpenSearch Service into four multi-index OpenSearch Service domains, using Migration Assistant for Amazon OpenSearch Service and a routing layer that directs each query to the shards that hold the data for that query.

Trace cascading decision failures with a blame graph on Amazon OpenSearch

Trace cascading decision failures with a blame graph on Amazon OpenSearch Service

When multiple AI agents collaborate on a decision and get it wrong, standard logs can’t tell you which agent caused it. This post shows how to build a blame graph on Amazon OpenSearch Service and Amazon Bedrock that measures influence between agents and walks backward from a failed decision to find the root cause.

How AppFolio transformed its data streaming architecture with Amazon MSK Express brokers

How AppFolio transformed its data streaming architecture with Amazon MSK Express brokers

Learn how AppFolio transformed its data streaming architecture by adopting Amazon MSK Express brokers, replacing hours-long rebalances and manual storage planning with a platform that scales automatically across workload-isolated clusters.

How GPU acceleration builds billion-scale vector indexes on Amazon OpenSearch Service

How GPU acceleration builds billion-scale vector indexes on Amazon OpenSearch Service

GPU-accelerated vector (k-NN) indexing on Amazon OpenSearch Service and OpenSearch Serverless lets you build billion-scale vector indexes in hours instead of days. This post goes deep on the decoupled GPU architecture, the CAGRA-to-HNSW conversion, a one-billion-vector benchmark, and operational best practices for production.

Centralized CloudTrail monitoring across 100+ AWS accounts

Centralized CloudTrail monitoring across 100+ AWS accounts

Learn how to build a centralized AWS CloudTrail monitoring solution on Amazon OpenSearch Service, with Terraform managing the full stack. It handles 200 GB/day of logs across 100+ accounts, provides automated threat detection, delivers on-demand SOC 2, PCI DSS, and HIPAA compliance reporting, and gives four teams isolated access.

How a team at Epic Games tuned Amazon OpenSearch Service for Fortnite analytics

How a team at Epic Games tuned Amazon OpenSearch Service for Fortnite analytics

Learn how a team at Epic Games tuned their Amazon OpenSearch Service cluster for Fortnite analytics. This post details how Epic Games partnered with AWS to right-size instances, rebalance shards, optimize index mappings, and upgrade engine and Graviton versions, improving query latency and throughput while reducing costs.

AI-powered cost optimization agent for Amazon Kinesis Data Streams

AI-powered cost optimization agent for Amazon Kinesis Data Streams

Learn how to deploy an open-source, AI-powered agent built on Amazon Bedrock that automatically analyzes every Amazon Kinesis Data Streams stream in your account, compares costs across the three capacity modes, and recommends the optimal mode to help you save over 60% on streaming costs on a schedule you choose.