AWS Big Data Blog
Category: Analytics
How United Airlines uses Amazon Redshift and AWS Glue Data Catalog federation to query Databricks-managed data
Learn how United Airlines uses AWS Glue Data Catalog federation to query Databricks Unity Catalog data directly from Amazon Redshift Serverless without duplicating data, using resource links and AWS Lake Formation for governance.
Scale down Kinesis Data Streams on-demand capacity with ODA warm throughput
Amazon Kinesis Data Streams now supports scaling down ingest capacity for on-demand Advantage streams with warm throughput. Learn how the scale-down works, how to monitor stream behavior with Amazon CloudWatch, and best practices for releasing excess capacity after transient traffic bursts.
How to migrate from Amazon CloudSearch to Amazon OpenSearch Serverless
Learn how to migrate an Amazon CloudSearch domain to Amazon OpenSearch Serverless: assess your configuration, create a collection with explicit index mappings, convert your documents and queries to the OpenSearch query DSL, configure security, load data with Amazon OpenSearch Ingestion, and validate before cutover.
Accelerating Spark queries with Iceberg materialized views
Accelerate slow, repetitive Apache Spark analytical queries on Apache Iceberg tables without rewriting any SQL. This post shows how automatic query rewrite in Amazon EMR and AWS Glue uses Iceberg materialized views in the AWS Glue Data Catalog to transparently substitute matching query plans, and how to design materialized views for the best speedup.
Build declarative ETL pipelines with AWS Glue 6.0
AWS Glue 6.0 introduces Spark Declarative Pipelines. In this post, you build a single declarative AWS Glue 6.0 job that turns raw order records into validated, aggregated tables through a bronze, silver, and gold sequence, without writing any orchestration logic.
How Sony LIV built real-time video streaming analytics with AWS
Real-time data analytics is transforming how streaming platforms serve their audiences. Learn how Sony LIV built a comprehensive, real-time streaming analytics solution on AWS using Amazon Kinesis Data Streams, Amazon EMR, and Apache Iceberg.
Network connectivity patterns for the next generation of Amazon OpenSearch Serverless
The next generation of Amazon OpenSearch Serverless uses standard AWS PrivateLink endpoints on the on.aws domain. This post shows nine connectivity patterns for private access, from a single VPC to multiple VPCs, cross-account, on-premises, and cross-Region, with the DNS resolution and data path for each.
How Moovit achieved 33% cost optimization through architectural modernization
Learn how Moovit modernized its data platform with a multi-engine lakehouse architecture: offloading heavy aggregation workloads from Amazon Redshift to Amazon EMR with Spark SQL, isolating workloads with Amazon Redshift Serverless, and cutting overall data pipeline cost by 33%.
Query Amazon S3 Tables from Amazon EMR Trino using the Iceberg REST endpoint
Learn how to query Amazon S3 Tables from Trino on Amazon EMR using the Apache Iceberg REST catalog endpoint. This post shows how to deploy the integration with AWS CloudFormation, configure the Trino catalog, and run SQL to create, query, and manage Apache Iceberg tables.
Build a dynamic streaming data lake with Apache Iceberg and Apache Flink
Learn how to build a dynamic streaming data lake on Amazon Managed Service for Apache Flink that adapts to new event types and schema changes without stopping the pipeline, using Apache Iceberg’s Dynamic Iceberg Sink for per-record table routing and automatic schema evolution.









