AWS Big Data Blog

Category: Advanced (300)

Query Amazon S3 Tables from Amazon EMR Trino using the Iceberg REST endpoint

Query Amazon S3 Tables from Amazon EMR Trino using the Iceberg REST endpoint

Learn how to query Amazon S3 Tables from Trino on Amazon EMR using the Apache Iceberg REST catalog endpoint. This post shows how to deploy the integration with AWS CloudFormation, configure the Trino catalog, and run SQL to create, query, and manage Apache Iceberg tables.

Building medallion architecture with Iceberg materialized views in Amazon SageMaker

Building medallion architecture with Iceberg materialized views in Amazon SageMaker

With Apache Iceberg materialized views in Amazon SageMaker, you can build a Bronze, Silver, and Gold medallion architecture as three SQL statements. This declarative approach folds transformation, orchestration, and incremental processing into per-layer definitions, with no ETL jobs, orchestrators, or change-data-capture code to maintain.

Build a dynamic streaming data lake with Apache Iceberg and Apache Flink

Build a dynamic streaming data lake with Apache Iceberg and Apache Flink

Learn how to build a dynamic streaming data lake on Amazon Managed Service for Apache Flink that adapts to new event types and schema changes without stopping the pipeline, using Apache Iceberg’s Dynamic Iceberg Sink for per-record table routing and automatic schema evolution.

Observing and evaluating production agents using OpenSearch Agent Health

Observing and evaluating production agents using OpenSearch Agent Health

Learn how to observe and evaluate production AI agents by combining an agent running on AWS with OpenSearch Agent Health. This post walks through deploying an agent and its observability pipeline to AWS, then using Agent Health to explore traces and run evaluations that measure and improve agent quality over time.

Accelerate Apache Spark debugging on Amazon EMR with AWS DevOps Agent

Accelerate Apache Spark debugging on Amazon EMR with AWS DevOps Agent

Extend AWS DevOps Agent to investigate Apache Spark failures on Amazon EMR. This post shows how to register the Apache Spark Troubleshooting Agent for Amazon EMR as a custom MCP capability provider over AWS PrivateLink, so a single agent chat session diagnoses a failing Spark job from an Amazon CloudWatch alarm to a line-numbered root cause.

Integrate Amazon Redshift and IAM Identity Center with enhanced VPC routing

Integrate Amazon Redshift and IAM Identity Center with enhanced VPC routing

Amazon Redshift now supports AWS IAM Identity Center authentication on clusters and workgroups that use enhanced VPC routing. Create two interface VPC endpoints to give your users single sign-on with their corporate credentials while keeping all authentication traffic on the AWS private network.

Build a real-time event pipeline with Spark Real-Time Mode on AWS Glue 6.0

Build a real-time event pipeline with Spark Real-Time Mode on AWS Glue 6.0

With AWS Glue 6.0, you can build real-time, near-real-time, and batch data pipelines on a single platform. Using a financial market-risk example, learn how to flag high-risk trades with sub-second latency using Spark Real-Time Mode, store heterogeneous pricing vectors with Apache Iceberg v3 Variant columns, and run batch analytics with Arrow-native UDFs.

How Picnic configured multiple OAuth providers for Amazon MQ

How Picnic configured multiple OAuth providers for Amazon MQ

Picnic runs RabbitMQ as the messaging backbone for hundreds of microservices on Amazon MQ for RabbitMQ. This post shows how to configure one broker to trust multiple OAuth 2.0 identity providers, Keycloak for operators and AWS IAM for services, so you can eliminate static credentials while maintaining separate identity paths for people and workloads.

Build with geospatial and variant types in Iceberg v3 on AWS Glue 6.0

Build with geospatial and variant types in Iceberg v3 on AWS Glue 6.0

AWS Glue 6.0 with Apache Spark 4.1 adds support for Apache Iceberg v3: native geospatial types, nanosecond-precision timestamps, the VARIANT type, and DEFAULT column values. This post builds a connected vehicle fleet telemetry pipeline that uses all four in a single Iceberg v3 table, from ingestion through spatial, nanosecond, and variant queries.

Enable cross-cloud analytics with Amazon S3 Tables and Google BigQuery, Part 1: IAM-based access control

Enable cross-cloud analytics with Amazon S3 Tables and Google BigQuery, Part 1: IAM-based access control

Your Google BigQuery users need to query data that lives in Amazon S3 Tables on AWS without copying it across clouds. This post shows how to connect BigQuery to Amazon S3 Tables through the AWS Glue Iceberg REST Catalog using IAM-based access control, so you keep one governed dataset and query it live from BigQuery.