AWS Big Data Blog

Category: AWS Glue

Cost-effective ETL with DuckDB and Amazon S3 Tables on AWS Glue

Cost-effective ETL with DuckDB and Amazon S3 Tables on AWS Glue

Learn how to pair DuckDB with AWS Glue 6.0 to run SQL-centric ETL on a single worker, reading Parquet from Amazon S3 and writing Apache Iceberg tables to Amazon S3 Tables. This post walks through a complete, runnable example and compares measured cost and runtime against an equivalent Apache Spark job on the same Glue runtime.

Accelerating Spark queries with Iceberg materialized views

Accelerating Spark queries with Iceberg materialized views

Accelerate slow, repetitive Apache Spark analytical queries on Apache Iceberg tables without rewriting any SQL. This post shows how automatic query rewrite in Amazon EMR and AWS Glue uses Iceberg materialized views in the AWS Glue Data Catalog to transparently substitute matching query plans, and how to design materialized views for the best speedup.

Build a real-time event pipeline with Spark Real-Time Mode on AWS Glue 6.0

Build a real-time event pipeline with Spark Real-Time Mode on AWS Glue 6.0

With AWS Glue 6.0, you can build real-time, near-real-time, and batch data pipelines on a single platform. Using a financial market-risk example, learn how to flag high-risk trades with sub-second latency using Spark Real-Time Mode, store heterogeneous pricing vectors with Apache Iceberg v3 Variant columns, and run batch analytics with Arrow-native UDFs.

Build with geospatial and variant types in Iceberg v3 on AWS Glue 6.0

Build with geospatial and variant types in Iceberg v3 on AWS Glue 6.0

AWS Glue 6.0 with Apache Spark 4.1 adds support for Apache Iceberg v3: native geospatial types, nanosecond-precision timestamps, the VARIANT type, and DEFAULT column values. This post builds a connected vehicle fleet telemetry pipeline that uses all four in a single Iceberg v3 table, from ingestion through spatial, nanosecond, and variant queries.

Introducing AWS Glue 6.0 for faster and more cost-effective data integration

Introducing AWS Glue 6.0 for faster and more cost-effective data integration

AWS Glue 6.0 is now available, lowering AWS Glue pricing by 30%, adding an AWS optimized build of Apache Spark 4.1, and introducing Apache Iceberg V3 capabilities suitable for enterprise adoption. This post covers the key capabilities and performance benefits, with code examples to help you get started.

Automate creating AWS Glue Data Catalog views with AWS SDK for data mesh use case

This post shows you how to use the Catalog objects API CreateTable() to programmatically create ATHENA and SPARK dialects using cross-account IAM definer roles, and how to add the ATHENA dialect programmatically for the views that were created earlier with only SPARK dialect.

Multi-cloud lakehouse architecture on AWS for Agentic AI, Part 1: Architecture and best practices

Multi-cloud lakehouse architecture on AWS for Agentic AI, Part 1: Architecture and best practices

This post focuses on explaining the architecture approach to build the open lakehouse architecture on AWS, unifying the metadata catalog across providers for the AI agents to access. In addition, it highlights the architecture trade-offs and best practices.