AWS Glue
Discover, prepare, and integrate all your data at any scale
Why AWS Glue?
AWS Glue is a serverless data integration service that makes it simpler, faster, and cheaper to discover, prepare, and combine data for analytics, AI and agents. Connect to more than 100 diverse data sources, manage data in a centralized catalog, and visually create, run, and monitor pipelines to load data into your data lakes, warehouses, and lakehouses — with no infrastructure to manage. AWS was recognized as a Leader in the 2025 Gartner Magic Quadrant for Data Integration Tools.
Integrate your data with AWS Glue in Amazon SageMaker
With AWS Glue in the next generation of Amazon SageMaker, you can manage and build your workloads in one place with cost-effective, serverless, and scalable data integration.
Benefits
AWS Glue provides all the capabilities needed for data integration, so you can start gaining insights quickly. The fully managed, serverless toolkit includes built-in ETL, automated schema discovery, a centralized Data Catalog, and cross-service integration.
Spark Declarative Pipelines eliminates hundreds of lines of orchestration boilerplate. Simply declare your transformations and the engine handles execution order, dependencies, error handling, and checkpointing automatically. Combined with generative AI-assisted code generation and AI-powered Spark job modernization, you build and ship pipelines faster than ever.
AWS Glue automatically scales even the most demanding data processing jobs from gigabytes to petabytes with no infrastructure to manage. No servers to provision, configure, or maintain — you pay only for the resources you use, billed by the second with no upfront costs.
Arrow-native Python UDFs and UDTFs eliminate serialization overhead, delivering significant performance improvements for complex PySpark transformations with zero code changes required.
Connect to hundreds of diverse data sources, seamlessly integrating data from databases, data warehouses, data lakes, and SaaS applications. Build ETL jobs using open-source Apache Spark, Python, and Scala. Your code runs anywhere Spark is supported, ensuring flexibility with no vendor lock-in. Integrate seamlessly with Amazon Athena, Amazon EMR, Amazon Redshift, and Amazon SageMaker.
Built on Apache Spark 4.1.1 with Python 3.12 and Scala 2.13, plus simultaneous support for Iceberg 1.10.0, Hudi 1.1.1, and Delta Lake 4.0.0. Full Apache Iceberg v3 support includes geometry/geography types for spatial analytics, nanosecond-precision timestamps, and Unknown type handling for resilient schema evolution.
AWS Glue serves as the data preparation and integration backbone for machine learning and AI agents — ensuring they receive clean, timely, and well-structured data at scale. Combine Glue's ETL capabilities with Amazon SageMaker for end-to-end ML workflows.
Iceberg v3 VARIANT Shredding in Glue delivers faster reads and easier schema management for semi-structured data like JSON, logs, and event streams — the exact formats that feed ML feature pipelines and AI training datasets.
Use Cases
Data lake ingestion
Ingest data from 100+ sources into your data lake using the latest open data standards. Glue's Iceberg v3 support handles schema evolution gracefully with Unknown type handling, nanosecond timestamps for IoT and financial data, and geometry types for spatial workloads.
ETL pipeline modernization
Replace brittle, manually orchestrated pipelines with Spark Declarative Pipelines — define what you want, not how to run it. Use generative AI to modernize legacy Spark jobs and migrate to Glue with minimal effort.
Real-time analytics
Detect fraud at the point of transaction, personalize experiences in real time, and alert on IoT anomalies within the same Glue environment you use for batch. Real-Time Mode delivers single-digit millisecond latency without separate streaming infrastructure.
ML data preparation
Build scalable data prep pipelines that feed clean, structured data to ML models and AI agents. VARIANT Shredding processes semi-structured training data faster — no flattening or duplicate copies needed.
Data quality and governance
Centralize metadata with the Glue Data Catalog, automate schema discovery with Crawlers, and enforce governance through native Lake Formation integration. Schema Registry supports AVRO, JSON, and Protocol Buffers for streaming data.
Semi-structured data processing
Store and query JSON, logs, event streams, and IoT telemetry natively with Iceberg VARIANT — eliminating the need for separate semi-structured databases. Organizations that maintained duplicate copies across systems can consolidate onto a single Iceberg-based architecture.
What's New
Stifel Financial
Built a modern data platform using AWS Glue and open data standards with an event-driven architecture for domain data products while centralizing metadata for discovery and sharing.
BMW
Processes 10 TB of data daily from 1.2 million vehicles, creating voice-activated in-vehicle assistants and deriving real-time insights from vehicle and customer telemetry.
"To stay innovative, we are focusing on creating new digital and connected experiences and driving change in our value chain toward improving both efficiency and effectiveness by enabling data-driven decisions."
— Kai Demtröder, Group VP of Data Transformation, AI, Data and DevOps, BMW
Stifel Financial
Built a modern data platform using AWS Glue and open data standards with an event-driven architecture for domain data products while centralizing metadata for discovery and sharing.
Did you find what you were looking for today?
Let us know so we can improve the quality of the content on our pages