AWS News Blog
Amazon S3 Tables now support all Apache Iceberg V3 data types
|
|
Amazon S3 Tables now support all data types in the Apache Iceberg V3 specification. You can create V3 tables or upgrade existing V2 tables to take advantage of V3 features like deletion vectors, row lineage, and new data types such as variant, nanosecond timestamps, unknown, geometry, and geography.
Apache Iceberg has become the open standard for managing large analytics datasets. It lets you manage petabyte-scale tables with features like schema evolution, hidden partitioning, and time travel queries, while keeping your data in open Parquet files in data lakes on object storage like Amazon S3. Amazon S3 Tables offer storage purpose-built to keep Iceberg tables performant and cost-effective as they grow, with fully managed features like automatic compaction, maintenance, replication, and Intelligent-Tiering.
Teams running analytics on Apache Iceberg V2 tables often hit the same limits as their data grows. A compliance request to delete 50,000 user records from a 2-billion-row table leaves behind positional delete files that slow queries until compaction runs. Semi-structured events land as JSON strings that every query has to parse. Geospatial coordinates and nanosecond-precision timestamps get encoded as strings or integers. Each workaround adds storage cost, query latency, and pipeline code. With V3, Iceberg solves these challenges by offering native support for semi-structured and geospatial data, faster row-level operations, and built-in row lineage for data governance.
Starting today, Amazon S3 Tables support all V3 data types, including variant, nanosecond timestamps, geometry, geography, and unknown, along with deletion vectors and row lineage. You can create new V3 tables or upgrade existing V2 tables in place, and S3 Tables continue to run compaction and maintenance for you.
Apache Iceberg V3
V3 is the latest version of the Iceberg specification. Among its many improvements, V3 introduces capabilities that address the most common pain points in V2. This includes:
Deletion vectors replace V2’s positional delete files with a compact binary format. That 50,000-row compliance delete now writes a single deletion vector file instead of thousands of small deletes, significantly reducing compaction time and delete file overhead.
Row lineage adds _row_id and _last_updated_sequence_number to each record automatically. Your downstream pipelines can query these fields to find changed rows without scanning the full table.
New data types let you store semi-structured, geospatial, and nanosecond-precision data natively instead of encoding it as strings or integers:
- Nanosecond timestamp(tz) for nanosecond-precision timestamps
- Geometry and geography for geospatial data
- Unknown for columns with no known type
Variant data type stores semi-structured data in columnar format. During writes, the engine shreds variant data into hidden columns and collects statistics. At query time, those statistics enable file pruning that significantly reduces I/O compared to parsing JSON strings.
The following sections walk through how to use these V3 capabilities in practice, with examples that show how to create tables, work with the new data types, and manage data at scale.
Getting started
A retail analytics team tracks user behavior across web and mobile apps. Each event has a different structure: page views include URLs and duration, purchases include items and amounts, and searches include query terms and result counts. With V3’s variant type, you store all event shapes in one table without predefined schemas:
CREATE TABLE my_catalog.namespace.clickstream (
event_id bigint,
event_time timestamp,
user_id string,
payload variant
)
USING iceberg
TBLPROPERTIES ('format-version' = '3')
Insert events with different payload shapes without worrying about schema evolution:
INSERT INTO my_catalog.namespace.clickstream VALUES
(1, current_timestamp(), 'user-42',
PARSE_JSON('{"action": "purchase", "amount": 99.99, "items": ["laptop_stand"]}')),
(2, current_timestamp(), 'user-17',
PARSE_JSON('{"action": "page_view", "url": "/products/webcam", "duration_ms": 4200}'));
Now query the variant column directly, without PARSE_JSON at read time. With Amazon EMR Spark, use variant_get:
SELECT
event_id,
user_id,
variant_get(payload, '$.action', 'string') AS action,
variant_get(payload, '$.amount', 'double') AS amount
FROM my_catalog.namespace.clickstream
WHERE variant_get(payload, '$.action', 'string') = 'purchase'
AND variant_get(payload, '$.amount', 'double') > 50.00
To enable deletion vectors for write operations, configure merge-on-read mode:
ALTER TABLE my_catalog.namespace.clickstream
SET TBLPROPERTIES (
'write.delete.mode' = 'merge-on-read',
'write.update.mode' = 'merge-on-read',
'write.merge.mode' = 'merge-on-read'
)
Now when you run a compliance delete, V3 writes a small deletion vector instead of rewriting data files:
DELETE FROM my_catalog.namespace.clickstream
WHERE user_id = 'user-42'
S3 Tables compaction handles these deletion vector files automatically on the next maintenance cycle.
Upgrading from V2
AWS provides backwards compatibility for both versions to minimize disruption during migration to V3. Existing V2 readers continue to work on upgraded tables until you’re ready to fully adopt V3 features. For more details, see the S3 Tables Iceberg V3 documentation.
Upgrade an existing table atomically without rewriting data:
ALTER TABLE my_catalog.namespace.existing_table
SET TBLPROPERTIES ('format-version' = '3')
On the next compaction cycle, S3 Tables remove old V2 delete files. New modifications use deletion vectors automatically. Row lineage fields initialize on the first data modification after the upgrade.
This is a one-way operation. The Apache Iceberg specification does not support downgrading from V3 to V2. Verify that all engines accessing the table support V3 before upgrading.
Using row lineage for incremental pipelines
After your table has V3 data, use row lineage to build efficient incremental pipelines:
SELECT *, _row_id, _last_updated_sequence_number
FROM my_catalog.namespace.clickstream
WHERE _last_updated_sequence_number > 42
This returns only rows modified after sequence number 42. Your downstream jobs can checkpoint this value and process only new changes on each run, instead of scanning the full table.
Compatibility across AWS analytics services
AWS offers the broadest native Apache Iceberg support of any major cloud provider, with Iceberg-compatible services at every layer of the data stack: ingestion, storage, catalog, and analytics. You can store and automatically optimize V3 tables in Amazon S3 Tables, write data with Amazon EMR Spark, integrate and manage data with AWS Glue, and run analytics with Amazon Redshift. To learn more about AWS analytics support for V3, see the Apache Iceberg on AWS prescriptive guidance.
Both S3 Tables and AWS Glue Data Catalog support the Iceberg REST Catalog (IRC) API, enabling interoperability across engines regardless of the catalog endpoint.
Things to know
- S3 Tables compaction fully supports V3 deletion vector files and preserves row lineage metadata.
- The new V3 data types (variant, nanosecond timestamps, geometry, geography, and unknown) require an engine built on Apache Spark 4.0 or later, such as AWS Glue 6.0 or later, or Amazon EMR release 8.1 or later.
- You can create V3 tables from the Amazon S3 console, AWS CLI, or any engine that supports the Iceberg REST Catalog API.
- The new V3 data types are supported only for tables that use the Parquet file format (not ORC or Avro).
- Columns of type variant, geometry, geography, or nanosecond timestamp can’t be included in a table’s sort order for compaction. Tables containing these columns still compact under the sort and Z-order strategies when the sort order uses columns of other types.
Now available
Amazon S3 Tables support for all Apache Iceberg V3 data types is now available in all AWS Regions where S3 Tables are supported. Apache Iceberg V3 support is available at no additional charge; standard S3 Tables pricing applies.
To get started, visit the Amazon S3 Tables documentation or create a table bucket from the Amazon S3 console. If you want to call APIs, search documentation, find regional availability, and check troubleshooting about this feature, try using the AWS MCP Server and plugins with your preferred AI tool. Send feedback to AWS re:Post or through your usual AWS Support contacts.
– Daniel Abib
