AWS Storage Blog
Category: Analytics
How Precisely transforms user experience with AI agents using Amazon S3 Vectors
At Precisely, the team is reimagining the user experience for its Data Integrity Suite by adding a conversational interface powered by AI agents to the traditional UI. With this enhancement, users can interact with the platform more naturally and intuitively (asking questions, making requests, and exploring data assets through dialogue) while still benefiting from the […]
Run Spark 31% faster and optimize compute costs with Amazon S3 Express One Zone on Amazon EMR
As your Spark datasets grow, storage latency often becomes the constraint, impeding application performance. Query runtimes stretch, and the bottleneck shifts from compute to how fast each node can read from Amazon Simple Storage Service (Amazon S3). We benchmarked this directly on Amazon EMR with TPC-DS at 3 TB scale. On an 8-node Graviton4 cluster […]
Optimize your self-managed PostgreSQL data warehouse with Amazon FSx for OpenZFS
A data warehouse is the analytical backbone of a modern enterprise, consolidating data from disparate sources into a single, authoritative view that enables complex queries, trend analysis, and confident decision-making. In financial services, this means sharper regulatory reporting, faster fraud detection, and deeper customer understanding. The operational reality is demanding. Enterprises juggle multiple source databases […]
Track healthcare data lineage in real time with Amazon S3 annotations
Organizations that store sensitive data need to answer a simple question: where did this data come from, and what happened to it? In healthcare, this isn’t optional. Regulations like HIPAA, GDPR, and SOX require organizations to produce this information on demand. When an auditor asks, for example, “Show me every transformation that touched patient record […]
Analyze Amazon S3 annotations at scale with materialized views
Customers managing large volumes of objects in Amazon Simple Storage Service (Amazon S3) often need to attach rich business context like compliance classifications, processing lineage, AI-generated labels, and more. Until now, this context lived in external databases or sidecar files that were stored as separate objects, which created complexity to manage and keep it up […]
Accelerate Amazon S3 Replication with automated S3 Batch Operations parallelization
As data volumes grow, organizations must move large datasets between storage locations to meet compliance requirements, optimize performance, and help meet data sovereignty requirements. However, migrating petabytes of data presents significant challenges: lengthy transfer times, complex coordination of parallel operations, data integrity verification, and substantial engineering overhead. Automation reduces resource consumption and operational complexity during […]
How Vanderbilt University scales digital archive discovery with Amazon S3 Metadata
Managing massive digital collections is hard. When you’re preserving decades of historical content and adding new materials daily, making that content discoverable matters more than the storage itself. Vanderbilt University Library discovered this firsthand while managing their extensive digital archives, including the renowned Vanderbilt Television News Archive (VTNA). Amazon S3 Metadata accelerates data discovery by […]
Query Amazon S3 access logs instantly with CloudWatch and S3 Tables
Knowing who accessed your data, when, and how is the foundation for security investigations, compliance audits, cost attribution, and performance troubleshooting. Detailed access logs capture every request: who made it, which resource was accessed, and what response was returned. In practice, though, they arrive as semi-structured records spread across different locations. Turning them into actionable […]
Amazon S3 audit logging, Part 3: Analyzing S3 Metadata journal tables for object lifecycle tracking
This is Part 3 of our three-part series on Amazon S3 audit logging. In Part 1, we covered server access logs for HTTP-level requests and performance analysis. In Part 2, we covered S3 data events in AWS CloudTrail for identity-focused security investigations. As data volumes grow and storage costs become a significant line item, organizations […]
Amazon S3 audit logging, Part 1: Analyzing server access logs with Amazon Athena for performance insights
Organizations storing sensitive data must maintain complete visibility into how it’s accessed, by whom, and what changes occur over time. Regulatory frameworks demand detailed audit trails, security teams need rapid answers during investigations, and finance teams require granular cost attribution. Yet as data grows from terabytes to petabytes, the scale that makes centralized storage attractive […]


