AWS Public Sector Blog
1,143 datasets and counting: The Registry of Open Data on AWS hits a milestone
The Amazon Web Services (AWS) Open Data Sponsorship Program makes high-value, cloud-optimized datasets freely available for analysis, helping researchers and developers worldwide access critical data without cost barriers.
The Registry of Open Data on AWS has surpassed a major milestone: over 1,100 datasets are now freely available to anyone. In the last 2 years, the Registry has more than doubled in size, growing from 556 to over 1,100 datasets—a 105% increase.
Anyone can use these datasets to analyze data on AWS, develop new cloud-based techniques and tools, or build communities that benefit from shared data access. Through this program, customers have made over 400 petabytes of high-value, cloud-optimized data publicly available.
All datasets are listed in the Registry of Open Data on AWS. This quarter, 261 new or updated datasets were released.
What are people doing with the Registry of Open Data on AWS?
Organizations are using the Registry of Open Data on AWS in many different ways, including:
- How AWS Open Data democratizes genomic data for the world. Every human deserves access to innovations that could save their life. Yet for decades, groundbreaking genomic research remained locked behind institutional walls, accessible only to well-funded laboratories. AWS addresses these challenges through open data and strategic programs that put powerful technology within reach of organizations supporting underserved populations.
- Transforming rare cancer research with Amazon Quick by integrating biomedical databases for breakthrough discoveries. Hosting the datasets on the AWS Registry of Open Data makes it widely accessible while removing the heavy computational barriers typically required to handle large biological datasets.
- Some of the most important datasets in human genomics are now available for researchers everywhere to use to make impactful health discoveries, without being limited by costly transfer fees. Researchers at the UC Santa Cruz Genomics Institute Computational Genomics Lab and the Broad Institute have deployed a mirror of the National Human Genome Research Institute (NHGRI) AnVIL Data Explorer’s open-access genomic datasets in the AWS Registry of Open Data.
- When Sid was diagnosed with osteosarcoma in November 2022, he pursued maximum diagnostic testing to explore all treatment options. After standard care, he pioneered parallel therapies and scaled this innovative approach for others. His comprehensive dataset hosted on AWS Open Data includes clinical records, molecular data, RNA and DNA sequencing, spatial transcriptomics, residual disease testing (MRD) testing, flow cytometry, imaging, and lab results. This freely shared resource accelerates global cancer research and drives progress toward better patient outcomes.
NASA launches 231 new datasets in the Registry of Open Data
National Aeronautics and Space Administration (NASA) made the decision to include these datasets—spanning decades of Earth observation, climate research, and satellite missions—in the Registry of Open Data on AWS, and it underscores its commitment to transparency, scientific collaboration, and democratizing access to critical environmental and space-related information.
By sharing projects from the 1990s AN and GR series to modern initiatives like NISAR, PACE, and SWOT, NASA ensures that researchers, policymakers, and the public can access high-quality data to advance climate science, disaster response, and resource management. This initiative fosters global partnerships and empowers communities to address challenges such as climate change and biodiversity loss through evidence-based solutions. They now have over 300 datasets within the Registry of Open Data.
NASA joins 30 other new or updated datasets on the Registry of Open Data on AWS in the following categories:
Climate and weather
- EMBER Modeling Files from U.S. Environmental Protection Agency
- IOOS MARACOOS Regional Ocean Modeling System (ROMS) “Doppio” Data Assimilative Reanalysis from National Oceanic Atmosphere Administration (NOAA)
- ECMWF AIFS ENS – dynamical.org Icechunk Zarr from dynamical.org
- East Coast Community Ocean Forecast System (ECCOFS) from NOAA
- Global aboveground biomass (AGB), 100m from CTrees
- Met Office Blended Probabilistic Forecast – Global gridded percentiles from Met Office
- Met Office Blended Probabilistic Forecast – Global gridded probabilities from Met Office
- Met Office Blended Probabilistic Forecast – Global spot percentiles from Met Office
- Met Office Blended Probabilistic Forecast – Global spot probabilities from Met Office
- Met Office Blended Probabilistic Forecast – UK Spot Percentiles from Met Office
- Met Office Blended Probabilistic Forecast – UK Spot Probabilities from Met Office
- Met Office Blended Probabilistic Forecast – UK gridded percentiles from Met Office
- Met Office Blended Probabilistic Forecast – UK gridded probabilities from Met Office
- DWD ICON-EU – dynamical.org Icechunk Zarr from dynamical.org
- NOAA North American Multi-Model Ensemble (NMME) from NOAA
- Met Office UK Marine Observations from Met Office
- NOAA JISAO’s Seasonal Coastal Ocean Prediction of the Ecosystem (J-SCOPE) from NOAA
- NOAA HRRR – dynamical.org Icechunk Zarr from dynamical.org
Geospatial
- Indiana Statewide Leaf-on Digital Aerial Imagery Catalog from Indiana Geographic Information Office
- Geospatial Information Center from Association for Promotion of Infrastructure Geospatial Information Distribution (AIGID)
- Data to Science Catalog from Geospatial Data Science Lab at Purdue University
- 231 datasets from NASA
Life sciences
- Sid Sijbrandij’s osteosarcoma dataset from Rare Cancer Research Foundation
- NCBI SRA Gene Feature RNA-Seq counts from National Institutes of Health (NIH)
- ESM Atlas — Protein Features and Structures from Biohub
- DynaCell from Biohub
- Allen Institute for Neural Dynamics – Extracellular Electrophysiology Hybrid Evaluation Benchmark from Allen Institute
- Intratelencephalic neuron connectivity paper supplemental data from Allen Institute
- BrainSeq – Neurogenomics to Drive Novel Target Discovery for Neuropsychiatric Disorders from Sage Bionetworks
- Multi-Anatomy Post-Surgical Magnetic Resonance Dataset (MAPSMR) from GE Healthcare
- Genoxus Annotation from Genoxus Labs
How can you make your data available?
The AWS Open Data Sponsorship Program covers storage costs for publicly available, high-value, cloud-optimized datasets. We work with data providers who seek to:
- Democratize access to data by making it available for analysis on AWS
- Develop new cloud-based techniques, formats, and tools that lower the cost of working with data
- Encourage the development of communities that benefit from access to shared datasets
Learn how to propose your dataset to the AWS Open Data Sponsorship Program.
Learn more about open data on AWS.
About the author
Kyle Cook
Kyle Cook is the Technical PM for AWS Open Data, driving initiatives to make high-value public datasets freely accessible in the AWS Cloud for researchers, developers, and enterprises worldwide—enabling breakthroughs in fields like climate science, genomics, AI, and machine learning (ML). He collaborates globally with AWS customers and internal teams to democratize data access, launching datasets through the AWS Open Data Registry and streamlining tools for millions of users.
TAGS: announcements, ASDI, AWS Data Exchange, AWS Open Data Sponsorship Program, AWS Public Sector, AWS Public Sector Partners, climate, datasets, geospatial data, life sciences, machine learning, news, open data, Registry of Open Data on AWS, weather
