
Sold by: OpenFold Consortium
Open data
|
Deployed on AWS
This dataset contains MSAs and predicted structures used to train OpenFold3 preview, an open-source, all-atom ligand, RNA and protein structure prediction software. This includes -
- PDB - 245k structures and alignments from the RCSB Protein Data Bank - https://www.rcsb.org/
- Long monomer distillation set - ~13 million long (sequence length >= 200 amino acids) monomers from the MGNIFY database - https://www.ebi.ac.uk/metagenomics/.
- Short monomer distillation set - 400k short (sequence length < 200 amino acid) monomers from the MGNIFY database - https://www.ebi.ac.uk/metagenomics/.
- Disordered set - AF2-predicted structures for unresolved segments missing from the PDB
- RNA - OF3p2-predicted RNA monomer structures generated from a clustered version of RFAM (current version)
For the distillation sets MSAs were generated using the AF3 protocol, and were used to predict structures with AlphaFold2, more details can be found in our whitepaper - https://portal.openfold.om[...]
Overview
This dataset contains MSAs and predicted structures used to train OpenFold3 preview, an open-source, all-atom ligand, RNA and protein structure prediction software. This includes -
- PDB - 245k structures and alignments from the RCSB Protein Data Bank - https://www.rcsb.org/
- Long monomer distillation set - ~13 million long (sequence length >= 200 amino acids) monomers from the MGNIFY database - https://www.ebi.ac.uk/metagenomics/ .
- Short monomer distillation set - 400k short (sequence length < 200 amino acid) monomers from the MGNIFY database - https://www.ebi.ac.uk/metagenomics/ .
- Disordered set - AF2-predicted structures for unresolved segments missing from the PDB
- RNA - OF3p2-predicted RNA monomer structures generated from a clustered version of RFAM (current version) For the distillation sets MSAs were generated using the AF3 protocol, and were used to predict structures with AlphaFold2, more details can be found in our whitepaper - https://portal.openfold.omsf.io/reports/of3p2_technical_report.pdf For a full description and an interactive data explorer, please visit https://portal.openfold.omsf.io/datasets
Features and programs
Open Data Sponsorship Program
This dataset is part of the Open Data Sponsorship Program, an AWS program that covers the cost of storage for publicly available high-value cloud-optimized datasets.
Pricing
This is a publicly available data set. No subscription is required.
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Legal
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
Delivery details
AWS Data Exchange (ADX)
AWS Data Exchange is a service that helps AWS easily share and manage data entitlements from other organizations at scale.
Open data resources
Available with or without an AWS account.
- How to use
- To access these resources, reference the Amazon Resource Name (ARN) using the AWS Command Line Interface (CLI). Learn more
- Description
- A repository of MSAs and 3D protein structural coordinates used to train OpenFold3.
- Resource type
- S3 bucket
- Amazon Resource Name (ARN)
- arn:aws:s3:::openfold3-data
- AWS region
- us-west-2
- AWS CLI access (No AWS account required)
- aws s3 ls --no-sign-request s3://openfold3-data/
Resources
Vendor resources
Support
Managed By
OpenFold Consortium
How to cite
OpenFold3 Training Data was accessed on DATE from https://registry.opendata.aws/openfold3 .
License
Similar products

Folding@home is a massively distributed computing project that uses biomolecular simulations to investigate the molecular origins of disease and accelerate the discovery of new therapies. Run by the Folding@home Consortium, a worldwide network of research laboratories focusing on a variety of different diseases, Folding@home seeks to address problems in human health on a scale that is infeasible by another other means, sharing the results of these large-scale studies with the research community through peer-reviewed publications and publicly shared datasets. During the COVID-19 epidemic, Folding@home focused its resources on understanding the vulnerabilities in SARS-CoV-2,[...]

Multiple sequence alignments (MSAs) for 140,000 unique Protein Data Bank (PDB) chains and 16,000,000 UniClust30 clusters. Template hits are also provided for the PDB chains and 270,000 UniClust30 clusters chosen for maximal diversity and MSA depth. MSAs were generated with HHBlits (-n3) and JackHMMER against MGnify, BFD, UniRef90, and UniClust30 while templates were identified from PDB70 with HHSearch, all according to procedures outlined in the supplement to the AlphaFold 2 Nature paper, Jumper et al. 2021. We expect the database to be broadly useful to structural biologists training or validating deep learning models for protein structure prediction and related tasks.

Community-sourced repository of coral reef image classification training data, including continually updated confirmed annotations from MERMAID

This is the data used to train the Boltz-1 model. It contains the following datasets:
Our pre-processed version of the Protein Data Bank
Our pre-processed version of the multiple sequence alignment data for each protein chain
The raw multiple sequence alginment data.
A pre-computed symmetry file for symmetry correction during training