Description

The 1000 Genomes Project is an international collaboration which has established the most detailed catalogue of human genetic variation, including SNPs, structural variants, and their haplotype context. The final phase of the project sequenced more than 2500 individuals from 26 different populations around the world and produced an integrated set of phased haplotypes with more than 80 million variants for these individuals.

Update Frequency

Not updated

License

Data from the 1000 Genomes Project is now available without embargo, following the final publication from the project. Use of the data should be cited in the usual way, with current details available at http://www.internationalgenome.org/faq/how-do-i-cite-1000-genomes-project.

Documentation

https://github.com/awslabs/open-data-docs/tree/main/docs/1000genomes

Managed By

National Institutes of Health

See all datasets managed by National Institutes of Health.

Contact

http://www.internationalgenome.org/contact

How to Cite

1000 Genomes was accessed on DATE from https://registry.opendata.aws/1000-genomes.

Usage Examples

Publications

Exploratory data analysis of genomic datasets using ADAM and Mango with Apache Spark on Amazon EMR by Alyssa Marrow

Resources on AWS

Description

http://www.internationalgenome.org/formats

Resource type

S3 Bucket

Amazon Resource Name (ARN)

arn:aws:s3:::1000genomes

AWS Region

us-east-1

AWS CLI Access (No AWS account required)

aws s3 ls --no-sign-request s3://1000genomes/