
Sold by: Zoey Werbin
Open data
|
Deployed on AWS
SoilMicrobeDB is a Kraken2 genome database with extensive representation of high-quality genomes of soil organisms, including uncultured and fungal species.
Overview
SoilMicrobeDB is a Kraken2 genome database with extensive representation of high-quality genomes of soil organisms, including uncultured and fungal species.
Features and programs
Open Data Sponsorship Program
This dataset is part of the Open Data Sponsorship Program, an AWS program that covers the cost of storage for publicly available high-value cloud-optimized datasets.
Pricing
This is a publicly available data set. No subscription is required.
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Legal
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
Delivery details
AWS Data Exchange (ADX)
AWS Data Exchange is a service that helps AWS easily share and manage data entitlements from other organizations at scale.
Open data resources
Available with or without an AWS account.
- How to use
- To access these resources, reference the Amazon Resource Name (ARN) using the AWS Command Line Interface (CLI). Learn more
- Description
- SoilMicrobeDB, a genome reference database for soil shotgun metagenomics
- Resource type
- S3 bucket
- Amazon Resource Name (ARN)
- arn:aws:s3:::kraken2-soil-microbe-database
- AWS region
- us-east-2
- AWS CLI access (No AWS account required)
- aws s3 ls --no-sign-request s3://kraken2-soil-microbe-database/
Resources
Vendor resources
Support
Contact
Managed By
Zoey Werbin
How to cite
SoilMicrobeDB genome database on AWS was accessed on DATE from https://registry.opendata.aws/soil_microbe_db .
License
There are no restrictions on the use of this data.
Similar products

The Alliance of Genome Resources is a consortium that integrates genomic, genetic, and molecular data from leading model organism databases including Drosophila melanogaster, Caenorhabditis elegans, Danio rerio (zebrafish), Mus musculus (mouse), Rattus norvegicus (rat), Saccharomyces cerevisiae (yeast), Xenopus laevis and Xenopus tropicalis (frogs), and human reference data. The Alliance provides comprehensive datasets including gene annotations, disease associations, expression data (bulk and single-cell RNA-Seq), protein and genetic interactions, orthology relationships, variants and alleles, and complete genome sequences with annotations. Data is organized into Alliance-wide integrated datasets and organism-specific collections, supporting comparative genomics, disease modeling, and functional genomics research.

The Genome Ark hosts genomic information for the Vertebrate Genomes Project (VGP) and other related projects. The VGP is an international collaboration that aims to generate complete and near error-free reference genomes for all extant vertebrate species. These genomes will be used to address fundamental questions in biology and disease, to identify species most genetically at risk for extinction, and to preserve genetic information of life.

Database for use with Kraken2 (taxonomic annotation of metagenomic sequencing reads) including all NCBI RefSeq genomes available in release V205
MonkDB is the AI-native database revolutionizing AWS workloads. Natively supporting SQL, NoSQL, Vector, Geospatial, Time Series, Full-Text Search, and Blob data, it queries anything in milliseconds even complex, high-volume, high-velocity data with the simplicity of SQL. Power real-time reconciliations, AI workflows, agentic flows, pattern detection, training, inferencing, advanced analytics and more effortlessly. Designed for any data type, MonkDB on AWS delivers seamless scalability and strict compliance for your most demanding applications. MonkDB Cloud Enterprise is available exclusively under the Enterprise pricing tier. Deploy now and transform your data game!