Overview
Production-Grade DeepVariant on AWS Batch
Run Google's DeepVariant variant caller at scale across large genomic datasets without building or managing your own pipeline infrastructure. This product provides a fully orchestrated, fault-tolerant mechanism to deploy and execute DeepVariant over large collections of files using AWS-native services.
How It Works
The pipeline follows a straightforward data flow designed for production genomic workloads:
- Data Ingestion from S3: Input files (BAM/CRAM and reference genomes) are consumed directly from your Amazon S3 buckets, eliminating the need for manual data transfers or staging.
- Processing via AWS Batch: DeepVariant jobs are submitted and executed through AWS Batch, which handles compute provisioning, job scheduling, and resource management automatically.
- Results Returned to S3: Variant call outputs (VCF/gVCF files) are written back to your designated S3 bucket, ready for downstream analysis or integration with other bioinformatics tools.
Key Capabilities
Scalability Through AWS Batch Process anywhere from a handful of samples to large population-scale cohorts. AWS Batch dynamically scales compute resources up or down based on your workload, so you only use what you need. Submit hundreds of jobs and let the architecture handle parallel execution across your dataset.
Built-In Fault Tolerance Production genomic pipelines cannot afford silent failures. This architecture includes automatic error detection, retry logic, and job-level fault tolerance. If a job fails due to a transient infrastructure issue, the system handles recovery without manual intervention. Errors are captured, logged, and reported so you maintain full visibility into every run.
Error Handling and Reporting Failed jobs are not lost or ignored. The pipeline tracks job status across your entire submission, surfaces errors with actionable detail, and ensures you know exactly which files succeeded and which require attention.
Simple Configuration Get started without extensive pipeline engineering. Configuration is designed to be straightforward, allowing you to point the pipeline at your S3 data and begin processing with minimal setup overhead.
Designed for Genomics Teams on AWS
This product is purpose-built for teams that need to run DeepVariant at production scale on AWS. Rather than spending weeks building custom orchestration around DeepVariant containers, configuring retry logic, managing job queues, and handling edge cases, you can leverage a pre-built architecture that addresses these operational concerns out of the box.
Typical workflows supported include:
- Germline variant calling across whole-genome sequencing (WGS) datasets
- Batch processing of large sample cohorts for research or clinical genomics programs
- Scalable reprocessing of archived sequencing data stored in S3
AWS Integration
This product leverages core AWS services to deliver a cloud-native genomics pipeline:
- Amazon S3 for input and output data storage
- AWS Batch for managed compute orchestration and job scheduling
- Native AWS scaling to grow compute capacity based on workload demands
All processing stays within your AWS account, and data remains in your S3 buckets throughout the pipeline.
Getting Started
After subscribing, configure the pipeline to point at your S3 input data and begin submitting DeepVariant jobs through AWS Batch. The architecture handles compute provisioning, job execution, fault recovery, and result delivery back to S3.
For technical questions, deployment guidance, or to discuss your specific genomic workflow requirements, reach out to the Orangenomix team through the support channels provided on this listing.
Highlights
- Scalable Batch Processing for Large Genomic Datasets: Process large collections of genomic files by leveraging AWS Batch to automatically grow compute resources as workload demands increase. Data is consumed directly from S3, processed through DeepVariant, and results are returned to S3, enabling high-throughput variant calling across many samples without manual infrastructure management.
- Built-In Fault Tolerance and Error Reporting: The architecture is designed to handle failures automatically so that production pipelines remain reliable. Errors are caught, handled, and reported, giving your team visibility into job status without needing to build custom monitoring around the variant-calling workflow.
- Secure, Production-Ready Container with Simple Configuration: The container image is maintained to be vulnerability-free, providing a hardened runtime for sensitive genomic workloads. Configuration is simple, allowing bioinformatics teams to move from subscription to running DeepVariant quickly without extensive setup or DevOps overhead.
Details
Introducing multi-product solutions
You can now purchase comprehensive solutions tailored to use cases and industries.
Features and programs
Financing for AWS Marketplace purchases
Pricing
Vendor refund policy
The product is free, so no refunds are provided for the software neither the infrastructure charges the user may incur.
How can we make this page better?
Legal
Vendor terms and conditions
Content disclaimer
Delivery details
Deploy with Cloudformation script
- Amazon ECS
Container image
Containers are lightweight, portable execution environments that wrap server application software in a filesystem that includes everything it needs to run. Container applications run on supported container runtimes and orchestration services, such as Amazon Elastic Container Service (Amazon ECS) or Amazon Elastic Kubernetes Service (Amazon EKS). Both eliminate the need for you to install and operate your own container orchestration software by managing and scheduling containers on a scalable cluster of virtual machines.
Version release notes
We are glad to share this pipeline with the general public. Enjoy!.
Additional details
Usage instructions
- Identify where your data lives on S3.
- Create and Drop your config file in the same folder.
- Drop a dummy '_READY' marker in the same folder.
- This will trigger processing in AWS Batch.
- Results end up in a folder named after input folder, but with the '_output' suffix.
- Monitor in Lambda.
- For more details check instructions.
Resources
Vendor resources
Support
Vendor support
Orangenomix provides support for the DeepVariant in AWS Batch product. For any questions, issues, or assistance with deployment and troubleshooting, please contact the support team at info@orangenomix.com .
Additional documentation and a getting started tutorial are available in the product's GitHub repository to help you deploy and configure the solution.
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
Similar products
![Variant Effect Predictor (VEP) and the Loss-Of-Function Transcript [...]](https://d1ewbp317vsrbd.cloudfront.net/605b6485-a848-4aac-8128-0a65ca0e6a59.png)
