NOTE: This deployment requires a specific SageMaker Inference AMI selection (al2-ami-sagemaker-inference-gpu-3-1).
Please use the example notebook provided at https://github.com/NVIDIA/digital-biology-examples/blob/main/examples/nims/msa-search/nim-msa-search-v2-1-0_aws_marketplace.ipynb for deploying the endpoint.
Generates a multiple sequence alignment from a query sequence and a protein sequence database search.
The MSA search NIM is powered by GPU MMSeqs2. GPU MMSeqs2 is a GPU-accelerated toolkit for protein database search and Multiple Sequence Alignment (MSA). While not a deep learning model, MMSeqs2 does require large protein databases for sequence similarity search.
The MSA Search NIM enables researchers and commercial entities in the Drug Discovery, Life Sciences, and Digital Biology fields to rapidly generate multiple sequence alignments (MSA). The output MSA can be used in downstream protein structure prediction and evolutionary analysis applications.
Highlights
The MSA Search NIM is based on GPU MMSeqs2, a highly-sensitive, highly-performant, GPU-accelerated package for a variety of sequence-alignment tasks. The NIM accuracy should match that of the open-source version of MMSeqs2 in most settings.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
You pay by the host hour based on the AWS instance type you run for protein sequence alignment. One option, ml.g5.12xlarge, runs in batch mode for processing groups of sequences together. The other eight options run in real-time mode for on-demand requests. These span the g6e, p4d, p4de, p5, p5e, and p5en instance families. Larger instance sizes carry more GPU and memory capacity, so cost scales with the instance you choose. You are billed only for the hours each instance runs.
Top-of-mind questions for buyers
What does one host hour cover for billing?
One host hour is one hour that a single instance of the chosen type runs the inference workload. You are billed per running instance-hour, not per sequence searched or per request. Cost accrues while the instance is active regardless of how many alignments you process during that hour.
How does the batch mode option differ from the real-time options in how work is processed?
The ml.g5.12xlarge batch option processes groups of sequences together in scheduled runs. The eight real-time options handle requests on demand as they arrive. Both bill by host hour on their instance type. Choose batch for bulk processing and real-time for interactive, request-driven alignment.
Am I charged when an instance is stopped or idle?
Software charges apply per host hour while the instance runs. A stopped instance does not accrue host-hour charges. Underlying AWS infrastructure fees, such as storage, may still apply separately depending on your setup. The software meter counts running time only.
docs.nvidia.com
Helpful?
Vendor refund policy
None
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
An Amazon SageMaker model package is a pre-trained machine learning model ready to use without additional training. Use the model package to create a model on Amazon SageMaker for real-time inference or batch processing. Amazon SageMaker is a fully managed platform for building, training, and deploying machine learning models at scale.
Deploy the model on Amazon SageMaker AI using the following options:
Real-time inference
Deploy the model as an API endpoint for your applications. When you send data to the endpoint, SageMaker processes it and returns results by API response. The endpoint runs continuously until you delete it. You're billed for software and SageMaker infrastructure costs while the endpoint runs. AWS Marketplace models don't support Amazon SageMaker Asynchronous Inference. For more information, see Deploy models for real-time inference .
Batch transform
Deploy the model to process batches of data stored in Amazon Simple Storage Service (Amazon S3). SageMaker runs the job, processes your data, and returns results to Amazon S3. When complete, SageMaker stops the model. You're billed for software and SageMaker infrastructure costs only during the batch job. Duration depends on your model, instance type, and dataset size. AWS Marketplace models don't support Amazon SageMaker Asynchronous Inference. For more information, see Batch transform for inference with Amazon SageMaker AI .
The model accepts JSON requests with parameters on /invocations and /ping APIs that can be used to control the generated text. See examples and field descriptions below.
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
Benchling helps scientists accelerate R&D with our R&D Cloud, making it easy to centralize scientific data, collaborate across teams, and access insights.
Benchling helps scientists accelerate R&D with our R&D Cloud, making it easy to centralize scientific data, collaborate across teams, and access insights.
The Steinegger Lab Dataset comprises biological databases and resources critical for protein sequence and structure analysis, developed to support ColabFold, MMseqs2, and Foldseek/Foldcomp—three high-performance computational tools widely used in bioinformatics.
The MMseqs2 dataset serves as the backbone for our fast structure prediction tool, ColabFold, and includes UniRef30, BFD, and the ColabFold environmental databases.
These datasets are specifically designed for the rapid generation of multiple sequence alignments (MSAs), which are essential for high-accuracy structure prediction.
Beyond MSA generation, these resources allow for fast taxonomy annotations and functional annotation, supporting a wide range of bioinformatics applications.
The Foldseek dataset includes preprocessed databases such as the AlphaFold Database (AFDB), PDB, SwissProt, and CATH, specifically designed for protein structure similarity searches.
These datasets encompass the majority of both experimental[...]
Places is a Mastercard dataset that provides a comprehensive view of all merchant locations accepting Mastercard credit cards in both the physical and digital world. Mastercard Places data is captured from aggregated and anonymized transaction data and matched to best-in-class third party location data listings. Then, Mastercard appends unique metadata, such as Cash Back or “in business” flags, which are derived from the transaction data.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.