Hugging Face Neuron Deep Learning AMI (DLAMI) makes it easy to use Amazon EC2 Inferentia & Trainium instances for training and inference of Hugging Face Transformers and Diffusers models, offering up to 50% cost-to-train savings over comparable GPU-based DLAMIs.
Hugging Face Neuron Deep Learning AMI (DLAMI) makes it easy to use Amazon EC2 Inferentia & Trainium instances for efficient training and inference of Hugging Face Transformers and Diffusers models.
With the Hugging Face Neuron DLAMI, scale your Transformers and Diffusion workloads quickly on Amazon EC2 while reducing your costs, with up to 50% cost-to-train savings over comparable GPU-based DLAMIs.
This DLAMI is the officially supported, and recommended solution by Hugging Face, to run training and inference on Trainium and Inferentia EC2 instances, and supports most Hugging Face use cases, including:
Fine-tuning and pre-training Transformers models like BERT, GPT, or T5
Running inference with Transformers models like BERT, GPT, or T5
Fine-tuning and deploying Diffusers models like Stable Diffusion
This DLAMI is provided at no additional charge to Amazon EC2 users.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
This software is free to use. You pay only for the AWS compute instance you run it on, billed per hour. The dimensions map to two hardware families. Training-optimized instances (trn1 and trn2 series) support model fine-tuning workloads. Inference-optimized instances (inf2 series) support model deployment. Within each family, larger instance sizes add more accelerators, memory, and bandwidth, which raises the hourly rate. You pick the instance that fits your workload and scale up or down by choosing a different size. There is no commitment term or fixed quantity.
Top-of-mind questions for buyers
What hardware does an instance dimension map to for billing?
Each dimension maps to one AWS accelerator EC2 instance you run for an hour. Training instances (trn1, trn1n, trn2) carry Neuron cores for fine-tuning. Inference instances (inf2) carry accelerators for model deployment. Larger sizes in each family add accelerators, memory, and bandwidth, which raises the hourly rate.
Am I charged when an instance is stopped or paused?
The software itself is free, so no software charge accrues. You pay only for the AWS compute instance while it runs, billed per hour. When you stop an instance, hourly compute charges stop. Attached storage may still incur separate AWS fees. Terminate instances you no longer need to avoid charges.
Should I pick a training or an inference instance for my workload?
Choose a trn1, trn1n, or trn2 instance to fine-tune or train models. Choose an inf2 instance to deploy trained models for inference. Both families work with the same tooling for loading models on AWS accelerators. Match the instance to the task, then scale by selecting a different size.
www.philschmid.de+1
Helpful?
Vendor refund policy
no refunds
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
An AMI is a virtual image that provides the information required to launch an instance. Amazon EC2 (Elastic Compute Cloud) instances are virtual servers on which you can run your applications and workloads, offering varying combinations of CPU, memory, storage, and networking resources. You can launch as many instances from as many different AMIs as you need.
Version release notes
v0.4.4: improved vLLM perf with on-device-sampling disable, fix speculation algo, PEFT update for GRPO
Launch instance on either trn1, inf2 or trn2 instance type.
Connect to the instance using the "EC2 Instance Connect" button in the AWS console.
User Name is/should be ubuntu
Test if neuron devices are accessible by running neuron-ls.
Test if Optimum Neuron library is installed by following the commands bellow:
python -c 'import torch_neuronx;import transformers;import accelerate;import optimum.neuron'
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
This product has charges associated with it for pre-configuration by Intuz Inc. This AMI comes with Hugging Face Transformers, Jupyter Notebook, and popular NLP models like GPT, BERT, and T5-ready to use for machine learning and natural language processing tasks.