HyperAccel delivers a fast and efficient LLM inference system on AWS F2 Instances. Our custom AI chip, the LLM Processing Unit (LPU), is the world's first hardware dedicated to end-to-end inference of large language models like Llama, HyperCLOVA X, and EXAONE.
HyperAccel provides a high-performance, low-latency inference platform for large language models (LLMs) on AWS F2 Instances. Our solution is powered by the LLM Processing Unit (LPU) - the world's first hardware engine purpose-built for full LLM inference.
Unlike GPU-based systems, LPUs are optimized for real-time performance with significantly lower power consumption and cost. This AMI includes a pre-configured software stack and FPGA bitstream that is fully integrated with the vLLM inference engine. It enables efficient deployment of popular open-source and commercial LLMs such as Meta Llama, NAVER HyperCLOVA X, and LG EXAONE. It provides high-throughput and low-latency inference through features, making it well-suited for running multi-billion parameter models.
With this instance, customers can easily deploy their own LLM-powered chatbot server using HyperAccel's FPGA acceleration, without requiring any prior hardware expertise or additional infrastructure setup.
Highlights
FPGA-based LPU architecture delivers high-performance LLM inference with lower latency and power consumption compared to GPUs.
Pre-configured AMI with a vLLM-integrated software stack and FPGA bitstream enables instant chatbot server deployment without hardware expertise.
Supports rapid deployment of the latest LLMs like Llama, HyperCLOVA X, and EXAONE through flexible Hugging Face integration.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
Try this product free for 14 days according to the free trial terms set by the vendor. Usage-based pricing is in effect for usage beyond the free trial terms. Your free trial gets automatically converted to a paid subscription when the trial ends, but may be canceled any time before that.
You pay by the hour for one instance type, the f2.6xlarge. Billing is usage-based, so charges accrue only for the hours the instance runs. There are no tiers or add-ons to choose from. This single dimension runs the LLM chatbot software on AWS F2 hardware. Your total cost scales directly with how long you keep the instance active. Start or stop it as needed to control spend.
Top-of-mind questions for buyers
What hardware does the f2.6xlarge hourly rate cover?
The rate covers one f2.6xlarge instance on AWS F2 hardware. This runs the LLM chatbot software on specialized processing units built for large language model inference. You get one running instance per hour billed. Each hour the instance stays active counts as one billable unit.
Am I charged when the f2.6xlarge instance is stopped or paused?
Software charges accrue only while the instance runs. Stopping the instance halts the hourly software billing. Underlying AWS storage or reserved resources may still incur separate AWS fees, but the software meter tracks running time only. Restart it whenever you need it.
What kinds of models can this instance run?
The software supports transformer-based large language models and multi-modal models. It works with common inference and model frameworks. This lets you run chatbot and inference workloads without changing your existing setup significantly. The same hourly rate applies regardless of which supported model you run.
docs.hyperaccel.ai+1
Helpful?
Vendor refund policy
If you believe you were charged in error or encountered technical issues that prevented product use, please contact our support team at support@hyperaccel.ai within 7 days of the charge date. Each refund request will be reviewed on a case-by-case basis. You may cancel your subscription at any time via the AWS Marketplace Console.
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
An AMI is a virtual image that provides the information required to launch an instance. Amazon EC2 (Elastic Compute Cloud) instances are virtual servers on which you can run your applications and workloads, offering varying combinations of CPU, memory, storage, and networking resources. You can launch as many instances from as many different AMIs as you need.
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
REQUIRES PRIVATE OFFER
To purchase LiteLLM Enterprise Self-Hosted, please reach out to sales@berri.ai for a Private Offer.
LiteLLM is an OpenAI compatible Proxy Server (LLM Gateway) to call 2,000+ LLM APIs using the OpenAI format Bedrock, Huggingface, VertexAI, TogetherAI, Azure OpenAI, OpenAI, etc. Get started with Opensource LiteLLM here: https://github.com/BerriAI/litellm (40,000+ Github Stars)
Transform complex documents into structured, schema-compliant JSON using the top-ranked self-hosted OCR model. Built for enterprise document automation, JSL Vision OCR Structured LLM extracts data from PDFs, forms, tables, and scanned documents while keeping sensitive information within your AWS environment.
Model with advanced clinical decision-support system built for structured medical reasoning rather than simple knowledge retrieval.
It analyzes symptoms, diagnostics, and longitudinal patient histories to guide complex diagnostic and treatment decisions in line with established clinical guidelines.
With multimodal vision capabilities, it can also interpret medical images and visual documents alongside text, broadening the range of clinical inputs.
This multimodal medical model delivers advanced clinical reasoning across both text and medical imagery in a highly efficient footprint. Trained on diverse medical thinking and patient-centered datasets, it understands complex clinical narratives while accurately interpreting X-rays, MRIs, CT scans, pathology slides, charts, diagrams, and structured medical records.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.