LLMWhisperer is a document parser that prepares documents for large language models.
It takes PDFs, scans, images, and just about any documents. It then returns text that preserves the layout of the source, so your LLMs receive maximum context.
It works in three parts: text extraction, document handling, and integration.
Text Extraction
LLMWhisperer reads a document and returns text that preserves its original structure, ready to pass to any model.
Preserves multi-column layouts, tables, and key-value pairs in place.
Returns tagged, layout-preserved output rather than a flat block of text.
Reaches 99.9% accurate raw text extraction across supported document types.
Document Handling
LLMWhisperer processes documents across various formats, with no manual pre-processing.
Reads scans, image PDFs, rotated pages, and handwriting as they come.
Extracts across 300+ languages, including non-Latin scripts.
Integration
LLMWhisperer runs as a REST API you call directly, built to drop into existing deployments or pipelines.
Extracts in synchronous or asynchronous mode, with Webhooks that signal when a job finishes.
Ships with Python and Node.js clients to shorten integration.
Processes pages in parallel to reduce manual work and save time.
Common use cases
Document pre-processing
Feeding scanned and image PDFs into LLM workflows
Converting handwritten forms into machine-readable text
Preparing invoices, contracts, and statements for downstream extraction
Building searchable text from scanned archives
Multilingual document processing across regions
Highlights
Layout-preserving mode: keeps multi-column text, tables, and forms in their original structure, so models read a document the way a person would.
300+ language coverage: extracts mixed-language and non-Latin documents, including Arabic, Chinese, Japanese, and Hindi.
Highlight API: returns the location of every extracted line in the source document, so you can highlight where a value came from during human review.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
Pricing is based on the duration and terms of your contract with the vendor, and additional usage. You pay upfront or in installments according to your contract terms with the vendor. This entitles you to a specified quantity of use for the contract duration. Usage-based pricing is in effect for overages or additional usage not covered in the contract. These charges are applied on top of the contract price. If you choose not to renew or replace your contract before the contract end date, access to your entitlements will expire.
Additional AWS infrastructure costs may apply. Use the AWS Pricing Calculator to estimate your infrastructure costs.
This contract gives you an annual commitment to process 300,000 pages of documents into LLM-ready text. You pay one yearly fee for that page volume. Pricing scales by pages, not by users or features. If you exceed the included 300,000 pages, each extra page is billed at the overage rate. Overage applies automatically, so processing continues without interruption. One page equals one physical document page, and multi-page files count each page separately. The two dimensions work together: the annual block covers your baseline volume, and overage handles any usage beyond it.
Top-of-mind questions for buyers
What counts as one page for billing purposes?
One page equals one physical page of a document. A 10-page PDF counts as 10 pages. A single image counts as one page. Multi-page scanned files, such as TIFFs, count each frame as a separate page.
What happens if I process more than 300,000 pages in the year?
You keep processing without interruption. Each page beyond the 300,000 included pages is billed at the overage rate per page. There are no cut-offs or throttling. Overage applies automatically on top of your annual commitment.
Am I charged when a page fails to process?
No. If the system returns an error for a technical reason, such as a timeout, unreadable file, or service issue, that page is not charged. You pay only for pages that process successfully, even if the extracted text is not perfect.
unstract.com
Helpful?
Vendor refund policy
This is a contract with usage-based pricing. Once the usage credits included in the contract are exhausted, you will be billed based on your actual usage.
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
SaaS delivers cloud-based software applications directly to customers over the internet. You can access these applications through a subscription model. You will pay recurring monthly usage fees through your AWS bill, while AWS handles deployment and infrastructure management, ensuring scalability, reliability, and seamless integration with other AWS services.
If you need assistance or encounter any issues during use, please contact us at support@unstract.com.
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.