TextIn xParse is a next-gen document intelligence product powered by LLMs, turning unstructured complex documents into structured, queryable data assets for relational and vector databases. It optimizes ETL pipelines and empowers high-quality RAG Q&A, serving 1000+ leading enterprises worldwide for diverse document processing needs. This version is passport parse version.
Rebuilt from the ground up with LLMs and beyond traditional OCR, TextIn xParse excels in processing all types of complex documents, breaking down arbitrary layouts into semantically complete paragraphs and restoring reading order for large model adaptability. Boasting industry-leading table recognition, it resolves merged cells, multi-page tables and borderless tables with ease, and integrates seamlessly with image processing to handle watermarked and curved documents. As an intelligent ETL solution, it enables zero-sample key information extraction, cross-document retrieval and intelligent document classification, solving large model pain points like unstable output and length truncation. TextIn xParse generates high-quality Chunks with semantic relationship labeling, coordinate and chapter information, boosting RAG Q&A accuracy and search efficiency, and supports one-click import to mainstream RAG frameworks including RagFlow, Dify and Coze. It builds a solid document infrastructure for enterprise scenarios like Knowledge Q&A, Agent Enablement, Data Entry and Data Cleaning, automating unstructured data processing, reducing manual workload and maximizing data asset value. Trusted by global leading enterprises, it delivers efficient, accurate document processing capabilities for mission-critical business scenarios.
Highlights
LLM-Powered Document Parsing: Beyond OCR, realizes accurate structured conversion of complex unstructured documents and adapts perfectly to large model application needs
High-Quality RAG Empowerment: Generates semantic-optimized Chunks to improve Q&A accuracy and search efficiency, supporting one-click access to mainstream RAG frameworks
Intelligent ETL Pipeline: Achieves zero-sample extraction, cross-document retrieval and intelligent classification, maximizing enterprise unstructured data asset value
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
You pay for this product by the hour on a single dimension tied to the t3a.large instance type. Billing follows usage, so charges accrue for each hour the software runs on that instance. There are no tiers or separate add-ons to choose. Your cost scales directly with how many hours you keep the instance running. This structure suits document parsing workloads where you start and stop processing as needed.
Top-of-mind questions for buyers
What resources do I get when running on the t3a.large instance?
You run the software on a t3a.large EC2 instance, a general-purpose machine with 2 vCPUs and 8 GiB of memory. This instance handles the document parsing engine. You pay the software rate for each hour this instance runs, alongside standard AWS infrastructure charges.
Am I charged when the t3a.large instance is stopped or paused?
Software charges accrue per hour the instance runs. When you stop the instance, the hourly software charge stops. Stopped instances may still incur AWS storage fees for attached volumes, but the software license meters running time only. This lets you control cost by starting and stopping as workloads require.
What document workloads does this hourly instance support?
The software parses documents into structured, model-ready output. It handles over 16 formats, including PDFs, scans, tables, presentations, emails, and screenshots. It recognizes tables with merged cells or multi-page spans. You can connect to cloud object storage, file transfer sources, network storage, and local file systems.
www.textin.ai
Helpful?
Vendor refund policy
nonrefundable
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
An AMI is a virtual image that provides the information required to launch an instance. Amazon EC2 (Elastic Compute Cloud) instances are virtual servers on which you can run your applications and workloads, offering varying combinations of CPU, memory, storage, and networking resources. You can launch as many instances from as many different AMIs as you need.
Version release notes
TextIn Document Parser AMI - Initial release.
Supports PDF, DOCX, TXT parsing and outputs clean, structured JSON data for AI systems.
Additional details
Usage instructions
TextIn Document Parser AMI Usage Instructions
Overview
This AMI provides a pre-installed, native document parsing service for Ubuntu 20.04 LTS. It converts PDF, DOCX, and TXT files into structured, AI-ready JSON data optimized for LLMs, Agents, and RAG systems.
Base OS: Ubuntu 20.04 LTS
Default username: ubuntu
SSH port: 22
Service port: 30006
Connect to Your EC2 Instance
From the Amazon EC2 Console, obtain the public IP or DNS of your instance.
Connect using SSH with your AWS key pair:
plaintext
ssh -i "your-key-pair.pem" ubuntu@<instance-public-ip>
Type yes to confirm the host key on first connection.
Verify the Service Status
The TextIn service starts automatically on boot.
Check service status:
plaintext
sudo systemctl status textin-parser.service
If inactive, start and enable it:
plaintext
sudo systemctl start textin-parser.service
sudo systemctl enable textin-parser.service
Verify health:
plaintext
curl http://localhost:30006/health
A healthy response returns:
plaintext
{"status":"healthy"}
Use the Document Parsing API
Submit documents for parsing:
plaintext
curl -X POST -F "file=@/path/to/your/document.pdf" http://<instance-public-ip>:30006/parse
Supported formats: PDF, DOCX, TXT
Output: Clean structured JSON
License Activation (Optional)
For enterprise features:
Run the license setup script:
plaintext
./1-install_licserver.sh
Send the generated machine fingerprint (seed.txt) to simon_liu@intsig.net to obtain a license.
Place your license file in /home/ubuntu/licFile/.
Apply the license:
plaintext
./2-apply_license.sh
Security Best Practices
Open only ports 22 (SSH) and 30006 (API) in your security group.
Restrict SSH access to trusted IP ranges.
Do not expose the service to 0.0.0.0/0 in production.
Troubleshooting
Service not running: Check logs with journalctl.
API unreachable: Verify port 30006 is open in the security group.
License issues: Confirm seed and license file match.
Last updated: March 2026
Support: For license and service support: simon_liu@intsig.net
Support
Vendor support
Contact email: sheng_song@intsig.net URL: https://www.textin.ai/contact Support time: 8 hours *5 workingdays Buyers can get professional and all-round technical support and after-sales service for TextIn products.
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
iText Suite BYOL provides broad functionality for manipulating and processing PDFs using the acclaimed iText Core PDF library, plus additional features enabled by iText add-ons. Automate your document workflows using a scalable and cost-effective RESTful API.
Turn imagination into art. Powered by the latest technology, our AI creates art and images based on simple text instructions.
This AI Art model can produce diverse painting styles, including oil, watercolor, modern, abstract, and more.
Our technology can also simulate Van Gogh, Monet, Picasso, and famous painters, or can attempt to create art in the style of famous paintings.
Please inquire about photorealistic styles.
Common applications include creating graphics for merchandise, book art, album covers, fan art, and simply helping people see their imaginations manifested in more tangible form.
Deepgram is the enterprise Voice AI platform for building and scaling real time voice applications on AWS. This product listing contains multiple versions of the nova-3 model which can each transcribe a set of languages. See version details for more information.
You will be billed $0.0077/min as described by https://deepgram.com/pricing. Private pricing available upon request.
Our APIs for Nova Speech to Text (STT) are natively available in the new SageMaker Bi-Directional Streaming API. Additional native touchpoints with Amazon Bedrock, Lex, and Amazon Connect make it simple to compose full voice experiences with the cloud services your teams already trust.
Build software resilience from a partner you can trust with application security as a service. Achieve all the advantages of security testing, vulnerability management, tailored expertise, and support without the need for additional infrastructure or resources.
Deepgram is the enterprise Voice AI platform for building and scaling real time voice applications on AWS. This product listing contains multiple versions of the aura-2 model which can each speak a set of languages and voices. See version details for more information. Deepgram charges are billed per request as described by https://deepgram.com/pricing
Our APIs for Nova Speech to Text (STT) are natively available in the new SageMaker Bi-Directional Streaming API. Additional native touchpoints with Amazon Bedrock, Lex, and Amazon Connect make it simple to compose full voice experiences with the cloud services your teams already trust.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.