Deploy a production-ready, self-hosted speech-to-text server on your container infrastructure. Convert audio into accurate text via a web UI and REST API. Runs efficiently on standard CPU instances entirely within your AWS account, making it suitable for cost-sensitive workloads, internal services, and private audio processing.
This product provides a fully self-hosted speech-to-text server packaged as a container image for AWS.
The server converts audio files into text and exposes both a web interface and a simple HTTP REST API for programmatic integration. It runs efficiently on standard CPU instances, delivering predictable performance and cost without requiring GPU hardware. All processing occurs within the customer's AWS environment, ensuring full control over data, networking, and security.
The container is designed for production use and supports configurable deployment options, including server ports and SSL settings, to match your workload requirements. The service is suitable for batch transcription, internal applications, and environments where external SaaS speech APIs are not permitted. Common use cases include call center transcription, media and meeting transcription, compliance-sensitive audio processing, and offline or private speech analytics pipelines. HTTPS can also be enabled using standard AWS patterns such as an Application Load Balancer with TLS termination, allowing the solution to integrate cleanly into existing AWS architectures.
Highlights
Self-hosted speech-to-text server running entirely within your AWS account, with no external data dependencies
Cost-effective transcription that runs on standard CPU instances without requiring GPU hardware
Production-ready container image with configurable deployment, Web and REST API access, and AWS-native security integration
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
Try this product free for 5 days according to the free trial terms set by the vendor. Usage-based pricing is in effect for usage beyond the free trial terms. Your free trial gets automatically converted to a paid subscription when the trial ends, but may be canceled any time before that.
Pricing is based on actual usage, with charges varying according to how much you consume. Subscriptions have no end date and may be canceled any time. Alternatively, you can pay upfront for a contract, which typically covers your anticipated usage for the contract duration. Any usage beyond contract will incur additional usage-based costs.
Additional AWS infrastructure costs may apply. Use the AWS Pricing Calculator to estimate your infrastructure costs.
This container is billed by usage, using a single dimension: Container Hours. You pay for each hour the container runs, with no upfront commitment or tiers. Costs scale directly with how long you keep the container active. The container delivers an HTTP API for transcribing audio, running on CPU. You control your total spend by managing how many hours you run the container.
Top-of-mind questions for buyers
What does one Container Hour cover, and how is it counted?
One Container Hour covers each hour the container runs the transcription server. Billing tracks running time, not the number of audio files or minutes transcribed. Partial hours follow AWS metering rules. Cost accrues only while the container is active.
Am I charged when the container is stopped or idle?
You are charged per hour the container runs. A stopped container does not accrue Container Hours. Idle running time still counts, since billing tracks active runtime rather than transcription volume. To limit spend, stop the container when it is not in use.
Does the hourly charge depend on how much audio I transcribe?
No. The single Container Hours dimension bills only for runtime. Transcribing many files in one hour costs the same as transcribing one. The server runs on CPU and exposes an HTTP API for audio uploads. Your bill scales with hours, not workload volume.
www.sigmodata.com
Helpful?
Vendor refund policy
Refunds are handled in accordance with AWS Marketplace refund policies. Buyers may request a refund within 48 hours of initial purchase or launch if the product does not function as described on supported EC2 instance types. To request a refund or support, contact support@sigmodata.com
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
Containers are lightweight, portable execution environments that wrap server application software in a filesystem that includes everything it needs to run. Container applications run on supported container runtimes and orchestration services, such as Amazon Elastic Container Service (Amazon ECS) or Amazon Elastic Kubernetes Service (Amazon EKS). Both eliminate the need for you to install and operate your own container orchestration software by managing and scheduling containers on a scalable cluster of virtual machines.
Version release notes
Create container image
Additional details
Usage instructions
This container runs a CPU-optimized Speech-to-Text server with:
Web UI for uploading audio and viewing transcriptions
Support description: Support is provided via email and web-based support channels.
Buyers can expect assistance with initial deployment, configuration, and troubleshooting of the AMI and speech-to-text service.
Response times and support scope may vary based on the customers subscription or private offer agreement.
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
CommX operational collaboration suite provides easy to use user friendly voice, video and chat collaboration. Stream and manage real-time media on your browser with low-latency technology. Elevate your multimedia monitoring capabilities and never miss a single detail. Record, stream and investigate audio, video, and telephony sources. CommX Discover enables you to take the next big step in telephony calls and video data management. You can generate maximum insights from all input sources in real time via a simplified, web-based user interface. Keep streaming, watch, and hear recordings, start investigating and be creative with AI-driven investigation tools: Video indexing, STT (speech to text), Speaker verification, Face recognition, and Video summarization.
Deepgram is the enterprise Voice AI platform for building and scaling real time voice applications on AWS. This product listing contains multiple versions of the nova-3 model which can each transcribe a set of languages. See version details for more information.
You will be billed $0.0077/min as described by https://deepgram.com/pricing. Private pricing available upon request.
Our APIs for Nova Speech to Text (STT) are natively available in the new SageMaker Bi-Directional Streaming API. Additional native touchpoints with Amazon Bedrock, Lex, and Amazon Connect make it simple to compose full voice experiences with the cloud services your teams already trust.
Deepgram is the enterprise Voice AI platform for building and scaling real time voice applications on AWS. This product listing contains multiple versions of the nova-3 model which can each transcribe a set of languages. See version details for more information.
You will be billed $0.0092/min as described by https://deepgram.com/pricing. Private pricing available upon request.
Our APIs for Nova Speech to Text (STT) are natively available in the new SageMaker Bi-Directional Streaming API. Additional native touchpoints with Amazon Bedrock, Lex, and Amazon Connect make it simple to compose full voice experiences with the cloud services your teams already trust.
Deepgram is the enterprise Voice AI platform for building and scaling real time voice applications on AWS. This product listing contains multiple versions of the flux model which can each transcribe a set of languages. See version details for more information.
You will be billed $0.0078/min as described by https://deepgram.com/pricing.
Our APIs for Nova Speech to Text (STT) are natively available in the new SageMaker Bi-Directional Streaming API. Additional native touchpoints with Amazon Bedrock, Lex, and Amazon Connect make it simple to compose full voice experiences with the cloud services your teams already trust.
Deepgram is the enterprise Voice AI platform for building and scaling real time voice applications on AWS. This product listing contains multiple versions of the flux model which can each transcribe a set of languages. See version details for more information.
You will be billed $0.0077/min as described by https://deepgram.com/pricing. Private pricing available upon request.
Our APIs for Nova Speech to Text (STT) are natively available in the new SageMaker Bi-Directional Streaming API. Additional native touchpoints with Amazon Bedrock, Lex, and Amazon Connect make it simple to compose full voice experiences with the cloud services your teams already trust.
Deepgram is the enterprise Voice AI platform for building and scaling real time voice applications on AWS. This product listing contains multiple versions of the nova-3 model which can each transcribe a set of languages. See version details for more information. Deepgram charges are billed per request as described by https://deepgram.com/pricing
Our APIs for Nova Speech to Text (STT) are natively available in the new SageMaker Bi-Directional Streaming API. Additional native touchpoints with Amazon Bedrock, Lex, and Amazon Connect make it simple to compose full voice experiences with the cloud services your teams already trust.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.