Deepdub eTTS is a cutting-edge neural text-to-speech model delivering ultra-realistic, human-like voices in 100+ languages and accents. Built for AWS SageMaker JumpStart, it enables developers and enterprises to generate expressive speech with natural prosody, emotion, and clarity, directly within their AWS environment. Easily deployable via SageMaker endpoints, Deepdub eTTS supports both streaming and batch workflows, making it ideal for media localization, conversational AI, eLearning, accessibility, and more. With low-latency inference, fine control over tone and style, and seamless AWS integration, Deepdub eTTS empowers you to create lifelike, engaging audio experiences at scale, without compromising on performance or security.
Deepdub eTTS is a next-generation neural text-to-speech model that produces speech almost indistinguishable from a real human voice. It combines advanced AI voice synthesis technology with a deep understanding of how people speak, resulting in output that is rich in emotion, accurate in pronunciation, and natural in pacing.
Key Capabilities:
Natural prosody and emotion: The model captures subtle vocal inflections, intonation, and timing, creating speech that feels authentic and engaging.
Extensive language and accent range: Supports more than 100+ languages, including regional accents and variations, enabling content to be tailored for audiences worldwide.
Variety of voice styles: Choose from a broad selection of voice styles to match different needs, from professional narrations and conversational tones to dynamic character performances.
Flexible usage: Works effectively for both real-time speech generation and large-scale batch processing.
Enterprise-level reliability: Built for high performance and scalability, while protecting data privacy and security.
Use Cases:
Media localization and dubbing: Translate and voice content while keeping the emotional impact of the original performance.
Interactive voice response (IVR): Deliver clear, engaging, and professional-sounding voices for automated call handling systems.
Conversational AI: Power chatbots, virtual assistants, and automated voice systems with natural, pleasant-sounding voices.
-Agentic AI applications: Provide lifelike voice output for AI agents that perform tasks, interact autonomously, and communicate naturally with users.
eLearning and training: Produce voiceovers that maintain learner attention through clear and expressive delivery.
Accessibility: Create high-quality audio for screen readers, audio books, and other assistive technologies.
Marketing and branding: Develop custom voice assets for advertising, promotions, and interactive experiences.
Deepdub eTTS reduces the time, cost, and complexity of creating realistic voice content while ensuring the highest possible audio quality. Its combination of linguistic accuracy, emotional expression, and scalability makes it a versatile solution for businesses, creators, and developers seeking professional-grade voice generation.
Highlights
Generate ultra realistic speech with natural rhythm, authentic prosody, and expressive emotional range. Deepdub eTTS delivers human like delivery that engages audiences, maintains clarity across different speaking styles, and adapts to various applications from professional narration and character dialogue to interactive AI driven voice experiences.
Support for more than 50 languages and a wide selection of regional accents enables Deepdub eTTS to deliver truly localized experiences. Whether for global media distribution, multilingual customer service, or region specific marketing campaigns, the model ensures cultural and linguistic authenticity that connects with audiences worldwide.
Built for performance, scalability, and flexibility, Deepdub eTTS handles real time streaming, batch processing, IVR systems, and AI driven applications with ease. Enterprise grade security and low latency processing ensure smooth integration into any workflow, allowing seamless deployment for large scale and mission critical voice generation needs.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
You pay by the hour based on the AWS instance your inference runs on. Two options exist, each tied to a specific instance type and processing mode. The ml.g5.2xlarge option runs in batch mode, which processes text-to-speech requests in groups. The ml.g6e.xlarge option runs in real-time mode, which generates audio as requests arrive for low-latency use. Billing accrues per host hour while the instance runs. Your cost scales with how long each instance stays active, so choose the mode that fits your workload.
Top-of-mind questions for buyers
What does one host hour cover for each instance option?
One host hour is one hour that a single instance runs the model. The ml.g5.2xlarge option bills each hour it processes text-to-speech in batch groups. The ml.g6e.xlarge option bills each hour it generates audio in real-time. Both meter running time on one instance, counted hourly.
Which option fits variable workloads versus steady, low-latency demand?
The ml.g5.2xlarge batch option processes requests in groups, which suits scheduled or bulk text-to-speech jobs. The ml.g6e.xlarge real-time option generates audio as requests arrive, which suits conversational or streaming use where fast response matters. Each meters per host hour independently on its own instance.
Am I charged when an instance is idle or stopped?
Charges accrue per host hour while the instance runs. A fully stopped instance does not accrue software charges. An idle but running instance still bills per hour, because billing tracks running time, not the number of requests processed. Underlying AWS infrastructure fees may still apply separately.
docs.deepdub.ai
Helpful?
Vendor refund policy
no refund
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
An Amazon SageMaker model package is a pre-trained machine learning model ready to use without additional training. Use the model package to create a model on Amazon SageMaker for real-time inference or batch processing. Amazon SageMaker is a fully managed platform for building, training, and deploying machine learning models at scale.
Deploy the model on Amazon SageMaker AI using the following options:
Real-time inference
Deploy the model as an API endpoint for your applications. When you send data to the endpoint, SageMaker processes it and returns results by API response. The endpoint runs continuously until you delete it. You're billed for software and SageMaker infrastructure costs while the endpoint runs. AWS Marketplace models don't support Amazon SageMaker Asynchronous Inference. For more information, see Deploy models for real-time inference .
Batch transform
Deploy the model to process batches of data stored in Amazon Simple Storage Service (Amazon S3). SageMaker runs the job, processes your data, and returns results to Amazon S3. When complete, SageMaker stops the model. You're billed for software and SageMaker infrastructure costs only during the batch job. Duration depends on your model, instance type, and dataset size. AWS Marketplace models don't support Amazon SageMaker Asynchronous Inference. For more information, see Batch transform for inference with Amazon SageMaker AI .
Comprehensive Documentation and Developer Tools:
Deepdub API clients receive in-depth documentation and access to developer tools that facilitate easy integration. Our extensive guides include API usage examples, integration instructions, and best practices. The developer portal also provides essential resources like code snippets and SDKs to assist in efficient setup and ongoing management.
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
The Deepdub API integrates our groundbreaking emotive-based Text-to-Speech technology, providing businesses with an efficient tool to create lifelike, emotionally resonant speech for a variety of applications. Designed for enterprise-scale use, this API supports extensive customization options, including accent control and advanced voice modification, ensuring that each audio output is perfectly tailored to meet specific content needs.
Deepdub Voice API powers multilingual AI agents with expressive, emotionally adaptive speech. Featuring licensed Hollywood-grade voices, real-time performance (under ~250ms), and enterprise-grade control for scalable, human-like interactions
Deepdub GO is a cutting-edge virtual AI studio designed to streamline the post-production dubbing process. This platform empowers creators to produce high-quality localized content quickly and efficiently by leveraging proprietary emotion-based text-to-speech technologies and professional voice creation.
Deepdub's managed services provide a comprehensive suite of dubbing and localization solutions tailored to meet the specific needs of studios, content creators and corporates. Leveraging advanced AI technologies and a team of industry experts, these services offer end-to-end management of the localization process to ensure high-quality, culturally relevant content for global audiences.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.