The Universal Sentence Encoder encodes text into high-dimensional vectors that can be used for text classification, semantic similarity, clustering and other natural language tasks.
The model is trained and optimized for greater-than-word length text, such as sentences, phrases or short paragraphs. It is trained on a variety of data sources and a variety of tasks with the aim of dynamically accommodating a wide variety of natural language understanding tasks. The input is variable-length text and the output is a 512-dimensional vector.
The model supports text in 16 languages (Arabic, Chinese-simplified, Chinese-traditional, English, French, German, Italian, Japanese, Korean, Dutch, Polish, Portuguese, Spanish, Thai, Turkish, Russian). The language of the text input does not need to be specified.
Highlights
Covers 16 languages, showing strong performance on cross-lingual retrieval.
The model is intended to be used for text classification, text clustering, semantic textual similarity retrieval, cross-lingual text retrieval, etc.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
You pay by the hour for each host you run, based on the instance type you choose. Pricing splits into two modes: batch inference and real-time inference. Batch mode covers m4 and m5 instance families across sizes from large to 24xlarge. Real-time mode adds the m5d family alongside m4 and m5. Larger instance sizes carry higher hourly rates because they provide more compute. You select the mode and instance size that fit your workload, and you are billed only for the hours each host runs. No per-token charges apply.
Top-of-mind questions for buyers
What is the difference between batch and real-time inference billing?
Both meter host-hours, but they suit different workloads. Batch mode processes grouped requests in scheduled runs, so you pay only while the batch job runs. Real-time mode keeps a host live to answer requests as they arrive, so charges accrue for the full period the endpoint stays active.
What does one host-hour cover, and am I billed when a host sits idle?
One host-hour is one running instance of the chosen type for one hour. Charges accrue for every hour a host stays active, even without incoming requests. Real-time endpoints run continuously until you stop them. Batch hosts run only during the job. Fully stopped hosts do not accrue software charges.
Are there extra charges based on how many sentences or tokens I process?
No. Billing is a flat hourly rate per running host, with no per-token or per-request fees. Your cost depends only on the instance type you select and the number of hours each host runs. Processing more sentences within those hours does not raise the software charge.
aumlabs.ai
Helpful?
Vendor refund policy
Thank you for purchasing USE Embedding API on AWS Marketplace. We strive to ensure customer satisfaction with our services. If you have issues accessing the service contact us at support@aumlabs.ai
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
An Amazon SageMaker model package is a pre-trained machine learning model ready to use without additional training. Use the model package to create a model on Amazon SageMaker for real-time inference or batch processing. Amazon SageMaker is a fully managed platform for building, training, and deploying machine learning models at scale.
Deploy the model on Amazon SageMaker AI using the following options:
Real-time inference
Deploy the model as an API endpoint for your applications. When you send data to the endpoint, SageMaker processes it and returns results by API response. The endpoint runs continuously until you delete it. You're billed for software and SageMaker infrastructure costs while the endpoint runs. AWS Marketplace models don't support Amazon SageMaker Asynchronous Inference. For more information, see Deploy models for real-time inference .
Batch transform
Deploy the model to process batches of data stored in Amazon Simple Storage Service (Amazon S3). SageMaker runs the job, processes your data, and returns results to Amazon S3. When complete, SageMaker stops the model. You're billed for software and SageMaker infrastructure costs only during the batch job. Duration depends on your model, instance type, and dataset size. AWS Marketplace models don't support Amazon SageMaker Asynchronous Inference. For more information, see Batch transform for inference with Amazon SageMaker AI .
The following table describes supported input data fields for real-time inference and batch transform.
Field name
Description
Constraints
Required
text
An array of text. The text can be in any of the supported languages: Arabic, Chinese-simplified, Chinese-traditional, English, French, German, Italian, Japanese, Korean, Dutch, Polish, Portuguese, Spanish, Thai, Turkish, Russian.
AWS Infrastructure
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
Upbound Universal Crossplane (UXP) is Upbound's official enterprise-grade Crossplane distribution. It's free, open source, and fully conformant with upstream Crossplane.
Parallel Universe is the industry's only SQL server to feature Parallel Query and Parallel Network Query (Distributed Query) with unprecedented speed, utilizing multiple CPU/cores and multiple server hardware while maintaining full compatibility with MySQL and Percona servers.
Parallel Universe is the industry's only SQL server to feature Parallel Query and Parallel Network Query (Distributed Query) with unprecedented speed, utilizing multiple CPU/cores and multiple server hardware while maintaining full compatibility with MySQL and Percona servers.
Infoblox Universal DDI Product Suite: the most advanced, comprehensive portfolio of DNS, DHCP and IP address management solutions for hybrid, multi-cloud environments.
AUM labs transformed our approach to AI deployment. The seamless integration and the ability to run inference APIs in our private cloud was a game changer for us and further helped us keep our data secure.