Overview
all-MiniLM-L6-v2 from the Sentence Transformers project is one of the most downloaded embedding models on HuggingFace, with over 160 million monthly downloads. Fine-tuned on 1 billion sentence pairs from 20+ datasets including MS MARCO, NQ, and StackExchange, its 6-layer architecture compresses to 90MB while delivering strong general-purpose semantic similarity.
At 90MB model weight, it cold-starts in seconds and sustains over 15,000 short sentence embeddings per minute on a single ml.m5.xlarge CPU instance, making it the right choice for high-volume batch ingestion, real-time classification pipelines, and any workload where throughput matters more than top-1 retrieval precision. Note: input text is truncated at 256 tokens.
Deploy as a SageMaker endpoint in your own AWS account. Works with pgvector, Amazon OpenSearch, Pinecone, Weaviate, and any vector store that accepts 384-dimensional float32 vectors. Integrates directly with LangChain HuggingFaceEmbeddings.
Primary use cases: nightly re-embedding of large document corpora, real-time log and event clustering, semantic deduplication across high-volume feeds, chat history summarization for context windows, and cost-sensitive similarity search where extreme throughput is the constraint.
Highlights
- 90MB model, 6-layer architecture -- 15,000+ sentences per minute on ml.m5.xlarge, one of the fastest CPU embedding endpoints available
- 160M+ monthly downloads -- the most widely deployed open-source embedding model, with extensive LangChain, LlamaIndex, and vector store integrations
- Flat $0.07/hr, 384-dim mean-pooled output -- drop-in for pgvector, OpenSearch, Pinecone, and Weaviate, no per-token charges
Details
Introducing multi-product solutions
You can now purchase comprehensive solutions tailored to use cases and industries.
Features and programs
Financing for AWS Marketplace purchases
Pricing
Dimension | Description | Cost/host/hour |
|---|---|---|
ml.m5.xlarge Inference (Real-Time) Recommended | Model inference on the ml.m5.xlarge instance type, real-time mode | $0.07 |
ml.m5.xlarge Inference (Batch) Recommended | Model inference on the ml.m5.xlarge instance type, batch mode | $0.07 |
Vendor refund policy
No refunds.
How can we make this page better?
Legal
Vendor terms and conditions
Content disclaimer
Delivery details
Amazon SageMaker model
An Amazon SageMaker model package is a pre-trained machine learning model ready to use without additional training. Use the model package to create a model on Amazon SageMaker for real-time inference or batch processing. Amazon SageMaker is a fully managed platform for building, training, and deploying machine learning models at scale.
Version release notes
Initial release
Additional details
Inputs
- Summary
sentence-transformers/all-MiniLM-L6-v2 on SageMaker. 384-dimensional mean-pooled embeddings, 90MB model, 160M+ monthly downloads. Embed over 15,000 sentences per minute on ml.m5.xlarge at $0.07/hr.
- Input MIME type
- application/json
Support
Vendor support
Contact support@waltsoft.net for deployment assistance.
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.