Overview
BGE Small EN v1.5 from BAAI is a lightweight English embedding model that scores 62.17 on the MTEB benchmark -- within 3% of the Large variant at one-tenth the model size. It produces 384-dimensional L2-normalized vectors from inputs up to 512 tokens.
Deploy it as a SageMaker endpoint in your own AWS account for bulk document ingestion, nightly re-embedding pipelines, or any workload where throughput matters more than the last point of retrieval precision. At $0.07/hr on ml.m5.xlarge you can embed over 10,000 short passages per minute.
Drop-in compatible with the same vector stores as the Large variant: pgvector, Amazon OpenSearch, Pinecone, Weaviate, Qdrant, and Chroma. No code changes required when scaling up to BGE Base or Large -- same 384-dim input, same L2-normalized output contract.
Primary use cases: bulk document ingestion before a one-time migration, high-frequency log or telemetry clustering, cost-sensitive semantic deduplication at scale, and RAG pipelines where you trade a small precision delta for 3x throughput.
Highlights
- MTEB score 62.17 at $0.07/hr -- embed 10,000+ passages per minute on ml.m5.xlarge, no per-token charges
- 384-dimensional L2-normalized vectors, 512-token input -- drop-in for pgvector, OpenSearch, Pinecone, and Weaviate
- Zero code changes to upgrade to BGE Base or Large -- same API, same vector dimensions in the S/M/L ladder
Details
Introducing multi-product solutions
You can now purchase comprehensive solutions tailored to use cases and industries.
Features and programs
Financing for AWS Marketplace purchases
Pricing
Dimension | Description | Cost/host/hour |
|---|---|---|
ml.m5.xlarge Inference (Real-Time) Recommended | Model inference on the ml.m5.xlarge instance type, real-time mode | $0.07 |
ml.m5.xlarge Inference (Batch) Recommended | Model inference on the ml.m5.xlarge instance type, batch mode | $0.07 |
Vendor refund policy
No refunds.
How can we make this page better?
Legal
Vendor terms and conditions
Content disclaimer
Delivery details
Amazon SageMaker model
An Amazon SageMaker model package is a pre-trained machine learning model ready to use without additional training. Use the model package to create a model on Amazon SageMaker for real-time inference or batch processing. Amazon SageMaker is a fully managed platform for building, training, and deploying machine learning models at scale.
Version release notes
Initial release
Additional details
Inputs
- Summary
BAAI BGE Small EN v1.5 on SageMaker. 384-dimensional embeddings, MTEB score 62.17, at $0.07/hr -- the budget tier for high-volume RAG ingestion and semantic search without per-token billing.
- Input MIME type
- application/json
Support
Vendor support
Contact support@waltsoft.net for deployment assistance.
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
Similar products
