Overview
BGE Base EN v1.5 from BAAI scores 63.55 on the MTEB benchmark -- just 0.68 points below the Large variant at one-third the memory footprint and roughly 3x the throughput on the same ml.m5.xlarge instance. It produces 768-dimensional L2-normalized vectors from inputs up to 512 tokens.
This is the production workhorse for RAG pipelines that need strong retrieval precision without the full cost of the Large model. At $0.08/hr it sits between BGE Small ($0.07/hr) and BGE Large ($0.10/hr), giving you a clear cost-quality ladder for different environments: Small for bulk ingestion, Base for production serving, Large for highest-precision search.
Drop-in compatible with all major vector stores: pgvector on Aurora/RDS, Amazon OpenSearch, Pinecone, Weaviate, Qdrant, Chroma, and Milvus. Integrates directly with LangChain (HuggingFaceBgeEmbeddings), LlamaIndex, and Haystack using the same configuration as the Large variant, switching only the model ID.
Primary use cases: production document retrieval for RAG, customer support knowledge base search, enterprise semantic search over internal wikis and documentation, product catalog similarity, and duplicate detection.
Highlights
- MTEB score 63.55 -- 95% of BGE Large precision at $0.08/hr, 3x throughput on the same ml.m5.xlarge instance
- 768-dimensional L2-normalized vectors, 512-token input -- drop-in for pgvector, OpenSearch, Pinecone, and Weaviate
- Runs in your own AWS VPC -- no external API calls, no token limits, no data leaving your account
Details
Introducing multi-product solutions
You can now purchase comprehensive solutions tailored to use cases and industries.
Features and programs
Financing for AWS Marketplace purchases
Pricing
Dimension | Description | Cost/host/hour |
|---|---|---|
ml.m5.xlarge Inference (Real-Time) Recommended | Model inference on the ml.m5.xlarge instance type, real-time mode | $0.08 |
ml.m5.xlarge Inference (Batch) Recommended | Model inference on the ml.m5.xlarge instance type, batch mode | $0.08 |
Vendor refund policy
No refunds.
How can we make this page better?
Legal
Vendor terms and conditions
Content disclaimer
Delivery details
Amazon SageMaker model
An Amazon SageMaker model package is a pre-trained machine learning model ready to use without additional training. Use the model package to create a model on Amazon SageMaker for real-time inference or batch processing. Amazon SageMaker is a fully managed platform for building, training, and deploying machine learning models at scale.
Version release notes
Initial release
Additional details
Inputs
- Summary
BAAI BGE Base EN v1.5 on SageMaker. 768-dimensional embeddings scoring 63.55 on MTEB -- 95% of Large quality at $0.08/hr. The mid-tier for production RAG when cost and quality both matter.
- Input MIME type
- application/json
Support
Vendor support
Contact support@waltsoft.net for deployment assistance.
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
Similar products
