Overview
BGE Large EN v1.5 from BAAI (Beijing Academy of AI) is a high-quality English text embedding model that scored 64.23 average on the MTEB benchmark across 56 tasks, compared to 60.99 for OpenAI text-embedding-ada-002. It produces 1024-dimensional dense vectors from inputs up to 512 tokens, ready for cosine similarity without post-processing.
This listing deploys the model as a SageMaker real-time endpoint in your own AWS account. Your documents never leave your VPC -- no external API calls, no third-party data access, no rate limits. You control the endpoint, the scaling, and the logs.
Integrates directly with LangChain (HuggingFaceBgeEmbeddings), LlamaIndex, and Haystack. Works with any vector store that accepts 1024-dimensional float32 vectors: pgvector on Aurora/RDS, Amazon OpenSearch, Pinecone, Weaviate, Qdrant, Chroma, and Milvus.
Deployed on ml.m5.xlarge (CPU). Flat hourly billing -- you pay for the compute, not per token. Batch transform available for offline document ingestion at the same instance rate.
Primary use cases: document retrieval for RAG pipelines, semantic search over internal knowledge bases, customer support ticket routing, duplicate detection, and product catalog similarity.
Highlights
- MTEB score 64.23 vs Ada-002's 60.99 -- higher retrieval accuracy on standard benchmarks, running entirely inside your AWS VPC
- 1024-dimensional vectors, 512-token inputs, L2-normalized output -- drop-in for pgvector, OpenSearch, Pinecone, and Weaviate with no post-processing
- Flat $0.10/hr on ml.m5.xlarge -- no per-token charges, no rate limits, no API key management
Details
Introducing multi-product solutions
You can now purchase comprehensive solutions tailored to use cases and industries.
Features and programs
Financing for AWS Marketplace purchases
Pricing
Dimension | Description | Cost/host/hour |
|---|---|---|
ml.m5.xlarge Inference (Real-Time) Recommended | Model inference on the ml.m5.xlarge instance type, real-time mode | $0.10 |
ml.m5.xlarge Inference (Batch) Recommended | Model inference on the ml.m5.xlarge instance type, batch mode | $0.10 |
Vendor refund policy
No refunds.
How can we make this page better?
Legal
Vendor terms and conditions
Content disclaimer
Delivery details
Amazon SageMaker model
An Amazon SageMaker model package is a pre-trained machine learning model ready to use without additional training. Use the model package to create a model on Amazon SageMaker for real-time inference or batch processing. Amazon SageMaker is a fully managed platform for building, training, and deploying machine learning models at scale.
Version release notes
Initial release
Additional details
Inputs
- Summary
BAAI BGE Large EN v1.5 deployed as a private SageMaker endpoint. Scored 64.23 on MTEB, outperforming OpenAI Ada-002 (60.99) on retrieval tasks. Runs in your own AWS account -- no external API calls, no token limits.
- Input MIME type
- application/json
Support
Vendor support
Contact support@waltsoft.net for deployment assistance.
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
Similar products
