Overview
This product has additional charges applied for the use of a preconfigured solution. AI Agent DevOps Server is a complete, production-ready platform for evaluating, monitoring, and improving AI agents before production. Built with GPU acceleration, it ships fully configured: Ollama with llama3.1:8b as LLM judge, Langfuse v2 self-hosted for observability, RAGAS evaluation framework, PostgreSQL 16 and Redis 7, FastAPI REST API with /evaluate, /compare and /traces endpoints, Python 3.12 in isolated venv, and systemd services with auto-start. Optimized for NVIDIA Tesla T4 GPUs on EC2 g4dn instances.
Key benefits include complete quality visibility with numeric scores stored in Langfuse, systematic A/B testing to compare agent versions objectively, and complete data privacy with no data leaving your account. Ideal for agent quality assurance before deployment, regression testing after prompt updates, continuous evaluation of production agents, and MLOps pipelines. Perfect for AI engineering teams, MLOps teams, and enterprises subject to HIPAA, GDPR, or SOC 2 seeking a private, GPU-accelerated platform for agent evaluation without per-token billing or vendor lock-in.
For more Information: Guide: https://bansir-img.s3.us-east-1.amazonaws.com/Guide+AI+Agent+DevOps+Server.html
Highlights
- Complete preconfigured AI agent evaluation platform with Ollama using llama3.1:8b as LLM judge, Langfuse v2 self-hosted for observability, RAGAS evaluation framework, PostgreSQL 16, Redis 7, FastAPI REST API, and systemd auto-start services. Deploy a production-grade MLOps evaluation pipeline in minutes instead of weeks, with /evaluate, /compare and /traces endpoints ready from first boot.
- Replace guesswork with numeric quality scores: every AI answer gets rated on faithfulness, answer relevancy, and overall score. Use /compare to A/B test two agent versions against the same question and get an objective winner with score difference. Track quality trends over time in the Langfuse dashboard instead of relying on subjective human review.
- 100% self-hosted evaluation with zero data leaving your infrastructure. Every trace, score, and prompt stays private, making it suitable for legal, healthcare, finance, and government workloads under HIPAA, GDPR, or SOC 2. Swap the judge model to any Ollama model (llama3.2:3b, qwen2.5:14b, mistral:7b), add custom metrics in Python, and integrate with your CI/CD pipeline via the REST API.
Details
Introducing multi-product solutions
You can now purchase comprehensive solutions tailored to use cases and industries.
Features and programs
Financing for AWS Marketplace purchases
Pricing
Dimension | Cost/hour |
|---|---|
g4dn.xlarge Recommended | $2.40 |
c6i.large | $1.60 |
c6id.large | $1.60 |
c6in.large | $1.60 |
c6i.xlarge | $2.40 |
c6id.xlarge | $2.40 |
c6in.xlarge | $2.40 |
c6i.2xlarge | $2.40 |
c6id.2xlarge | $2.40 |
c6in.2xlarge | $2.40 |
Vendor refund policy
For this offering, Bansir does not offer refund, you may cancel at anytime.
How can we make this page better?
Legal
Vendor terms and conditions
Content disclaimer
Delivery details
64-bit (x86) Amazon Machine Image (AMI)
Amazon Machine Image (AMI)
An AMI is a virtual image that provides the information required to launch an instance. Amazon EC2 (Elastic Compute Cloud) instances are virtual servers on which you can run your applications and workloads, offering varying combinations of CPU, memory, storage, and networking resources. You can launch as many instances from as many different AMIs as you need.
Version release notes
ver 2026
Additional details
Usage instructions
-
LAUNCH: Subscribe and launch on g4dn.xlarge (minimum) or g4dn.2xlarge (recommended for production multi-user). Storage requirements: minimum 25 GB EBS (AMI base uses ~20 GB); recommended 30 GB for regular use; 50-100 GB for production with high trace volume in Langfuse. Open Security Group ports 22, 3000, 8000 restricted to your IP.
-
VERIFY: SSH as ubuntu and run sudo systemctl status devops-api docker ollama. All must show active. Health check: curl -s http://localhost:8000/health
-
ACCESS: Swagger API at http://IP:8000/docs . Langfuse UI at http://IP:3000 (first user becomes admin). Judge model is llama3.1:8b by default.
-
EVALUATE: Use POST /evaluate in Swagger with question, answer and optional context. Returns faithfulness, relevancy and overall scores.
-
COMPARE: Use POST /compare in Swagger to A/B test two versions of your AI agent. Returns winner and score difference.
-
DOCUMENTATION: Full bilingual guide (EN/ES) at https://bansir-img.s3.us-east-1.amazonaws.com/Guide+AI+Agent+DevOps+Server.html
Resources
Vendor resources
Support
Vendor support
Remote support support@bansircloud.com
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.