Listing Thumbnail

    AI Agent DevOps Server Langfuse RAGAS LLM Judge Stack

     Info
    Sold by: Bansir LLC 
    Deployed on AWS
    This product has additional charges for a preconfigured AI Agent DevOps solution containing: Ollama with llama3.1:8b as LLM judge model, Langfuse v2 self-hosted for observability and trace management, PostgreSQL 16 for metadata storage, Redis 7 for cache and queues, RAGAS evaluation framework, FastAPI REST API with endpoints for single-answer evaluation and A/B comparison of AI agent versions, Python 3.12 with isolated venv, and systemd services with auto-start. 100% self-hosted agent quality evaluation on GPU Tesla T4 with no data leaving your infrastructure.

    Overview

    This product has additional charges applied for the use of a preconfigured solution. AI Agent DevOps Server is a complete, production-ready platform for evaluating, monitoring, and improving AI agents before production. Built with GPU acceleration, it ships fully configured: Ollama with llama3.1:8b as LLM judge, Langfuse v2 self-hosted for observability, RAGAS evaluation framework, PostgreSQL 16 and Redis 7, FastAPI REST API with /evaluate, /compare and /traces endpoints, Python 3.12 in isolated venv, and systemd services with auto-start. Optimized for NVIDIA Tesla T4 GPUs on EC2 g4dn instances.

    Key benefits include complete quality visibility with numeric scores stored in Langfuse, systematic A/B testing to compare agent versions objectively, and complete data privacy with no data leaving your account. Ideal for agent quality assurance before deployment, regression testing after prompt updates, continuous evaluation of production agents, and MLOps pipelines. Perfect for AI engineering teams, MLOps teams, and enterprises subject to HIPAA, GDPR, or SOC 2 seeking a private, GPU-accelerated platform for agent evaluation without per-token billing or vendor lock-in.

    For more Information: Guide: https://bansir-img.s3.us-east-1.amazonaws.com/Guide+AI+Agent+DevOps+Server.html 

    Highlights

    • Complete preconfigured AI agent evaluation platform with Ollama using llama3.1:8b as LLM judge, Langfuse v2 self-hosted for observability, RAGAS evaluation framework, PostgreSQL 16, Redis 7, FastAPI REST API, and systemd auto-start services. Deploy a production-grade MLOps evaluation pipeline in minutes instead of weeks, with /evaluate, /compare and /traces endpoints ready from first boot.
    • Replace guesswork with numeric quality scores: every AI answer gets rated on faithfulness, answer relevancy, and overall score. Use /compare to A/B test two agent versions against the same question and get an objective winner with score difference. Track quality trends over time in the Langfuse dashboard instead of relying on subjective human review.
    • 100% self-hosted evaluation with zero data leaving your infrastructure. Every trace, score, and prompt stays private, making it suitable for legal, healthcare, finance, and government workloads under HIPAA, GDPR, or SOC 2. Swap the judge model to any Ollama model (llama3.2:3b, qwen2.5:14b, mistral:7b), add custom metrics in Python, and integrate with your CI/CD pipeline via the REST API.

    Details

    Delivery method

    Delivery option
    64-bit (x86) Amazon Machine Image (AMI)

    Latest version

    Operating system
    Ubuntu 26.04

    Deployed on AWS
    New

    Introducing multi-product solutions

    You can now purchase comprehensive solutions tailored to use cases and industries.

    Multi-product solutions

    Features and programs

    Financing for AWS Marketplace purchases

    AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
    Financing for AWS Marketplace purchases

    Pricing

    AI Agent DevOps Server Langfuse RAGAS LLM Judge Stack

     Info
    Pricing is based on actual usage, with charges varying according to how much you consume. Subscriptions have no end date and may be canceled any time. Alternatively, you can pay upfront for a contract, which typically covers your anticipated usage for the contract duration. Any usage beyond contract will incur additional usage-based costs.
    Additional AWS infrastructure costs may apply. Use the AWS Pricing Calculator  to estimate your infrastructure costs.

    Usage costs (26)

     Info
    Dimension
    Cost/hour
    g4dn.xlarge
    Recommended
    $2.40
    c6i.large
    $1.60
    c6id.large
    $1.60
    c6in.large
    $1.60
    c6i.xlarge
    $2.40
    c6id.xlarge
    $2.40
    c6in.xlarge
    $2.40
    c6i.2xlarge
    $2.40
    c6id.2xlarge
    $2.40
    c6in.2xlarge
    $2.40

    Vendor refund policy

    For this offering, Bansir does not offer refund, you may cancel at anytime.

    How can we make this page better?

    Tell us how we can improve this page, or report an issue with this product.
    Tell us how we can improve this page, or report an issue with this product.

    Legal

    Vendor terms and conditions

    Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA) .

    Content disclaimer

    Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.

    Usage information

     Info

    Delivery details

    64-bit (x86) Amazon Machine Image (AMI)

    Amazon Machine Image (AMI)

    An AMI is a virtual image that provides the information required to launch an instance. Amazon EC2 (Elastic Compute Cloud) instances are virtual servers on which you can run your applications and workloads, offering varying combinations of CPU, memory, storage, and networking resources. You can launch as many instances from as many different AMIs as you need.

    Version release notes

    ver 2026

    Additional details

    Usage instructions

    1. LAUNCH: Subscribe and launch on g4dn.xlarge (minimum) or g4dn.2xlarge (recommended for production multi-user). Storage requirements: minimum 25 GB EBS (AMI base uses ~20 GB); recommended 30 GB for regular use; 50-100 GB for production with high trace volume in Langfuse. Open Security Group ports 22, 3000, 8000 restricted to your IP.

    2. VERIFY: SSH as ubuntu and run sudo systemctl status devops-api docker ollama. All must show active. Health check: curl -s http://localhost:8000/health 

    3. ACCESS: Swagger API at http://IP:8000/docs . Langfuse UI at http://IP:3000  (first user becomes admin). Judge model is llama3.1:8b by default.

    4. EVALUATE: Use POST /evaluate in Swagger with question, answer and optional context. Returns faithfulness, relevancy and overall scores.

    5. COMPARE: Use POST /compare in Swagger to A/B test two versions of your AI agent. Returns winner and score difference.

    6. DOCUMENTATION: Full bilingual guide (EN/ES) at https://bansir-img.s3.us-east-1.amazonaws.com/Guide+AI+Agent+DevOps+Server.html 

    Resources

    Vendor resources

    Support

    Vendor support

    AWS infrastructure support

    AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.

    Similar products

    Customer reviews

    Ratings and reviews

     Info
    0 ratings
    5 star
    4 star
    3 star
    2 star
    1 star
    0%
    0%
    0%
    0%
    0%
    0 reviews
    No customer reviews yet
    Be the first to review this product . We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.