ArgusRL verifies AI-generated responses with evidence-based atomic claim verdicts, citations, and confidence scores for AI engineering and MLOps teams.support model evaluation, regression testing and post-training workflows.
ArgusRL is an evidence-based evaluation solution that independently verifies AI-generated responses against external evidence. It delivers deterministic verdicts with supporting evidence, citations and confidence scores for model evaluation, regression testing and post-training workflows.
Try ArgusRL Before You Subscribe: Explore the ArgusRL Playground at no cost. Submit prompts, inspect API responses and review verification results.
About TrustScale
ArgusRL is built by TrustScale, an AI training, evaluation, and assurance company. TrustScale combines deep expertise in human annotation, AI evaluation methodologies, and scalable verification systems, with data operations spanning multiple languages, to deliver reliable evaluation signals for AI engineering teams.
Why ArgusRL?
Reinforcement Learning from Human Feedback (RLHF) remains the gold standard but it is costly, time-intensive and difficult to scale.
LLM-as-a-Judge automates evaluation at scale but an AI judging an AI remains probabilistic and inherits the same biases and failure modes as the model being evaluated.
ArgusRL bridges this gap by combining the scalability of automated evaluation with independent, evidence backed verification. Every claim is checked against external evidence, results are deterministic, repeatable and auditable, enabling consistent benchmarking, regression testing and release validation.
Benchmark Results: In customer benchmark evaluations, ArgusRL achieved more than 90% agreement with human reviewer assessments on the evaluated dataset. Results vary by dataset, domain, evidence availability, and evaluation configuration.
How Teams Use ArgusRL
ArgusRL is built for AI engineering and MLOps teams that need a consistent, repeatable evaluation verification signal for model development and production AI.
How it Works:
Submit: Submit a prompt/response pair, multi-turn conversation, or batch file through the API.
Verify: ArgusRL decomposes responses into atomic claims and independently verifies each claim against external evidence.
Review Results: Each API response returns structured claim-level results, including Supported, Contradicted, or No Evidence verdicts, supporting citations, confidence scores, and machine-readable JSON for downstream workflows.
Use the Results to:
Benchmark and compare model performance
Run repeatable regression tests
Generate post-training labels or reward signals
Route uncertain responses for human review
Monitor production AI quality over time
Highlights
Independent Verification: Verdicts are grounded in external evidence, so the verification signal stays independent of the model under test and doesn't inherit its biases or failure modes.
Claim-Level Granularity: Instead of a single pass/fail score per response, ArgusRL returns a verdict, citations and a confidence score for every atomic claim, so you can pinpoint exactly what failed and why.
Built for Automation: Standard APIs and structured JSON outputs drop into CI/CD pipelines, eval harnesses and monitoring stack.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
This service uses one usage-based pricing dimension. You pay per atomic claim processed. A claim is a single factual statement that the service extracts from an AI-generated response. When you submit a query and its AI answer, the service breaks the answer into separate claims and checks each one against external evidence. Your cost scales directly with how many claims you process. More responses, or responses with more factual statements, mean more claims and higher usage. There are no seats or fixed terms tied to this dimension.
Top-of-mind questions for buyers
What exactly counts as one atomic claim for billing?
An atomic claim is a single factual statement the service extracts from an AI answer. When you submit a query and its response, the engine decomposes that response into separate claims. Each one is checked against external evidence and counted individually. One response can contain several claims, so each adds to your usage.
Am I charged if a claim comes back unverifiable or has no supporting evidence?
Each claim is processed regardless of its verdict. The service returns supported, unsupported, or unverifiable results, plus a confidence score. Every claim it decomposes and evaluates counts toward usage. The verdict does not change whether a claim is billed. Processing effort, not outcome, drives the count.
Do polling or listing requests add to my claim count?
Billing meters atomic claims processed, not API calls. You submit a detection request, then poll a status endpoint until results are ready. Polling for results and listing past submissions do not create new claims. Only the factual statements decomposed from each submitted response count toward usage.
api.trustscale.ai
Helpful?
Vendor refund policy
Charges are based on claims processed and are non-refundable once billed, except for verified metering or billing errors reported within 60 days. Canceling stops future billing but does not refund prior usage. Contact argushelp@trustscale.ai for billing questions.
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
API-Based Agents and Tools integrate through standard web protocols. Your applications can make API calls to access agent capabilities and receive responses.
Additional details
Usage instructions
API
fter subscribing, you will be redirected to the ArgusRL registration page to complete setup and receive your API key (prefixed ts_api_).
Store it securely; it authenticates every API call via the x-api-key header.
The API is asynchronous: POST /detect with your query and response, receive a submission_id, then poll GET /detect/{submission_id} until status is completed to retrieve claim-level results.
For API issues, billing questions, integration support, or refund requests, reach our team at the email above.
Getting Started
After subscribing through AWS Marketplace, you will be redirected to our own fulfillment page to receive API credentials to begin making evaluation calls. Visit the API documentation for authentication setup, request formats, and endpoint details.
Steps to your first evaluation:
Subscribe to ArgusRL through AWS Marketplace
Obtain your API key from the onboarding process
Review the API documentation for request format and authentication
Submit your first prompt/response pair for verification
Review the structured JSON response with claim-level verdicts
Need More Than the API?
For the full enterprise solution including domain-specific rule sets, custom evaluation guidelines, and human annotation services, contact argushelp@trustscale.ai.
Security and Data Handling
ArgusRL uses encrypted connections (TLS) for all API communications. For details on data retention, access controls, and compliance certifications, contact argushelp@trustscale.ai.
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.