Noveum is the AI-native reliability platform for production AI agents and LLM apps. It traces every run, evaluates quality with 106 calibrated scorers, simulates voice and chat scenarios, and ships validated fixes as pull requests.
Noveum gives teams building AI agents, RAG pipelines and LLM applications one platform to see, measure and improve how their AI behaves in production.
Tracing: the open-source NovaTrace SDKs for Python and TypeScript capture every LLM call, tool call, retrieval step and agent hop as a structured trace, with latency, token usage and cost. Integrations cover LangChain, LangGraph, LiveKit, Pipecat and CrewAI in a few lines of code.
Evaluation: NovaEval scores sampled production traces with 106 calibrated LLM-as-judge scorers across 18 categories, including 17 dedicated voice scorers. Every verdict comes with its reasoning, and datasets are built automatically from real application logs.
Simulation: NovaSynth forward-simulates voice (SIP) and chat scenarios with personas, accents, interruptions and tool virtualization to validate end-to-end outcomes before release.
Autonomous fixing: NovaPilot reads failing evaluations, isolates root causes and ships fixes as pull requests, each backtested and validated in simulation.
Billing is organization-flat: one price per organization, not per seat. Each plan includes a monthly credit allowance for evaluation, simulation and analysis, plus span and storage allowances. Subscribing through AWS Marketplace consolidates Noveum on your AWS bill.
Highlights
Trace every LLM call, tool call, RAG step and agent hop with open-source Python and TypeScript SDKs, including latency, token and cost metrics.
Evaluate production traces with 106 calibrated LLM-as-judge scorers, simulate voice and chat scenarios, and catch regressions before users do.
NovaPilot isolates root causes of failing evaluations and ships validated fixes as pull requests. Organization-flat pricing with no seat fees.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
Pricing is based on the duration and terms of your contract with the vendor. This entitles you to a specified quantity of use for the contract duration. If you choose not to renew or replace your contract before it ends, access to these entitlements will expire.
Additional AWS infrastructure costs may apply. Use the AWS Pricing Calculator to estimate your infrastructure costs.
These four plans form a tiered contract. Each step up raises your monthly allowances across four dimensions: credits, spans, storage, and seats. Pro gives the smallest allotments, and Growth, Max, then Scale raise each limit together. Scale also removes the seat cap and makes seats unlimited. Billing is per organization, not per person. Credits act as one shared currency for eval scoring, voice and text testing, and trace analysis. Spans measure captured activity from your AI calls. Storage holds that data. Pick the tier whose allowances match your expected monthly usage.
Top-of-mind questions for buyers
What is a Noveum Credit, and how do different activities consume it?
A credit is a shared unit for metered activities. One premium LLM scorer check on up to 4,000 input tokens costs 1 credit. Voice agent testing costs 100 credits per call minute. Text conversations cost 12 credits each. Trace analysis costs 8 credits per trace, with a minimum of 800 credits per report. Rule-based scorers cost nothing.
What happens when I use up my monthly credits, spans, or storage?
Credit-consuming features stop when your balance hits zero. The interface shows a message, and API calls return a quota error. You can buy a one-time add-on credit pack that never expires, upgrade your plan, or opt into usage-based overage billing for credits, spans, and storage. Overage is off by default.
How do credits, spans, and storage combine on one plan?
Each plan bundles all three allowances together per organization, not per seat. Credits meter your eval scoring, voice and text testing, and trace analysis. Spans count captured activity from AI calls. Storage holds that data. Credits usually drive active usage, while spans and storage grow with the volume of traced traffic you retain.
noveum.ai
Helpful?
Vendor refund policy
Subscriptions can be cancelled at any time and stay active until the end of the current billing period. Refunds for unused time are considered case by case within 14 days of purchase; contact hello@noveum.ai with your AWS account ID and agreement ID.
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
SaaS delivers cloud-based software applications directly to customers over the internet. You can access these applications through a subscription model. You will pay recurring monthly usage fees through your AWS bill, while AWS handles deployment and infrastructure management, ensuring scalability, reliability, and seamless integration with other AWS services.
Email support is included with every plan at hello@noveum.ai, with responses within one business day. Documentation, SDK guides and integration recipes are available at https://noveum.ai/docs. Enterprise plans include a dedicated support channel and onboarding assistance.
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
Observe Inc. is redefining modern AI-powered observability at scale. The only company to offer a platform built on an open data lake with a proprietary Knowledge Graph and AI SRE, Observe enables users to troubleshoot faster at drastically lower cost.
Tackle underperforming factors at the sight by visualizing IT performance in real-time. Get a better view of multiple environments instantly at ease, enabling swift recognition and response.
Deepchecks is an AI evaluation platform, for evaluating LLM-based applications during research, ci/cd and production. Deepchecks' robust automatic scoring, version comparison, and auto-calculated metrics (such as relevance and grounded in context) enables AI teams to efficiently detect, troubleshoot, and improve their application's performance. It supports single-step workflows, multi-step workflows, chat, and agentic use cases.
Observe.AI combines Interaction Intelligence and Operations Agents to analyze every customer interaction, automate quality and performance management, and uncover insights that improve customer experience and contact center operations. Teams can identify trends, evaluate performance, coach more effectively, and turn conversation data into action at scale.
Professional service to build AI model invocation dashboards on AWS, visualizing usage, latency, errors, token consumption, and cost across Amazon Bedrock and related services.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.