Deepchecks is an AI evaluation platform, for evaluating LLM-based applications during research, ci/cd and production. Deepchecks' robust automatic scoring, version comparison, and auto-calculated metrics (such as relevance and grounded in context) enables AI teams to efficiently detect, troubleshoot, and improve their application's performance.
Deepchecks enables AI application developers and stakeholders to continuously validate LLM-based applications including characteristics, performance metrics, and potential pitfalls throughout the entire lifecycle from pre-deployment and internal experimentation to production.
Using the built-in and customizable properties, robust automatic scoring, version comparison, root cause analysis and production monitoring capabilities of Deepchecks, AI teams can achieve a comprehensive understanding of the application's performance, troubleshoot and improve it.
Highlights
Simulate expert human review with Deepchecks' multi-faceted automatic scoring system, transforming hours of manual evaluation into a one-click process
Detect issues using Deepchecks' pre-built Property Bank, combined with custom properties created by you - catch hallucinations, incomplete or incorrect responses, deviations from company policy, and more
Monitor your LLM apps in production, leveraging insights and configurations from the pre-deployment testing phase
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
Pricing is based on the duration and terms of your contract with the vendor. This entitles you to a specified quantity of use for the contract duration. If you choose not to renew or replace your contract before it ends, access to these entitlements will expire.
Additional AWS infrastructure costs may apply. Use the AWS Pricing Calculator to estimate your infrastructure costs.
You choose one of three contract tiers, each priced as units. The tiers scale by capacity limits, so you pick based on how many applications and users you need to support. Deepchecks Starter covers up to 3 applications, 10 users, and 80M tokens uploaded monthly. Deepchecks Standard raises the limits to 10 applications and 30 users. Deepchecks Pro extends coverage to 30 applications and 70 users and adds enterprise-grade tools. As your application count and user base grow, you move to the tier that fits those limits.
Top-of-mind questions for buyers
What counts as one application, and how are the token limits measured?
An application is one LLM-powered pipeline you evaluate, such as a RAG chatbot, an agent workflow, or a summarization tool. Tokens are the text units uploaded for evaluation. Deepchecks Starter caps uploads at 80M tokens monthly. Standard and Pro raise the application and user limits instead.
What happens if I outgrow my tier's application or user limits?
Each tier sets fixed ceilings on applications and users. When your application count or user base passes those ceilings, you move up to the tier that fits. The limits apply to your whole account, not just added users. Deepchecks Pro also adds enterprise-grade tools beyond raised capacity.
Which capacity limit matters most when choosing a tier?
You are bound by whichever limit you hit first: applications, users, or, for Starter, monthly tokens. Deepchecks Starter fits small teams with few applications. Standard suits more applications and users. Pro supports the highest counts and adds enterprise-grade tools. Pick the tier that covers all three limits.
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
SaaS delivers cloud-based software applications directly to customers over the internet. You can access these applications through a subscription model. You will pay recurring monthly usage fees through your AWS bill, while AWS handles deployment and infrastructure management, ensuring scalability, reliability, and seamless integration with other AWS services.
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
Deepchecks is an AI evaluation platform, for evaluating LLM-based applications during research, ci/cd and production. Deepchecks' robust automatic scoring, version comparison, and auto-calculated metrics (such as relevance and grounded in context) enables AI teams to efficiently detect, troubleshoot, and improve their application's performance. It supports single-step workflows, multi-step workflows, chat, and agentic use cases.
Ask Sage is a secure, model and cloud-agnostic generative AI platform trusted by 15,000 government teams and 2,500 companies. Users have access to any commercial and government approved Large Language Model and various open-source models with ability to ingest CUI data, and integrate with existing systems via APIs. IL5/IL6 authorized with Top Secret deployments. Ask Sage is built with zero-trust architecture and label-based access control. Deploy on any cloud, on-premise, or air-gapped.
Ask Sage is a secure, model and cloud-agnostic generative AI platform trusted by 15,000 government teams and 2,500 companies. Users have access to any commercial and government approved Large Language Model and various open-source models with ability to ingest CUI data, and integrate with existing systems via APIs. IL5/IL6 authorized with Top Secret deployments. Ask Sage is built with zero-trust architecture and label-based access control. Deploy on any cloud, on-premise, or air-gapped.
Ask Sage is a secure, model and cloud-agnostic generative AI platform trusted by 15,000 government teams and 2,500 companies. Users have access to any commercial and government approved Large Language Model and various open-source models with ability to ingest CUI data, and integrate with existing systems via APIs. IL5/IL6 authorized with Top Secret deployments. Ask Sage is built with zero-trust architecture and label-based access control. Deploy on any cloud, on-premise, or air-gapped.
Assess a transition from OpenAI to AWS with Applogika's comprehensive Migration Assessment – 75%-100% funded by AWS. Discover cost savings, enhanced model tuning, and compliance benefits within 4-6 weeks. Our assessment provides a detailed comparison, bespoke migration roadmap, and executive go/no go decision guide, ensuring your AI strategy aligns with security, performance, and cost-efficiency goals.
Complimentary - GenPrompt (AI Portal ) deployment.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.