Listing Thumbnail

    Datafold

     Info
    Sold by: Datafold 
    Automated testing for Data Engineers. Datafold is the fastest way to validate dbt model changes during development, deployment & migrations. With Datafold, data engineers can audit their work in minutes without writing tests or custom queries. Integrated into CI, Datafold enables data teams to deploy with full confidence, ship faster, and leave tedious QA and firefighting behind.
    4.5

    Overview

    Datafold is the fastest way to validate dbt model changes during development, deployment & migrations. With Datafold, data engineers can audit their work in minutes without writing tests or custom queries. Integrated into CI, Datafold enables data teams to deploy with full confidence, ship faster, and leave tedious QA and firefighting behind.

    Automate proactive testing for all data transformations. Supercharge your dbt workflows with seamless dbt Cloud and Core integrations.

    Know exactly what will happen to data and data applications once the code is deployed, right in the pull request. Identify breaking changes, sudden metric shifts and edge cases before they do any damage to the business.

    Stop guessing what this regex does or arguing if that CASE WHEN statement has correct logic. No more custom scripts and audit spreadsheets to fill.

    Stop surprising your data users with unexpected metric changes and broken dashboards. Easily share impact reports with everyone and give heads up before deploying the changes to production.

    Manual data testing is hard, tedious, and error-prone. Focus on what matters and not on writing boilerplate tests, custom scripts and filling out audit spreadsheets.

    With full visibility into every change, everyone, not just data team, can contribute, because testing and reviewing code is so easy!

    Highlights

    • Development testing for dbt
    • Deployment testing for dbt
    • Migration testing

    Details

    Sold by

    Delivery method

    Deployed on AWS
    New

    Introducing multi-product solutions

    You can now purchase comprehensive solutions tailored to use cases and industries.

    Multi-product solutions

    Features and programs

    Trust Center

    Trust Center
    Access real-time vendor security and compliance information through their Trust Center powered by Drata or Vanta. Review certifications and security standards before purchase.

    Financing for AWS Marketplace purchases

    AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
    Financing for AWS Marketplace purchases

    Pricing

    Pricing is based on the duration and terms of your contract with the vendor, and additional usage. You pay upfront or in installments according to your contract terms with the vendor. This entitles you to a specified quantity of use for the contract duration. Usage-based pricing is in effect for overages or additional usage not covered in the contract. These charges are applied on top of the contract price. If you choose not to renew or replace your contract before the contract end date, access to your entitlements will expire.
    Additional AWS infrastructure costs may apply. Use the AWS Pricing Calculator  to estimate your infrastructure costs.

    12-month contract (2)

     Info
    Dimension
    Description
    Cost/12 months
    Overage cost
    10 Developers
    10 Provisioned Developers
    $30,000.00
    5 Developers
    5 Provisioned Developers
    $15,000.00
    -

    Vendor refund policy

    All fees are non-cancellable and non-refundable except as required by law.

    How can we make this page better?

    Tell us how we can improve this page, or report an issue with this product.
    Tell us how we can improve this page, or report an issue with this product.

    Legal

    Vendor terms and conditions

    Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA) .

    Content disclaimer

    Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.

    Usage information

     Info

    Delivery details

    Software as a Service (SaaS)

    SaaS delivers cloud-based software applications directly to customers over the internet. You can access these applications through a subscription model. You will pay recurring monthly usage fees through your AWS bill, while AWS handles deployment and infrastructure management, ensuring scalability, reliability, and seamless integration with other AWS services.

    Resources

    Support

    AWS infrastructure support

    AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.

    Product comparison

     Info
    Updated weekly
    By Datafold
    By Paradime Labs, Inc.
    By dbt Labs

    Accolades

     Info
    Top
    50
    In Data Preparation, ELT/ETL
    Top
    100
    In Analytic Platforms
    Top
    10
    In Data Analytics, ELT/ETL, Business Intelligence & Advanced Analytics

    Customer reviews

     Info
    Sentiment is AI generated from actual customer reviews on AWS and G2
    Reviews
    Functionality
    Ease of use
    Customer service
    Cost effectiveness
    24 reviews
    Insufficient data
    0 reviews
    Insufficient data
    Insufficient data
    Insufficient data
    Insufficient data
    Positive reviews
    Mixed reviews
    Negative reviews

    Overview

     Info
    AI generated from product descriptions
    dbt Integration
    Seamless integration with dbt Cloud and Core for automated testing of data transformations and model changes
    Automated Data Validation
    Automated testing capability that validates data model changes without requiring manual test writing or custom queries
    CI/CD Pipeline Integration
    Integration into continuous integration workflows to enable automated validation during development, deployment, and migration phases
    Change Impact Analysis
    Detection and identification of breaking changes, metric shifts, and edge cases in data transformations before production deployment
    Data Quality Reporting
    Generation of impact reports and visibility into data changes for stakeholder communication and deployment decision-making
    AI-Native Code Development Environment
    Integrated development environment with AI capabilities for coding data pipelines using dbt and Python, featuring built-in warehouse access and column-level lineage context.
    State-Aware Pipeline Scheduling
    Scheduler supporting state-aware execution of dbt and Python data pipelines with column-level impact analysis for CI testing.
    Data Lineage and Impact Analysis
    Column-level lineage tracking and impact analysis capabilities for understanding data dependencies and transformation effects across pipelines.
    Warehouse Cost Optimization
    AI-agent based monitoring and optimization system operating continuously to reduce warehouse operational costs.
    Data Pipeline Orchestration Integration
    Support for orchestrating multi-tool data workflows including Fivetran ingestion, data transformation pipelines, and downstream application refreshes for Tableau and PowerBI.
    Integrated Development Environment
    IDE built for dbt with SQL Runner capable of executing Jinja templates and guided version control enforcement for git best practices
    Job Scheduling and Orchestration
    Custom scheduling capabilities for production jobs with incremental testing triggered on change or before deployment
    CI/CD Pipeline Support
    Continuous integration and continuous deployment functionality for automated dbt project workflows
    Enterprise Security and Compliance
    SOC2 Type II compliance certification, single sign-on (SSO) authentication, and role-based access control
    Monitoring and Alerting
    Built-in monitoring and alerting capabilities for tracking job execution and system health

    Contract

     Info
    Standard contract
    No

    Customer reviews

    Ratings and reviews

     Info
    4.5
    25 ratings
    5 star
    4 star
    3 star
    2 star
    1 star
    76%
    20%
    4%
    0%
    0%
    0 AWS reviews
    |
    25 external reviews
    External reviews are from G2  and PeerSpot .
    JohnBosco Obi

    Data diffing has caught regressions early and now reporting and usability still need improvement

    Reviewed on Jul 18, 2026
    Review provided by PeerSpot

    What is our primary use case?

    My main use case for Datafold  is to automatically compare data differences through data diffing, where I can compare datasets row-by-row, column-by-column, to surface unexpected changes.

    I also use it for CI/CD integration for data pipelines, where we embed data quality checks into GitHub  and GitLab  pull requests. Another interesting feature I use Datafold  for is column-level lineage to trace how specific columns flow through SQL transformations across the data stack we have.

    Recently, for data diffing, there was one time we wanted to conduct a lot of comparisons on a very large dataset to be sure that all was in check before we migrated to Snowflake , which is a complex and resource-heavy process.

    We used Datafold to do the automatic checks across rows and columns to ensure that when we completed the migration, no data was lost and the migration into Snowflake  was very smooth. That was in March, and it was very useful.

    For a friend who uses Datafold in their enterprise, an engineer made a change from an SQL model and did not realize it would cascade and alter downstream dashboard metrics.

    However, with the help of Datafold integrated into the pull request workflow, the data diff actually surfaced exactly which columns changed, which rows were affected, and by how much.

    What is most valuable?

    The data diff stands out the most for me as the best feature because it gives me that column-level and row-level comparison between any two datasets, highlighting what changed, even down to characters and whitespaces.

    It also tells me when something has failed, and it makes debugging faster and more precise.

    Column-level lineage derived from SQL static analysis helps me trace how columns flow through transformations across the entire pipeline.

    Datafold has positively impacted my organization by helping us catch data regressions before they reach production.

    Even while we are still building the pipeline and mapping out how the data will flow, we catch any data regressions and alterations before they reach production. This helps us avoid a lot of debugging when the migration has happened or when we reach production level and data goes live.

    Additionally, it significantly reduces manual testing time, saving us time that we can put into useful work. Furthermore, it builds our team's confidence when pushing code changes, as the data reviewers, data annotators, and engineers can actually see the data impact of every change, not just the code change.

    What needs improvement?

    The first pain point for me is that the reporting capabilities are weak. The ease of setup is also challenging, particularly for those who are not tech-savvy or do not know how to navigate it. Most importantly, there is no free trial, so you cannot deploy and test it to see the efficiency and use case before purchasing.

    I feel the features need more attention because sometimes finding the data I want to use or compare is challenging, especially when looking for it inside the SaaS.

    Automated workflows also break sometimes, but whenever that happens, the good thing is that it gives you an error report so you know where the issue is coming from.

    For how long have I used the solution?

    I have been using Datafold for about six months.

    What other advice do I have?

    I give Datafold a seven out of ten because it is excellent at its core specialization, which is data diffing and CI/CD integration, and its migration automation capabilities are strong and very AI-powered. However, I lose three points due to the weak reporting.

    Even though it provides reports when there is a breakage in your push or migration, I sometimes cannot get the full scope of what I want when producing an actual report. Additionally, the lack of a free trial is a downside.

    Regarding Datafold's governance and security, I rate it high because its migration agent uses LLMs for SQL translation and validation, which is a strong point.

    Additionally, there is a self-hosted deployment option available, so organizations with strict data residency or compliance requirements can run Datafold entirely within their own cloud environment, which could be either AWS , GCP , or Azure , ensuring that data never leaves the perimeter of the organization.

    Datafold's output is highly accurate and very reliable. The migration agent's accuracy, which utilizes LLMs to convert SQL dialects, is excellent.

    Furthermore, the data diff, which is the most reliable AI-adjacent feature, is deterministic, not generative, and it compares actual data values mathematically rather than using inference.

    Because it does the comparison mathematically, the outputs are highly accurate and consistently reliable.

    Datafold is deployed in my organization as a cloud, specifically as a SaaS, which is fully managed by Datafold. This means that hosting, maintenance, and automatic updates are all managed by Datafold, making it simpler for us and easier to get started.

    The advice I would give others looking to use Datafold is that whoever is handling it, perhaps the head of IT, should have a sit-down with the analysts to ensure it fits into the stack that the organization is already conversant with.

    Datafold is purpose-built for SQL and warehouse-based analytic pipelines with DBT, so if the current stack does not include a data warehouse and DBT, I would advise them to evaluate alternatives first. I also recommend using it for CI/CD quality, not just for general observability, because Datafold excels at pre-merge testing and data diffs.

    Therefore, if the primary need is broad production or observability, the organization should also check out other options.

    I rate Datafold a seven out of ten overall.

    Information Technology and Services

    Right one for testing

    Reviewed on Apr 20, 2023
    Review provided by G2
    What do you like best about the product?
    Awesome workflow, which is really a great feature of Datafold. Traditional way of doing data transfers is laborious. But this one helps to the best of capabilities. My fellow team was impressed.
    What do you dislike about the product?
    Nothing as such to say about negative here. But breadth of usage can be enhanced. This is not a drawback for sure. May be in coming days, will explore more and revert. For now all good
    What problems is the product solving and how is that benefiting you?
    Integration of data and managing data quality are two things that pose a big challenge in front of us. But Datafold made it easier to look and saved our time and effort
    Telecommunications

    Review for Datafold

    Reviewed on Apr 19, 2023
    Review provided by G2
    What do you like best about the product?
    Makes life easier with SQL code reviews helping find the hidden changes we made we dint know to our data.It gives ability to quality check the things in our own way with very much lesser errors compared to manual testing.
    What do you dislike about the product?
    Nothing in particular, a brief guide with documentation would have been justified.
    Helps in data managing of huge chunks of table, rows and records of data in day to day usage.
    What problems is the product solving and how is that benefiting you?
    Makes life way easier for automated testing
    Data accuracy and quality kpi achievement.easily integrates with modern data stacks such as Amazon redshift, snowflake,gitlab and GitHub
    Sanjay D.

    Great platform for improving data quality

    Reviewed on Apr 19, 2023
    Review provided by G2
    What do you like best about the product?
    Datafold is a valuable data testing and monitoring tool that helps users ensure the quality and accuracy of their pipelines. It detects errors, improves data quality, and helps troubleshoot problems quickly, increasing confidence in the data.
    What do you dislike about the product?
    Datafold is an amazing platform I like all the features provided by this, but there is one limitation I observed that it has a steep learning curve, limited integration options, and a commercial subscription requirement these factors challenges for us in adopting and using this tool effectively.
    What problems is the product solving and how is that benefiting you?
    Datafold offers a complete picture of data pipeline health, performance, and quality for data teams to monitor and ensure smooth operations.
    Datafold automates the process of detecting data pipeline issues and ensures data quality. It saves time and resources by automating data testing, which would otherwise be time-consuming and error-prone if done manually.
    Madhusudhan M.

    "Best tool for data quality and consistency"

    Reviewed on Apr 18, 2023
    Review provided by G2
    What do you like best about the product?
    It is simple and User friendly UI. It provides real-time updates.It includes best range in build tests and checks for data quality, consistency and saving time.
    What do you dislike about the product?
    There is no negative thing about the software but the price is little bit high compared to other softwares in the market.
    What problems is the product solving and how is that benefiting you?
    It is very much useful for data engineers for automation testing.Data Quality monitoring is too easy without any hard failures compared to other tools which are inflexible and takes hard failures.
    View all reviews