Sold by

Databricks Data Intelligence Platform
The Databricks Data Intelligence Platform unlocks the power of data and AI for your entire organization. Enjoy up to $400 in usage credits during your 14-day free trial. Cancel anytime. After your trial ends, you will automatically be enrolled into a Databricks pay-as-you-go plan.
Reviews (1387)
Anonymous
Comprehensive AI Platform with Unified Governance
Reviewed on Aug 06, 2026
Review provided by G2
What do you like best about the product?
I use Databricks as the backbone for all my data and AI work. I love that everything is in a single place and all connects, which allows me to manage everything from raw data digestion to deploying AI products. Unity Catalog provides a single governance source of truth with permissions throughout the platform, which is incredibly useful. Genie Spaces offer an interface in natural language for non-technical stakeholders, saving me time as I don't need to create additional Tableau reports. I can ship everything as code rather than clicking through notebooks manually, which lets me scale much faster. Databricks is a comprehensive platform and is ahead of the curve when it comes to artificial intelligence.
What do you dislike about the product?
The biggest issue I have is that I'm not applied for admin so I don't have visibility into a lot of my costs or access to control more things in the platform. I also wish there was a way in Genie code to make it more clear which aspects are within my control and which are not, and to have an integrated system to notify my platform admin when there are issues.
What problems is the product solving and how is that benefiting you?
I use Databricks as the backbone for all my data and AI work. It's an all-in-one platform, eliminating the need for handoffs between different tools. Unity Catalog gives me a single governance source, ensuring consistency and saving my time.
Jonathan B.
Amazing Platform for Running AI at Massive Scale
Reviewed on Aug 06, 2026
Review provided by G2
What do you like best about the product?
Amazing platform to run AI on a massive scale without complexity
What do you dislike about the product?
The pricing can be challenging if you don't have enough knowledge
What problems is the product solving and how is that benefiting you?
It’s one of the best ways to run AI and ETL on large amounts of data. Databricks also comes with industry-ready features that make it easier to work at scale.
Gina S.
Unified Data Management with Governance Clarity
Reviewed on Aug 05, 2026
Review provided by G2
What do you like best about the product?
I like how Databricks merges our warehousing and lakes into one architecture, eliminating the need to maintain separate storage and compute for operational reporting and predictive work on sensor data. This unified source of truth helps keep data duplication and sprawl in check while focusing on the operational picture. The Lakehouse concept of bringing warehousing and lake environments under one roof won me over. Unity Catalog simplifies governance by providing a clear view of data access, and Delta Live Tables efficiently manages our telemetry pipelines, reducing the effort needed for ensuring their reliability. I also find the initial setup quite easy with no bigger issues encountered.
What do you dislike about the product?
My one reservation is the cost side. The DBU model can climb quickly, and if cluster policies and auto-scaling limits go unattended, spend has a way of getting ahead of you. We saw this play out during an early stretch of heavy streaming, which pushed us to tighten our boundaries and keep a firm hand on utilization. Budget discipline rewards constant attention rather than a set-and-forget approach. I'd want clearer, real-time visibility into DBU spend at the cluster and workload level, with configurable alerts that flag when a job or fleet's consumption starts trending past its expected envelope before the bill lands, so budget governance becomes proactive rather than a monthly reconciliation exercise.
What problems is the product solving and how is that benefiting you?
Databricks ends siloization by merging warehousing and lakes, so I manage data from one governed source of truth. It allows me to scale telemetry across client fleets and focus on operations without reconciling parallel systems.
Nahid S.
Streamlined Data Management with a Learning Curve
Reviewed on Aug 04, 2026
Review provided by G2
What do you like best about the product?
I really appreciate that Databricks doesn't force a trade-off between control and convenience. My developers can access the cluster configuration for detailed tuning when necessary, but for routine tasks, the managed notebooks, job scheduling, and Unity Catalog take care of everything without needing a platform specialist. I like that Databricks borrows from software engineering practices instead of sticking to the classic analytics silo. The integration with version control, CI-friendly job definitions, and separated environments mean our data work follows the same review and release processes as the rest of our codebase. This consistency makes it easier to manage as the team grows.
What do you dislike about the product?
I find the on-ramp steeper than it should be for people who aren't already fluent in the ecosystem. When someone on my team picks it up for the first time, there's a long stretch where they're productive with the UI but not with what's happening underneath. Spark is the main pain point. When a job degrades or fails, the reasons are rarely obvious, and tracing a problem back to a skewed join or a bad shuffle usually eats far more of a senior developer's attention than it deserves. Richer query plan visualization, with plain-language explanations of why the optimizer made a given choice, would remove a lot of that guesswork and stop these investigations from landing on the same two or three people every time.
What problems is the product solving and how is that benefiting you?
Databricks provides a governed data repository, making pipeline changes reviewable in pull requests. I can plan data work in sprints alongside engineering tasks, eliminating separation between data and software development.
aravind k.
Reliable Platform for Building Scalable Data Pipelines
Reviewed on Aug 03, 2026
Review provided by G2
What do you like best about the product?
What I like most about Databricks is how its user friendly UI brings the entire data engineering workflow into one platform. In my project, data lands in the Bronze layer through snaplogic, and we use databricks to transform it into Silver and Gold data products. The serverless compute option has significantly reduced the effort of managing infrastructure, while Unity Catalog makes governance and access control straightforward. I also find AI Genie and the built-in monitoring features useful for investigating pipeline failures, checking job runtimes, and debugging issues much faster.From a pricing perspective, the pay-for-what-you-use model works well for us, especially with serverless compute and auto-scaling, as we avoid paying for idle infrastructure. The documentation is comprehensive, onboarding new team members is relatively straightforward, and there is a strong knowledge base and community that helps resolve issues quickly. Overall, it has made our ETL development and day 2 day operations more efficient.
What do you dislike about the product?
One downside is that troubleshooting can sometimes be difficult when a pipeline fails, as the error messages aren't always detailed enough and you may need to dig through multiple logs to find the root cause. The platform also has a lot of features, so it can take some time for new users to become comfortable with everything. While serverless is convenient, costs can increase if compute usage isn't monitored properly. I'd also like to see faster UI responsiveness in some areas, especially when navigating large job histories or catalogs.
What problems is the product solving and how is that benefiting you?
Databricks has helped me and my team simplify our data engineering workflow by providing a single platform for data ingestion, transformation, governance, and analytics. We use it to process data from Bronze to Silver and Gold layers, which has made our ETL pipelines more reliable and easier to manage. Features like serverless compute, Unity Catalog, and built-in monitoring have reduced operational effort, improved collaboration across teams, and helped us deliver data products faster.
Transportation/Trucking/Railroad
Scalable, powerful, and full of integrations — ready for AI
Reviewed on Jul 31, 2026
Review provided by G2
What do you like best about the product?
Ease of working with data, with scalability and high processing power. It has many integrations, which facilitates access to data from different sources. Furthermore, it is natively prepared for AI.
What do you dislike about the product?
The interface is technical and the price is steep for some products.
What problems is the product solving and how is that benefiting you?
Democratize the data in different areas and decentralize the processing, making access broader and distributing activities better.
Diana C.
Databricks Streamlines ETL and Analytics with Scalable Notebooks
Reviewed on Jul 29, 2026
Review provided by G2
What do you like best about the product?
I've been using Databricks as part of our data engineering workflow to build and maintain ETL pipelines, analyze large datasets, and support reporting requirements. One of the things I like most is that it brings data engineering, analytics, and notebooks into a single workspace. Instead of switching between multiple tools, I can write PySpark code, validate transformations, collaborate with teammates, and schedule jobs from the same platform. This has made day-to-day development more organized, especially when working on multiple data pipelines.
Another feature I rely on frequently is the notebook environment. It's convenient for developing and testing transformations before moving them into production. During development, I often use notebooks to inspect sample data, troubleshoot failed transformations, and validate business logic with SQL and PySpark. The ability to mix code, markdown documentation, and query results in one place also makes it easier for team members to understand the implementation during code reviews or knowledge transfer sessions.
I also appreciate the platform's scalability. Some of our data processing jobs involve millions of records, and Databricks handles distributed processing efficiently without requiring us to manage the underlying infrastructure directly. Features like cluster management, job scheduling, and integration with cloud storage reduce operational overhead. That said, cluster startup times can occasionally delay quick debugging sessions, and managing compute resources carefully is important to avoid unnecessary costs. Overall, Databricks has helped simplify large scale data processing while giving enough flexibility for both development and production workloads.
Another feature I rely on frequently is the notebook environment. It's convenient for developing and testing transformations before moving them into production. During development, I often use notebooks to inspect sample data, troubleshoot failed transformations, and validate business logic with SQL and PySpark. The ability to mix code, markdown documentation, and query results in one place also makes it easier for team members to understand the implementation during code reviews or knowledge transfer sessions.
I also appreciate the platform's scalability. Some of our data processing jobs involve millions of records, and Databricks handles distributed processing efficiently without requiring us to manage the underlying infrastructure directly. Features like cluster management, job scheduling, and integration with cloud storage reduce operational overhead. That said, cluster startup times can occasionally delay quick debugging sessions, and managing compute resources carefully is important to avoid unnecessary costs. Overall, Databricks has helped simplify large scale data processing while giving enough flexibility for both development and production workloads.
What do you dislike about the product?
While Databricks has been reliable for our data engineering workloads, there are a few areas where I think it could be improved. One challenge I've experienced is cluster startup time. When I only need to test a small code change or validate a transformation, waiting for a cluster to start can interrupt the development flow. It's not a major issue for scheduled production jobs, but during active development and debugging, those extra minutes add up.
Another limitation is cost management. Since compute resources are tied to cluster usage, it's important to monitor cluster configurations and ensure they are shut down when not needed. We've had situations where development clusters remained active longer than expected, resulting in higher cloud costs. The platform provides tools to manage this, but it still requires teams to establish good governance and usage policies. I also found that some configuration settings for jobs, permissions, and clusters have a learning curve, especially for new team members who are unfamiliar with the Databricks environment.
From a day to day perspective, debugging distributed Spark jobs can sometimes be challenging. While the logs provide useful information, identifying the exact cause of failures often requires navigating through multiple execution logs and Spark UI details. For straightforward issues this isn't a problem, but troubleshooting more complex pipeline failures can take time. Despite these limitations, none of them outweigh the benefits of the platform, and most challenges can be managed with proper cluster configuration, monitoring, and team practices.
Another limitation is cost management. Since compute resources are tied to cluster usage, it's important to monitor cluster configurations and ensure they are shut down when not needed. We've had situations where development clusters remained active longer than expected, resulting in higher cloud costs. The platform provides tools to manage this, but it still requires teams to establish good governance and usage policies. I also found that some configuration settings for jobs, permissions, and clusters have a learning curve, especially for new team members who are unfamiliar with the Databricks environment.
From a day to day perspective, debugging distributed Spark jobs can sometimes be challenging. While the logs provide useful information, identifying the exact cause of failures often requires navigating through multiple execution logs and Spark UI details. For straightforward issues this isn't a problem, but troubleshooting more complex pipeline failures can take time. Despite these limitations, none of them outweigh the benefits of the platform, and most challenges can be managed with proper cluster configuration, monitoring, and team practices.
What problems is the product solving and how is that benefiting you?
Databricks has helped address one of the biggest challenges in our data engineering workflow: processing and transforming large volumes of data efficiently. Before the data reaches reporting or downstream applications, we need to ingest data from multiple sources, apply business rules, clean inconsistent records, and create curated datasets. Databricks provides a single platform where we can develop, test, and run these data pipelines using PySpark and SQL instead of managing multiple disconnected tools. This has made our development process more consistent and easier to maintain.
A practical example is one of our daily ETL pipelines that processes data from different source systems before loading it into curated tables for reporting. We use Databricks notebooks during development to validate transformations on sample data and then schedule the same logic as production jobs. If a pipeline fails, the job history and execution logs help us identify the stage where the failure occurred, making troubleshooting more efficient than manually tracing scripts across different servers. Having notebooks, job scheduling, and cluster management in one platform has reduced the effort required to manage these workflows.
From a business perspective, the biggest benefit is faster availability of reliable data for reporting and analytics. Our team spends less time managing infrastructure and more time implementing business logic and improving data quality. While optimizing Spark jobs and monitoring cluster costs still require attention, Databricks has streamlined our daily workflow by providing a scalable environment for developing, testing, and running data pipelines. This has improved collaboration within the team and made it easier to deliver data that downstream users can trust.
A practical example is one of our daily ETL pipelines that processes data from different source systems before loading it into curated tables for reporting. We use Databricks notebooks during development to validate transformations on sample data and then schedule the same logic as production jobs. If a pipeline fails, the job history and execution logs help us identify the stage where the failure occurred, making troubleshooting more efficient than manually tracing scripts across different servers. Having notebooks, job scheduling, and cluster management in one platform has reduced the effort required to manage these workflows.
From a business perspective, the biggest benefit is faster availability of reliable data for reporting and analytics. Our team spends less time managing infrastructure and more time implementing business logic and improving data quality. While optimizing Spark jobs and monitoring cluster costs still require attention, Databricks has streamlined our daily workflow by providing a scalable environment for developing, testing, and running data pipelines. This has improved collaboration within the team and made it easier to deliver data that downstream users can trust.
Gambling & Casinos
Feature-Rich, Intuitive UI with Great AI Assistance and Easy Integrations
Reviewed on Jul 29, 2026
Review provided by G2
What do you like best about the product?
many features with intuitive ui, great ai assistance and you easily integrate with other services. I like the facts that yo can orchestrate jobs and notebooks easily, run queries on large volumes of data and the fact that the AI genie can help you a lot understand and implement better and faster tasks.
What do you dislike about the product?
sometimes the documentation is not easily obtainable. I have also come across cases where the catalogue quick search did not yield my table, but the table existed. generally I am happy
What problems is the product solving and how is that benefiting you?
doing large scale queries, making it fast and easy to generate reports and dashboards for both technical and non technical people. bridges the gap between tech and product/business
Reetika P.
Easy API Data Pulls and Collection Management, Plus AI-Powered Coding
Reviewed on Jul 28, 2026
Review provided by G2
What do you like best about the product?
It easily pulls data from the API, and within the same dataset we can manage our collections. We also have the option to write code using the AI.
What do you dislike about the product?
In our current setup, BigQuery SQL queries run with predictable costs that are easy to control. With Databricks, though, if a data engineer spins up an oversized cluster or leaves a node running after processing dealer posts or telematics logs, compute costs can ramp up quickly and may go unnoticed.
What problems is the product solving and how is that benefiting you?
For our projects, we use it to pull the source, or raw, data from the APIs and then transfer that same data into BigQuery. It essentially acts as a middleman for us.
jimena m.
Multiservice platform
Reviewed on Jul 24, 2026
Review provided by G2
What do you like best about the product?
I usually use it to make integrations between different data sources.
What do you dislike about the product?
I don't like the Genie Code, it's not good and I prefer to rely on other AI tools.
What problems is the product solving and how is that benefiting you?
The main issues we have solved so far are that we have migrated several workflows from another tool to Databricks, and the execution time has decreased considerably.