Listing Thumbnail

    IBM watsonx.data as a Service - GenAI Ready Data Lakehouse for AWS

     Info
    Deployed on AWS
    IBM watsonx.data is an open, hybrid data lakehouse with built-in data fabric and multi-engine optimization to prepare structured and unstructured data for AI.
    4.4

    Overview

    IBM watsonx.data as a Service is an open, hybrid-cloud data lakehouse on AWS that combines lakehouse storage with integrated data fabric capabilities for governance, lineage, and data quality. Using open formats such as Apache Iceberg and Parquet, and engines including Presto SQL and Apache Spark, the platform provides governed access to structured, semi-structured, and unstructured data across hybrid, multi-cloud, and on-premises environments.

    watsonx.data is GenAI-ready, automating ingestion, preparation, and retrieval of unstructured data to fuel accurate generative AI. With vector search and multi-model capabilities through Cassandra (Astra DB) and Milvus, watsonx.data supports advanced RAG, similarity search, and real-time operational workloads. Internal testing shows improved accuracy over vector-only RAG by leveraging retrieval governance and integrated metadata.

    watsonx.data offers enterprise-grade deployment flexibility and security, including VPC-based deployments, AWS PrivateLink, and support for FedRAMP (Medium) and HIPPA for AWS GovCloud. Native AWS integrations, such as AWS Lake Formation and the Common Policy Gateway (CPG) for unified access control, enable real-time policy synchronization and full auditability. With multi-engine optimization across Presto and Spark, organizations can reduce data warehouse costs while scaling analytics and AI across their AWS footprint.

    Q: How does watsonx.data integrate with AWS-native services?

    The platform integrates with AWS Lake Formation for access management and metadata alignment, supports AWS PrivateLink for secure connectivity, and uses the Common Policy Gateway (CPG) for unified access control with real-time policy synchronization and full audit tracking.

    Q: What security and compliance capabilities are available?

    watsonx.data offers enterprise-grade deployment flexibility and security, including VPC-based deployments, AWS PrivateLink, and support for FedRAMP (Medium) and HIPPA for AWS GovCloud. to support regulated workloads.

    Q: What deployment options does watsonx.data support?

    IBM watsonx.data supports SaaS on AWS, in-customer VPC deployments on AWS and Azure, multi-cloud architectures, and on-premises deployments on Red Hat OpenShift. On-premises deployments can take advantage of existing IBM Power and IBM Fusion HCI environments to deliver optimized performance, while maintaining flexibility for data residency, security, and compliance requirements.

    Q: How does watsonx.data improve GenAI and RAG accuracy?

    watsonx.data enhances generative AI results by combining governed retrieval with integrated vector databases such as Milvus and Cassandra (Astra DB), enabling fusion of unstructured, structured, and metadata-rich context. Internal testing shows higher answer correctness compared to vector-only RAG by applying data fabric governance and optimized retrieval strategies.

    Highlights

    • Unify hybrid-cloud analytics through a single entry point: Access all enterprise data across AWS, on-premises, and multi-cloud environments through a shared metadata layer that supports open table formats such as Apache Iceberg and Parquet, enabling consistent analytics and governance without ETL.
    • Deploy and connect to AWS data sources in minutes: Begin querying data quickly by connecting AWS storage (e.g. Amazon S3) and analytics environments - including Db2 Warehouse on AWS and Netezza on AWS - within minutes, supported by built-in governance, security automation, and multi-engine execution through Presto and Spark.
    • Reduce the cost of your data warehouse by up to 50% through workload optimization: Lower analytics spend by offloading and optimizing workloads across fit-for-purpose engines (Presto, Spark) and storage tiers, enabling measurable cost reductions of up to 50% when augmenting traditional warehouse workloads.

    Details

    Delivery method

    Deployed on AWS
    New

    Introducing multi-product solutions

    You can now purchase comprehensive solutions tailored to use cases and industries.

    Multi-product solutions

    Features and programs

    Buyer guide

    Gain valuable insights from real users who purchased this product, powered by PeerSpot.
    Buyer guide

    Financing for AWS Marketplace purchases

    AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
    Financing for AWS Marketplace purchases

    Pricing

    IBM watsonx.data as a Service - GenAI Ready Data Lakehouse for AWS

     Info
    Pricing is based on the duration and terms of your contract with the vendor, and additional usage. You pay upfront or in installments according to your contract terms with the vendor. This entitles you to a specified quantity of use for the contract duration. Usage-based pricing is in effect for overages or additional usage not covered in the contract. These charges are applied on top of the contract price. If you choose not to renew or replace your contract before the contract end date, access to your entitlements will expire.
    Additional AWS infrastructure costs may apply. Use the AWS Pricing Calculator  to estimate your infrastructure costs.

    12-month contract (4)

     Info
    Dimension
    Description
    Cost/12 months
    Extra-small Watsonx.data installation
    Watsonx.data Resource Units annual Contract "pack" of 2000 Resource Units
    $2,000.00
    Small Watsonx.data installation
    Watsonx.data Resource Units annual Contract "pack" of 20000 Resource Units
    $20,000.00
    Medium Watsonx.data installation
    Watsonx.data Resource Units annual Contract "pack" of 50000 Resource Units
    $50,000.00
    Large Watsonx.data installation
    Watsonx.data Resource Units annual Contract "pack" of 100000 Resource Units
    $100,000.00

    Additional usage costs (1)

     Info

    The following dimensions are not included in the contract terms, which will be charged based on your usage.

    Dimension
    Cost/unit
    Overage charge for overconsumption of contracted resource units
    $1.10

    Vendor refund policy

    All orders are non-cancellable and all fees and other amounts that you pay are non-refundable.

    Custom pricing options

    Request a private offer to receive a custom quote.

    How can we make this page better?

    Tell us how we can improve this page, or report an issue with this product.
    Tell us how we can improve this page, or report an issue with this product.

    Legal

    Vendor terms and conditions

    Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA) .

    Content disclaimer

    Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.

    Usage information

     Info

    Delivery details

    Software as a Service (SaaS)

    SaaS delivers cloud-based software applications directly to customers over the internet. You can access these applications through a subscription model. You will pay recurring monthly usage fees through your AWS bill, while AWS handles deployment and infrastructure management, ensuring scalability, reliability, and seamless integration with other AWS services.

    Support

    Vendor support

    This product includes enterprise-grade support designed for fast deployment and low operational risk. Customers have access to comprehensive public documentation, step-by-step integration guides, and architecture references aligned with AWS best practices. Technical support is available through defined support channels with documented SLAs, and our team actively assists with onboarding, configuration, and troubleshooting.

    AWS infrastructure support

    AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.

    Product comparison

     Info
    Updated weekly

    Accolades

     Info
    Top
    50
    In Data Warehouses
    Top
    10
    In Databases & Analytics Platforms, ML Solutions, Data Analytics
    Top
    10
    In Data Analysis

    Customer reviews

     Info
    Sentiment is AI generated from actual customer reviews on AWS and G2
    Reviews
    Functionality
    Ease of use
    Customer service
    Cost effectiveness
    Positive reviews
    Mixed reviews
    Negative reviews

    Overview

     Info
    AI generated from product descriptions
    Open Table Format Support
    Supports open table formats including Apache Iceberg and Parquet for consistent analytics and governance across hybrid-cloud environments without requiring ETL processes.
    Multi-Engine Query Optimization
    Provides multi-engine optimization across Presto SQL and Apache Spark to execute queries across structured, semi-structured, and unstructured data with workload-specific optimization.
    Vector Database Integration
    Integrates vector search and multi-model capabilities through Cassandra (Astra DB) and Milvus to support advanced retrieval-augmented generation (RAG), similarity search, and real-time operational workloads.
    Enterprise Security and Compliance
    Offers VPC-based deployments, AWS PrivateLink connectivity, and compliance support for FedRAMP (Medium) and HIPAA for AWS GovCloud environments.
    Unified Access Control and Governance
    Implements integrated data fabric with governance, lineage, and data quality capabilities, including AWS Lake Formation integration and Common Policy Gateway (CPG) for unified access control with real-time policy synchronization and audit tracking.
    Lakehouse Architecture
    Unified data foundation built on lakehouse architecture providing open, unified foundation for data and governance with support for open standards and formats
    Data Intelligence Engine
    Powered by Data Intelligence Engine that enables organization-wide access to data and insights across all users and roles
    Multi-Workload Unification
    Consolidates data engineering, analytics, business intelligence, data science and machine learning workloads on a single common platform
    Collaborative Development Environment
    Native collaboration capabilities enabling data teams to collaborate across entire data and AI workflow
    Open Source Foundation
    Built on open source data projects and open standards to maximize flexibility and interoperability with existing data ecosystems
    Workload Auto-scaling
    Intelligently autoscales workloads up and down across hybrid and public cloud environments for optimized cloud infrastructure utilization.
    Multi-function Analytics Platform
    Provides integrated data warehouse, machine learning, and custom analytics capabilities with unified analytic functions to eliminate data silos.
    Shared Data Experience (SDX)
    Implements security and governance policies that are set once and applied consistently across all data and workloads, with portability across supported infrastructures.
    Data Lifecycle Management
    Manages complete data lifecycle functions including ingestion, transformation, querying, optimization, and predictive analytics across multiple cloud environments.
    Unified Security and Governance
    Ensures all workloads share common security, governance, and metadata with capabilities for data discovery, curation, and self-service access controls.

    Contract

     Info
    Standard contract
    No
    No
    No

    Customer reviews

    Ratings and reviews

     Info
    4.4
    180 ratings
    5 star
    4 star
    3 star
    2 star
    1 star
    56%
    40%
    3%
    1%
    0%
    3 AWS reviews
    |
    177 external reviews
    External reviews are from G2  and PeerSpot .
    Nishant V.

    Flexible and Scalable Data Platform for Analytics and AI

    Reviewed on Aug 08, 2026
    Review provided by G2
    What do you like best about the product?
    What I like most about IBM watsonx.data is how it provides a flexible, scalable data platform that brings data from different sources together in one place. Its open architecture, integration with various data engines, and support for AI and analytics make it easier to manage, access, and use data efficiently.
    What do you dislike about the product?
    The main drawback is that IBM watsonx.data can have a fairly steep learning curve, particularly for users who are new to its architecture and ecosystem. It can also feel complex to configure and manage day to day, and the overall cost may be a concern for smaller organizations.
    What problems is the product solving and how is that benefiting you?
    IBM watsonx.data helps address the challenge of managing and accessing data across multiple sources and platforms. It offers a unified environment for data integration, governance, analytics, and AI workloads, which makes it easier to locate and use trusted data. For me, this means fewer data silos, simpler data access and management, and better efficiency when working on analytics and AI-related tasks.
    Aliasgar B.

    Clean, Smooth UI with Excellent Onboarding and Infrastructure Visuals

    Reviewed on Aug 04, 2026
    Review provided by G2
    What do you like best about the product?
    The UI is clean and easy to navigate, especially for someone using the platform for the first time. The getting started guides and onboarding flow helped me understand the different components without needing to spend much time reading documentation .
    One feature I liked was the infrastructure section. It provides a visual interface that feels similar to tools like n8n,
    The storage integration experience is also well designed. It supports connecting to services like Amazon S3, Redis, PostgreSQL, MySQL, and other data sources from the interface.
    In free mode i not able to add components but it was good and performance was aslo good every click feels smooth
    What do you dislike about the product?
    The biggest issue I encountered was around Spark engine management. When I tried stopping the Spark server from the infrastructure page, I repeatedly received error messages even after pausing the engine and related services. The error messages weren't very descriptive, so it was difficult to understand what was actually wrong or how to resolve it. Better diagnostics and more user-friendly error messages would improve the experience. I also noticed IBM documents several known Spark UI and engine-related limitations, so I hope these areas continue to improve.

    Another area that could be improved is the Query Workspace. While it's functional, the interface feels a bit too compact, especially on smaller screens. More spacing and a cleaner layout would make writing and reviewing queries more comfortable.
    What problems is the product solving and how is that benefiting you?
    As a student, I mostly used watsonx.data to learn. From what I understood, it's useful for companies that have data spread across different storage systems and want a single place to manage and query it, especially for AI and analytics use cases. It also helped me understand how enterprise data platforms work in practice.
    Eric B.

    Clean, Unobtrusive UI with Seamless Integrations and On-Demand AI Insights

    Reviewed on Jul 29, 2026
    Review provided by G2
    What do you like best about the product?
    I like that the UI stays out of the way, the integrations keep our data connected overall behind the scenes, and its most noticeable AI feature is there whenever I need an additional layer of insight.
    What do you dislike about the product?
    Well it wasn’t perfect from the start. AI occasionally requires a second thought before I move forward with its decisions. And it does demand some solid attention to make complete sense to us.
    What problems is the product solving and how is that benefiting you?
    We were putting too much effort into finding, preparing, and then validating data before making any analysis. Now that our data is synced with the best of the features, the entire process feels more connected, making it simpler for us to make informed decisions about data.
    SHIWAM T.

    Seamless Data Integration with Stellar Performance

    Reviewed on Jul 29, 2026
    Review provided by G2
    What do you like best about the product?
    I like how IBM watsonx.data unifies data from multiple sources into a single lakehouse platform while delivering fast query performance. Its strong data integration capabilities and open lakehouse architecture allow us to work with data in place instead of moving or duplicating it. I also appreciate that the platform scales well as our data grows, supports a wide range of analytics workloads, and integrates smoothly with AI business intelligence tools. The initial setup process was relatively straightforward, with well-documented installation and configuration steps, and connecting common data sources was uncomplicated.
    What do you dislike about the product?
    For me, everything is good.
    What problems is the product solving and how is that benefiting you?
    I use IBM watsonx.data to consolidate data from multiple sources into one platform, improving access and analysis. It eliminates silos and enhances query performance for large datasets, providing faster insights without data duplication.
    MOUNEES KUMAR C.

    Great Platform for Unified Data and Analytics

    Reviewed on Jul 27, 2026
    Review provided by G2
    What do you like best about the product?
    You can use this response (more than 40 characters):

    > What I like best about IBM watsonx.data is its ability to manage and analyze large volumes of structured and unstructured data efficiently. Its open data lakehouse architecture, scalability, and support for AI and analytics make it a powerful platform for modern data-driven applications.
    What do you dislike about the product?
    You can use this balanced review:

    > One drawback of IBM watsonx.data is that the initial setup and configuration can be complex for new users. Some advanced features also have a learning curve, and performance tuning may require technical expertise to get the best results.
    What problems is the product solving and how is that benefiting you?
    You can use this response:

    > IBM watsonx.data helps solve the challenge of managing and analyzing large volumes of data from multiple sources in one platform. It improves query performance, reduces data management complexity, and supports AI and analytics workloads, enabling faster insights and more efficient decision-making.
    View all reviews