Overview
IBM watsonx.data as a Service is an open, hybrid-cloud data lakehouse on AWS that combines lakehouse storage with integrated data fabric capabilities for governance, lineage, and data quality. Using open formats such as Apache Iceberg and Parquet, and engines including Presto SQL and Apache Spark, the platform provides governed access to structured, semi-structured, and unstructured data across hybrid, multi-cloud, and on-premises environments.
watsonx.data is GenAI-ready, automating ingestion, preparation, and retrieval of unstructured data to fuel accurate generative AI. With vector search and multi-model capabilities through Cassandra (Astra DB) and Milvus, watsonx.data supports advanced RAG, similarity search, and real-time operational workloads. Internal testing shows improved accuracy over vector-only RAG by leveraging retrieval governance and integrated metadata.
watsonx.data offers enterprise-grade deployment flexibility and security, including VPC-based deployments, AWS PrivateLink, and support for FedRAMP (Medium) and HIPPA for AWS GovCloud. Native AWS integrations, such as AWS Lake Formation and the Common Policy Gateway (CPG) for unified access control, enable real-time policy synchronization and full auditability. With multi-engine optimization across Presto and Spark, organizations can reduce data warehouse costs while scaling analytics and AI across their AWS footprint.
Q: How does watsonx.data integrate with AWS-native services?
The platform integrates with AWS Lake Formation for access management and metadata alignment, supports AWS PrivateLink for secure connectivity, and uses the Common Policy Gateway (CPG) for unified access control with real-time policy synchronization and full audit tracking.
Q: What security and compliance capabilities are available?
watsonx.data offers enterprise-grade deployment flexibility and security, including VPC-based deployments, AWS PrivateLink, and support for FedRAMP (Medium) and HIPPA for AWS GovCloud. to support regulated workloads.
Q: What deployment options does watsonx.data support?
IBM watsonx.data supports SaaS on AWS, in-customer VPC deployments on AWS and Azure, multi-cloud architectures, and on-premises deployments on Red Hat OpenShift. On-premises deployments can take advantage of existing IBM Power and IBM Fusion HCI environments to deliver optimized performance, while maintaining flexibility for data residency, security, and compliance requirements.
Q: How does watsonx.data improve GenAI and RAG accuracy?
watsonx.data enhances generative AI results by combining governed retrieval with integrated vector databases such as Milvus and Cassandra (Astra DB), enabling fusion of unstructured, structured, and metadata-rich context. Internal testing shows higher answer correctness compared to vector-only RAG by applying data fabric governance and optimized retrieval strategies.
Highlights
- Unify hybrid-cloud analytics through a single entry point: Access all enterprise data across AWS, on-premises, and multi-cloud environments through a shared metadata layer that supports open table formats such as Apache Iceberg and Parquet, enabling consistent analytics and governance without ETL.
- Deploy and connect to AWS data sources in minutes: Begin querying data quickly by connecting AWS storage (e.g. Amazon S3) and analytics environments - including Db2 Warehouse on AWS and Netezza on AWS - within minutes, supported by built-in governance, security automation, and multi-engine execution through Presto and Spark.
- Reduce the cost of your data warehouse by up to 50% through workload optimization: Lower analytics spend by offloading and optimizing workloads across fit-for-purpose engines (Presto, Spark) and storage tiers, enabling measurable cost reductions of up to 50% when augmenting traditional warehouse workloads.
Details
Introducing multi-product solutions
You can now purchase comprehensive solutions tailored to use cases and industries.
Features and programs
Buyer guide

Financing for AWS Marketplace purchases
Pricing
Dimension | Description | Cost/12 months |
|---|---|---|
Extra-small Watsonx.data installation | Watsonx.data Resource Units annual Contract "pack" of 2000 Resource Units | $2,000.00 |
Small Watsonx.data installation | Watsonx.data Resource Units annual Contract "pack" of 20000 Resource Units | $20,000.00 |
Medium Watsonx.data installation | Watsonx.data Resource Units annual Contract "pack" of 50000 Resource Units | $50,000.00 |
Large Watsonx.data installation | Watsonx.data Resource Units annual Contract "pack" of 100000 Resource Units | $100,000.00 |
The following dimensions are not included in the contract terms, which will be charged based on your usage.
Dimension | Cost/unit |
|---|---|
Overage charge for overconsumption of contracted resource units | $1.10 |
Vendor refund policy
All orders are non-cancellable and all fees and other amounts that you pay are non-refundable.
Custom pricing options
How can we make this page better?
Legal
Vendor terms and conditions
Content disclaimer
Delivery details
Software as a Service (SaaS)
SaaS delivers cloud-based software applications directly to customers over the internet. You can access these applications through a subscription model. You will pay recurring monthly usage fees through your AWS bill, while AWS handles deployment and infrastructure management, ensuring scalability, reliability, and seamless integration with other AWS services.
Resources
Support
Vendor support
This product includes enterprise-grade support designed for fast deployment and low operational risk. Customers have access to comprehensive public documentation, step-by-step integration guides, and architecture references aligned with AWS best practices. Technical support is available through defined support channels with documented SLAs, and our team actively assists with onboarding, configuration, and troubleshooting.
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.


Standard contract
Customer reviews
Flexible and Scalable Data Platform for Analytics and AI
Clean, Smooth UI with Excellent Onboarding and Infrastructure Visuals
One feature I liked was the infrastructure section. It provides a visual interface that feels similar to tools like n8n,
The storage integration experience is also well designed. It supports connecting to services like Amazon S3, Redis, PostgreSQL, MySQL, and other data sources from the interface.
In free mode i not able to add components but it was good and performance was aslo good every click feels smooth
Another area that could be improved is the Query Workspace. While it's functional, the interface feels a bit too compact, especially on smaller screens. More spacing and a cleaner layout would make writing and reviewing queries more comfortable.
Clean, Unobtrusive UI with Seamless Integrations and On-Demand AI Insights
Seamless Data Integration with Stellar Performance
Great Platform for Unified Data and Analytics
> What I like best about IBM watsonx.data is its ability to manage and analyze large volumes of structured and unstructured data efficiently. Its open data lakehouse architecture, scalability, and support for AI and analytics make it a powerful platform for modern data-driven applications.
> One drawback of IBM watsonx.data is that the initial setup and configuration can be complex for new users. Some advanced features also have a learning curve, and performance tuning may require technical expertise to get the best results.
> IBM watsonx.data helps solve the challenge of managing and analyzing large volumes of data from multiple sources in one platform. It improves query performance, reduces data management complexity, and supports AI and analytics workloads, enabling faster insights and more efficient decision-making.