Skip to content
IBM watsonx.data logo

IBM watsonx.data

Data Infrastructure · www.ibm.com/products/watsonx-data

Is this your tool? Claim this listing →

Overview

IBM watsonx.data is a hybrid, open data lakehouse that lets organizations store structured, semi-structured, and unstructured data once and query it with multiple fit-for-purpose engines (Presto, Spark, Db2, Netezza) instead of copying it between separate warehouses and lakes. It layers built-in governance and cataloging over open table formats (Iceberg, Delta Lake, Hudi) so the same governed data can feed BI reporting and AI/LLM workloads. It is available as fully managed SaaS on IBM Cloud and AWS, as self-managed software, or as a client-managed VPC (BYOC) deployment.

The problem IBM watsonx.data solves

Enterprises with data scattered across warehouses, data lakes, and multiple clouds usually end up copying data between systems just to run BI queries or train AI models, which duplicates storage costs and creates governance gaps. watsonx.data solves that by layering multiple open-format query engines over the same governed lakehouse storage, so teams query and govern data where it already sits instead of migrating it first.

Decision context

Use these points to test whether the product fits your operation, not just whether it has a long feature list.

  • Published starting price: Custom pricing. Confirm user, usage, and feature limits for the plan you would actually buy.
  • Deployment: cloud, on_premise, hybrid. Check security, data-residency, and access requirements for every team that will use it.
  • Verified integrations include Presto (query engine), Apache Spark, IBM Db2, IBM Netezza, Apache Iceberg, Delta Lake. Validate sync direction and plan limits for the connections that matter.
  • This record was last checked on 7/31/2026; pricing and features can change.

How to evaluate IBM watsonx.data

A listing helps create a shortlist; a trial with the team’s real workflow decides whether the tool fits. Use this reading with the structured facts and confirm changes with the vendor.

Workflow fit

The record describes it as a fit for Enterprises running hybrid or multi-cloud data estates who want to query data in place instead of migrating it between systems, Data/platform teams that need to feed the same governed dataset into both BI reporting and AI/LLM pipelines, Organizations trying to offload analytics workloads from an expensive data warehouse to cut storage/compute costs. Check that this context matches the volume, roles, and processes your team needs it to support.

Pilot questions

  • Can IBM watsonx.data complete the critical workflow without manual work outside the product?
  • Do the recorded connections (Presto (query engine), Apache Spark, IBM Db2, IBM Netezza) support the sync direction, permissions, and volume we need?
  • What user, usage, storage, support, or security limits appear after the headline starting price?

Evidence and freshness

This record was checked on 7/31/2026. That date tells you when the record was reviewed, not that the vendor has left its terms unchanged since then.

Best for

  • Enterprises running hybrid or multi-cloud data estates who want to query data in place instead of migrating it between systems
  • Data/platform teams that need to feed the same governed dataset into both BI reporting and AI/LLM pipelines
  • Organizations trying to offload analytics workloads from an expensive data warehouse to cut storage/compute costs

Not a fit if

  • Small teams or startups without dedicated data engineering resources to manage multi-engine setup and tuning
  • Companies wanting a simple, single-engine, minimal-configuration analytics tool rather than a hybrid lakehouse platform

Why it’s listed

  • Widely used open data lakehouse product from a major enterprise vendor (IBM), positioned as core Data Infrastructure content
  • Offers a genuine SaaS deployment path alongside self-managed and hybrid options
  • Actively developed with frequent releases and public roadmap/feedback channels

Pricing

Medium (Balanced)

Custom pricing

General-purpose compute instance sized for moderate query workloads.

  • Balanced compute/storage ratio
  • Scales independently of storage
  • Pay-as-you-go consumption billing

Medium (Storage Optimized)

Custom pricing

Medium compute tier tuned for storage-heavy workloads.

  • Higher storage throughput
  • Consumption-based RU billing

Large (Balanced)

Custom pricing

Larger compute instance for higher-concurrency analytics and AI workloads.

  • Higher query concurrency
  • Balanced compute/storage

Large (Storage Optimized)

Custom pricing

Largest published tier, optimized for heavy storage and big-data workloads.

  • Highest storage throughput tier
  • Designed for large-scale lakehouse workloads

Features

AI featuresBuilt to make governed data directly accessible to AI/LLM workloads (integrates with watsonx.ai) alongside BI
Data governance & lineageBuilt-in governance via IBM Knowledge Catalog and Apache Ranger integration
Data pipelines / ETLSupports data movement/ELT via Spark and connectors, not a dedicated pipeline-orchestration product
Pre-built connectorsPresto/Spark connectors to Db2, Netezza, Kafka, MySQL, PostgreSQL, MongoDB, S3, and more
Public APIREST APIs documented for provisioning, catalog, and query operations
Role-based access controlRole-based access control delivered through the integrated governance/catalog layer
Scheduling & triggersAchievable via Spark jobs and external orchestration, no native built-in scheduler emphasized
Self-hosting / on-premAvailable as self-managed software on-premises or as client-managed VPC (BYOC), in addition to SaaS
Workflow automationPrimarily API/pipeline-driven automation rather than a native workflow-builder

Integrations

Presto (query engine)Apache SparkIBM Db2IBM NetezzaApache IcebergDelta LakeApache HudiMilvus (vector database)DataStax CassandraApache KafkaMySQLPostgreSQLMongoDBAWS S3IBM Cloud Object StorageApache Ranger

Security & compliance

SOC 2ISO 27001

Pros & cons

Pros

  • Unifies data from multiple sources without complex migrations or duplication, per G2 reviewers
  • Open lakehouse architecture delivers strong performance for analytics, reporting, and AI workloads while remaining cost-efficient
  • Clean, organized UI/UX for navigating datasets, managing workloads, and monitoring operations
  • Flexibility to run different fit-for-purpose query engines depending on workload

Cons

  • Initial setup and configuration can feel complex for new users or smaller teams
  • Some integrations and advanced features have a learning curve, with documentation that could be clearer
  • Tuning performance or cost across multiple engines and hybrid environments requires more expertise than simpler platforms
  • TrustRadius reviewers note data import speed and clarity of pricing model as areas for improvement

What we found

4.3/5
93 reviews aggregatedLast checked 2026-07-31

Rated 4.3/5 on G2 (93 reviews) for open lakehouse architecture and cost/performance flexibility; TrustRadius reviewers echo strong hybrid/multi-cloud data integration but flag pricing clarity and UI usability.

Ratings and review counts come from public review platforms. We link to the original source and keep the underlying review text out of this profile.

User reviews

Written by Audyense accounts · moderated before publishing

No user reviews yet.

Used IBM watsonx.data? Be the first to tell other buyers what actually worked.