Skip to content
Databricks logo

Databricks

Data Infrastructure · www.databricks.com

Is this your tool? Claim this listing →

Overview

Databricks is a cloud-based data and AI platform built by the original creators of Apache Spark, Delta Lake, MLflow, and Unity Catalog. It combines data engineering, SQL analytics, machine learning, and generative AI on a single "lakehouse" architecture that runs on AWS, Azure, or GCP.

The problem Databricks solves

Data teams building analytics and AI products typically end up assembling a separate data warehouse, a Spark cluster, an ML platform, and a BI connector layer, each with its own security model and copies of the data. Databricks solves that by putting data engineering pipelines, SQL analytics, governance, and ML/AI model training and serving on one lakehouse, so teams work against a single governed copy of the data instead of reconciling it across systems.

Decision context

Use these points to test whether the product fits your operation, not just whether it has a long feature list.

  • Published starting price: Free. Confirm user, usage, and feature limits for the plan you would actually buy.
  • Deployment: cloud. Check security, data-residency, and access requirements for every team that will use it.
  • Verified integrations include Fivetran, dbt, Alation, Microsoft Power BI, Tableau, Rivery. Validate sync direction and plan limits for the connections that matter.
  • This record was last checked on 7/31/2026; pricing and features can change.

How to evaluate Databricks

A listing helps create a shortlist; a trial with the team’s real workflow decides whether the tool fits. Use this reading with the structured facts and confirm changes with the vendor.

Workflow fit

The record describes it as a fit for Data engineering and data platform teams building large-scale ETL/ELT pipelines, Organizations wanting a single governed platform for both BI/SQL analytics and ML/AI, Enterprises with dedicated cloud/data infrastructure teams to manage usage-based costs. Check that this context matches the volume, roles, and processes your team needs it to support.

Pilot questions

  • Can Databricks complete the critical workflow without manual work outside the product?
  • Do the recorded connections (Fivetran, dbt, Alation, Microsoft Power BI) support the sync direction, permissions, and volume we need?
  • What user, usage, storage, support, or security limits appear after the headline starting price?

Evidence and freshness

This record was checked on 7/31/2026. That date tells you when the record was reviewed, not that the vendor has left its terms unchanged since then.

Best for

  • Data engineering and data platform teams building large-scale ETL/ELT pipelines
  • Organizations wanting a single governed platform for both BI/SQL analytics and ML/AI
  • Enterprises with dedicated cloud/data infrastructure teams to manage usage-based costs

Not a fit if

  • Small teams or startups without data engineering expertise or budget to monitor consumption-based billing
  • Companies wanting a simple, flat-fee, fully self-hosted/on-premises deployment

Why it’s listed

  • Category-defining lakehouse platform combining data engineering, warehousing, and AI/ML in one system
  • Built by the original creators of Apache Spark, Delta Lake, MLflow, and Unity Catalog
  • Used by 20,000+ organizations including a majority of the Fortune 500

Pricing

Free Edition

Free

Free access to core Databricks tools for individual learning, prototyping, and small-scale exploration.

  • Notebooks (SQL/Python/R)
  • Limited serverless compute
  • Unity Catalog basics
  • Community support

Standard compute (Jobs)

$0 per DBU

Pay-as-you-go pricing for automated Jobs compute, the lowest-cost general-purpose workload type.

  • Automated/scheduled workloads
  • Per-second billing
  • Autoscaling clusters
  • Delta Lake storage

SQL Serverless

$1 per DBU

Serverless SQL warehouses for BI dashboards and ad hoc analytics.

  • Serverless SQL warehouses
  • Instant elastic scaling
  • BI tool connectors
  • Query result caching

Enterprise

$0 per DBU (Jobs, Enterprise tier)

Adds enhanced security, compliance controls, and premium support on top of usage-based compute pricing.

  • Unity Catalog governance
  • HIPAA/PCI-eligible workspaces
  • Private connectivity (PrivateLink)
  • Premium support SLAs
  • Committed-use discounts

Features

Workflow automationDatabricks Workflows/Lakeflow Jobs natively orchestrate notebooks, pipelines, and ML tasks with dependencies and retries.
Pre-built connectorsNative ingestion via Auto Loader plus Partner Connect integrations, but not as broad a native connector catalog as dedicated iPaaS/ELT tools.
Data pipelines / ETLDelta Live Tables / Lakeflow Declarative Pipelines is a core product for building and managing ETL pipelines.
Scheduling & triggersWorkflows support cron scheduling, file-arrival triggers, and event-based triggers.
Data governance & lineageUnity Catalog provides centralized governance, lineage, and cataloging across all workspaces.
Role-based access controlUnity Catalog supports role-based, row/column-level, and attribute-based (ABAC) access control.
AI featuresMosaic AI, Databricks Assistant, AI/BI Genie, and Model Serving provide built-in generative AI and ML tooling.
Public APIExtensive REST APIs are documented for workspaces, jobs, clusters, and Unity Catalog objects.
Self-hosting / on-premSaaS-only; compute runs inside the customer's cloud account, but there is no on-premises or self-hosted deployment option.

Integrations

FivetrandbtAlationMicrosoft Power BITableauRiveryLabelboxProphecyArcionConfluent (Apache Kafka)LookerQlikApache AirflowCollibraInformatica

Security & compliance

SOC 2 Type IIISO 27001HIPAAPCI DSSGDPR

Pros & cons

Pros

  • Strong at unifying data engineering, data science, and BI/analytics under one governed platform
  • Delta Lake and Unity Catalog provide robust data reliability, lineage, and access control
  • Scales well for very large datasets and complex ML/AI workloads
  • Native integrations across major BI tools (Tableau, Power BI, Looker) and ingestion tools (Fivetran, dbt)

Cons

  • Steep learning curve, especially for teams new to Spark or lakehouse concepts
  • Usage-based DBU pricing is complex and can scale unpredictably
  • Non-technical users find the interface and setup less approachable than simpler BI tools
  • Managing and forecasting compute cost requires ongoing FinOps discipline

What we found

4.6/5
658 reviews aggregatedLast checked 2026-07-31

Reviewers praise Databricks for unifying data engineering, analytics, and ML/AI workflows, while citing a steep learning curve and unpredictable usage-based costs as the main drawbacks.

Ratings and review counts come from public review platforms. We link to the original source and keep the underlying review text out of this profile.

User reviews

Written by Audyense accounts · moderated before publishing

No user reviews yet.

Used Databricks? Be the first to tell other buyers what actually worked.