Databricks
Built the dominant enterprise lakehouse platform for data engineering and large-scale machine learning.
Mission
To simplify and democratize data and AI on any cloud.
Founded By
Key Products & Research
- Databricks Lakehouse Platform
- MLflow (open-source)
- Apache Spark (originated)
Headquarters
San Francisco, California, USA
Founded
2013
Status
Active
Contribution to AI
Iterative algorithms were what MapReduce did worst: every pass over the data went back to disk, so training a model across a cluster meant fighting the framework. Spark, begun at Berkeley's AMPLab in 2009 and handed to the Apache Software Foundation in 2013, held working sets in memory between passes and turned distributed model fitting into an ordinary operation, with MLlib shipping the algorithms in the same box. The company built around that engine then went after the unstandardised parts of the pipeline. MLflow, released openly in 2018, took experiment tracking, model packaging and a versioned registry — things every team improvised badly and separately — and gave them one interface, which is now how a great many organisations record a trained model and move it into production. Delta Lake, opened in 2019, wrote a transaction log over Parquet files so that a data lake could give ACID guarantees and serve batch and streaming work from a single copy; the 2021 lakehouse paper named the architecture that resulted and argued that warehouse and lake need not be two systems with an export step between them. Dolly 2.0 in April 2023 was a smaller but pointed contribution: a 12-billion-parameter model tuned on fifteen thousand instruction records written by employees and released under a permissive licence, the first such dataset carrying no terms inherited from a commercial API. DBRX followed in 2024.
Drafted with AI and edited by hand (claude-opus-5, reviewed 2026-08).