Scale AI
The dominant AI data labeling company powering virtually every major AI lab's training data pipeline.
Mission
To accelerate the development of AI applications through high-quality data.
Founded By
Key Products & Research
- Scale Data Engine
- Scale Donovan (Defense)
- RLHF Data pipelines
Headquarters
San Francisco, California, USA
Founded
2016
Status
Active
Contribution to AI
Training data was, for most of the field's history, something a research group produced grudgingly in-house or bought as piecework from Mechanical Turk with no guarantees attached. What changed after 2016 was that annotation became a specified, priced service with error rates you could hold a vendor to, delivered through an API — human judgement ordered the way compute was ordered. The first proof was in autonomous driving, where the hard part is not a bounding box on an image but a consistent object tracked across lidar sweeps, radar returns and six cameras at once. The nuScenes dataset, released in 2018 by nuTonomy with Scale's sensor-fusion annotation, put over a million 3D boxes into public hands and became a standard against which perception systems were measured. The same apparatus was redirected when reinforcement learning from human feedback became the default post-training step: preference comparisons and expert-written demonstrations turned into a purchasable commodity, which is why frontier labs could align models without building annotation organisations themselves. Two consequences outlasted the commercial story. The Washington Post's 2023 reporting on Remotasks workers in the Philippines forced the contractor workforce beneath model quality into open argument about pay and regulation. And Meta's $14.3 billion for a 49 per cent stake in June 2025, after which rival labs pulled work over confidentiality, priced the data supply chain as strategic territory rather than back-office cost.
Drafted with AI and edited by hand (claude-opus-5, reviewed 2026-08).