LAION
German non-profit that released the massive open image-text datasets used to train Stable Diffusion and other generative models.
Mission
To make large-scale machine learning models, datasets, and code available to the general public.
Founded By
Key Products & Research
- LAION-400M dataset
- LAION-5B dataset
- OpenCLIP
Headquarters
Hamburg, Germany
Founded
2021
Status
Active
Contribution to AI
By 2021 CLIP and DALL·E had shown that noisy image-text pairs scraped from the web were enough to teach a model broad visual competence, but the corpora behind both were private, so no outside group could check a claim, vary an ingredient, or train anything comparable. The German association that formed that year answered with 400 million pairs, then with 5.85 billion, of which 2.32 billion were English — released not as images but as URLs with metadata, a design choice that made web-scale distribution practical and later became the question a Hamburg court answered in 2024 when it found the scraping covered by the text-and-data-mining exception for scientific research. What followed depended entirely on the corpus being public. OpenCLIP reproduced CLIP openly and then went past it, with a ViT-G/14 crossing 80 per cent zero-shot accuracy on ImageNet; the accompanying scaling study found that OpenAI's models and theirs scaled differently despite identical architectures, isolating training data as a variable rather than a constant. Stable Diffusion was trained on a filtered slice of the same collection, which is why image generation arrived as downloadable weights instead of an API. The 2023 discovery of abuse-material links, and the audited re-release that removed them, demonstrated something no closed corpus can: a dataset outsiders can inspect is a dataset that can be corrected.
Drafted with AI and edited by hand (claude-opus-5, reviewed 2026-08).