Together AI
Open-source cloud platform providing API access and collaborative training infrastructure for foundation models.
Mission
To run and train the world's most capable open models with shared compute.
Founded By
Key Products & Research
- Together API
- RedPajama (dataset)
- FlashAttention (co-developed)
Headquarters
San Francisco, California, USA
Founded
2022
Status
Active
Contribution to AI
LLaMA arrived in early 2023 as a description rather than an artefact: gated weights, and a training corpus specified only as a table of proportions in a paper. RedPajama rebuilt it — 1.2 trillion tokens assembled to that published recipe across seven slices, from Common Crawl and C4 down to arXiv and StackExchange, with the collection scripts released under Apache 2.0 so the recipe could be re-run rather than merely trusted. The corpus passed 190,000 downloads and turned up in the pretraining of Snowflake's Arctic, Salesforce's XGen and AI2's OLMo. The November 2023 successor changed the object again: 30 trillion tokens of web text in five languages, shipped raw with more than forty pre-computed quality signals attached, so filtering became a decision the downstream trainer makes and can document instead of one baked in by whoever built the corpus. The accompanying 3B and 7B models were trained on 3,072 V100s on Oak Ridge's Summit under a DOE INCITE allocation, showing that public supercomputing time could yield openly licensed foundation models. The FlashAttention line grew up in the same place — exact attention computed blockwise, memory linear rather than quadratic in sequence length, the second version reaching roughly 72 percent model FLOPs utilisation on A100s. Long context stopped being a luxury priced in hardware.
Drafted with AI and edited by hand (claude-opus-5, reviewed 2026-08).