Large OTS Catalog
Move faster with a deep off-the-shelf catalog spanning pre-training, post-training, and RL—including video, documents, and global web text beyond Common Crawl. We can deliver exabytes of data and 60T+ corpus-ready tokens.
Exabite (formerly Calaveras) provides pre-training and post-training data for frontier AI labs. Trillion-dollar corporations are training the most powerful AI models on earth with our data.
Schedule a chatTotal OTS data
Corpus-ready tokens (fully incremental to CommonCrawl, GitHub, and other common sources)
Supply of new novel datasets spanning pretraining, SFT, RL, and subject-specific data.
01 / WHAT WE PROVIDE
Move faster with a deep off-the-shelf catalog spanning pre-training, post-training, and RL—including video, documents, and global web text beyond Common Crawl. We can deliver exabytes of data and 60T+ corpus-ready tokens.
Our ethical scraping infrastructure is built for petabyte- and exabyte-scale procurements. We deliver faster, cheaper, and at higher volume than any competitor.
We continually develop new ways to source hard-to-find data for pre-training, post-training, RL, and evals. Ask for our latest catalog!
hi@exabite.ai02 / EXABITE IS FOR YOU
Our team is entirely technical, from Stanford, MIT, Magic.dev, and Pika. We are laser focused on helping you build the best model.
We have invested substantially in building some of the best technical scraping infrastructure in our industry. Ask for a sample!
03 / PRICE MATCH
For equivalent internet-sourced data on equivalent timelines, we aim to beat any competitor quote by 10–20%.
Request a price match