The AMPL-aligned federation, by source category. Token counts are drawn from public corpus documentation — Common Corpus (PleIAs), Common Pile (EleutherAI), Institutional Books (Harvard IDI). Click any segment to see details.
Public domain books, newspapers from cultural heritage repositories, and open projects like Wikisource and Project Gutenberg. The largest source category in the federation — primarily multilingual periodical archives and pre-1929 book corpora cleared through copyright expiration.