mittagessen/common_corpus_hist_equalized
收藏资源简介:
--- license: cc-by-4.0 task_categories: - text-generation language: - multilingual size_categories: - 100B<n<1T --- # Historical Multilingual Corpus A ~250 billion byte corpus for quick-and-dirty pretraining of multi-lingual small language models, with a focus on historical language use. ## Source The corpus is derived from the first 8 subcorpora of the [pleias/common_corpus](https://huggingface.co/datasets/pleias/common_corpus) dataset. ## Filtering Records are included if they satisfy at least one of the following conditions: - The `date` metadata field has a value below 1900 - The record belongs to the "Open Culture" section (`open_type == "Open Culture"`) The date metadata in Common Corpus is incomplete and quite noisy, so this filter is intentionally broad. Including all "Open Culture" material serves both to boost token counts for the long tail of under-represented languages and to capture additional historical texts whose date field may be missing or unreliable. Texts shorter than 5 bytes are discarded. ## Language-Balanced Sampling To prevent high-resource languages from dominating the corpus, a two-pass sampling procedure is applied: 1. **Survey pass**: All qualifying records are scanned to compute the total byte count per language. 2. **Sampling pass**: A greedy redistribution algorithm assigns per-language sampling rates. Languages whose total bytes fall below an equal share of the target budget are included at a rate of 1.0 (i.e. kept in full); the unused budget is then redistributed equally among the remaining larger languages. This flattens the distribution without discarding any material from smaller languages. During extraction, texts are chunked to a maximum sequence length of 2048 bytes before sampling. The output is written as sharded Parquet files with three columns: `date`, `language` , `text`. ## Token Counts Token counts below are in bytes, corresponding to a byte-level tokenizer. Total corpus size is approximately 250 billion bytes across 233 languages. | Language | Bytes | |---|--:| | Japanese | 20,999,995,461 | | Spanish | 20,997,789,191 | | Latin | 20,997,343,886 | | Italian | 20,992,480,543 | | German | 20,992,083,134 | | English | 20,984,812,739 | | French | 20,983,366,882 | | Arabic | 18,767,527,791 | | Polish | 17,377,379,181 | | Portugueuse | 9,871,842,591 | | Greek | 9,600,804,653 | | Russian | 8,876,882,395 | | Dutch | 7,574,057,626 | | Danish | 5,962,129,190 | | Estonian | 4,583,027,488 | | Modern Korean | 2,811,227,977 | | Czech | 2,734,123,149 | | Hanmun | 1,466,447,521 | | Swedish | 1,419,288,031 | | Hungarian | 1,194,813,398 | | Finnish | 1,061,429,796 | | Yiddish | 885,569,094 | | Hebrew | 821,881,598 | | Maltese | 804,837,459 | | Serbian | 564,214,023 | | Lithuanian | 551,434,902 | | Welsh | 539,896,905 | | Croatian | 469,382,597 | | Unknown | 463,336,523 | | Slovenian | 411,084,510 | | Latvian | 308,837,113 | | Basque | 301,128,057 | | Ukrainian | 290,815,455 | | Slovak | 288,067,504 | | Scottish Gaelic | 274,719,158 | | Korean | 251,724,238 | | Standard Malay (Latin script) | 239,658,095 | | Armenian | 138,972,286 | | Icelandic | 138,392,846 | | Traditional Chinese | 136,869,575 | | Catalan | 119,049,784 | | Romanian | 101,410,270 | | Urdu | 72,274,639 | | Scots | 63,827,026 | | Bulgarian | 61,047,959 | | Norwegian Nynorsk | 49,064,088 | | Xhosa | 46,626,488 | | Indonesian | 44,218,973 | | Sanskrit | 41,712,961 | | Irish | 41,095,001 | | Yoruba | 40,925,704 | | Malagasy | 40,134,138 | | Syriac | 38,651,591 | | Sumerian | 36,918,751 | | Manx | 31,209,369 | | Haitian Creole (Latin script) | 30,754,301 | | Tamil | 28,133,863 | | Akkadian | 27,697,317 | | Limburgish | 27,444,299 | | Galician | 26,627,651 | | Portuguese | 26,428,383 | | Romansh | 25,356,932 | | Vietnamese | 24,195,518 | | Afrikaans | 21,312,154 | | Oromo | 20,975,830 | | Hawaiian | 20,279,514 | | Norwegian | 20,190,659 | | Kinyarwanda | 19,304,296 | | Hmong | 19,069,065 | | Wolof | 18,802,482 | | Bosnian | 17,566,978 | | Persian | 16,550,962 | | Chinese | 15,803,723 | | Sardinian | 15,542,739 | | Breton | 15,270,604 | | Uzbek | 15,112,151 | | Tongan | 13,279,169 | | Standard Latvian | 13,041,011 | | Sundanese | 12,970,237 | | Malay | 12,398,291 | | Tagalog | 11,856,549 | | Bislama | 11,854,939 | | Amharic | 11,711,600 | | Maori | 11,664,094 | | Tsonga | 11,254,685 | | Samoan | 11,153,278 | | Greenlandic | 10,950,668 | | Esperanto | 10,850,923 | | Somali | 10,667,263 | | Thai | 10,499,460 | | Kannada | 10,472,855 | | Tigrinya | 10,330,468 | | Akan | 10,235,472 | | Ge'ez | 10,091,426 | | Javanese | 9,313,178 | | Haitian Creole | 8,933,441 | | Interlingue | 8,932,357 | | Early Modern Korean | 8,564,803 | | Interlingua | 8,510,180 | | Iloko | 8,157,612 | | Luganda | 7,983,100 | | Corsican | 7,868,777 | | Malayalam | 7,496,409 | | West Frisian | 7,431,406 | | Luxembourgish | 7,262,942 | | Swahili | 6,846,489 | | Volapuk | 6,836,334 | | Georgian | 6,822,434 | | Telugu | 6,806,828 | | Mauritian Creole | 6,514,111 | | Occitan | 6,500,306 | | Kirundi | 6,451,187 | | Wolaytta | 6,308,168 | | Sidamo | 6,249,679 | | Northern Uzbek | 6,113,033 | | Afar | 5,978,287 | | Khasi | 5,904,573 | | Turkish | 5,773,144 | | Venetian | 5,665,803 | | Shona | 5,454,088 | | Guarani | 5,445,202 | | Venda | 5,414,129 | | Swahil | 5,359,239 | | Igbo | 5,258,852 | | Norwegian Bokmal | 5,239,792 | | Hausa | 5,120,042 | | Albanian | 4,790,895 | | Lingala | 4,701,942 | | Fijian | 4,566,609 | | Quechua | 4,424,783 | | Fulah | 4,102,657 | | Zhuang | 4,094,907 | | Tatar | 3,995,081 | | Egyptian | 3,994,255 | | Nauru | 3,982,505 | | Southern Sotho | 3,969,464 | | Hindi | 3,894,317 | | Turkmen | 3,497,302 | | Waray | 3,464,105 | | Faroese | 3,450,747 | | Cebuano | 3,251,926 | | Dagbani | 2,965,618 | | Southern Dagaare | 2,719,170 | | Ewe | 2,679,726 | | Seychellois Creole | 2,641,931 | | Zulu | 2,531,741 | | Bengali | 2,343,110 | | Tswana | 2,206,756 | | Ikposo | 2,195,630 | | Tocharian B | 1,920,800 | | Aymara | 1,852,927 | | Swati | 1,798,633 | | Azerbaijani | 1,682,743 | | Chichewa | 1,682,146 | | Sinhala | 1,578,087 | | Sicilian | 1,453,598 | | Sango | 1,410,399 | | qpc | 1,376,146 | | Acoli | 1,334,513 | | Nyankole | 1,159,596 | | Kazakh | 1,154,807 | | Masai | 1,122,233 | | Soga | 1,104,316 | | Macedonian | 928,431 | | Pashto | 907,757 | | Papiamento | 891,944 | | Tocharian A | 751,214 | | Gujarati | 715,915 | | Tibetan | 656,921 | | Earlier Egyptian | 602,294 | | Belarusian | 561,177 | | Lombard | 559,045 | | Tajik | 552,997 | | Mongolian | 515,725 | | Kikuyu | 515,323 | | Demotic Egyptian | 503,150 | | Kyrgyz | 398,363 | | Bau | 365,751 | | Inupiaq | 348,788 | | Fanti | 333,958 | | Pangasinan | 332,354 | | xur | 320,659 | | Late Egyptian | 262,904 | | Tocharian (Skt.; TB) | 171,567 | | Burmese | 166,540 | | Etruscan | 156,174 | | Nigerian Fulfulde | 144,350 | | Old Persian | 135,553 | | Kabyle | 86,973 | | Marathi | 85,470 | | Northern Sotho | 83,380 | | Southern Ndebele | 80,794 | | Minangkabau | 79,589 | | Tok Pisin | 76,740 | | Japanese (Japanese script) | 67,911 | | Bambara | 64,763 | | Bihari | 54,857 | | Kurdish | 47,699 | | Tocharian (Skt.; TA) | 47,500 | | Balinese | 38,152 | | Friulian | 33,458 | | Elamite | 29,020 | | Kimbundu | 24,996 | | Plateau Malagasy | 20,780 | | Iranian Persian | 14,836 | | Sindhi | 14,599 | | Punjabi | 14,033 | | Asturian | 10,881 | | Latgalian | 10,802 | | Crimean Tatar | 9,980 | | Hurrian | 9,323 | | Ugaritic | 8,433 | | Alemannic | 7,583 | | Tocharian | 6,543 | | Aramaic | 5,710 | | Odia | 3,629 | | Middle Korean | 2,764 | | Bashkir | 2,576 | | Magahi | 2,147 | | Hittite | 2,100 | | qeb | 1,770 | | qcu-949 | 1,310 | | qpe | 459 | | xur-946 | 399 | | qcu | 235 | | qur | 229 | | Uyghur | 200 | | Nepali | 55 | | hlu | 52 | ## Format The dataset is stored as sharded Parquet files. Each row contains: | Column | Type | Description | |---|---|---| | `date` | int64 | Year from the source metadata (may be missing or inaccurate) | | `language` | string | Language label from the source metadata | | `text` | string | Text chunk, up to 2048 bytes | ## Intended Use Pretraining of byte-level multi-lingual language models, particularly small models where a curated but imperfect corpus is acceptable. The historical bias makes this corpus especially suitable for models intended to process older texts, archival material, and low-resource languages that are better represented in historical collections. ## Limitations - **Documents chunking/truncation** Documents are chunked into 2048 byte runs. - **Columns** apart from **language, date, and text** are discarded. - **Noisy language classification.** The language labels are inherited from the upstream Common Corpus and are frequently inaccurate. Some languages appear under multiple identifiers -- for example, the two Tocharian languages (A and B) are split across five separate labels ("Tocharian A", "Tocharian B", "Tocharian", "Tocharian (Skt.; TA)", "Tocharian (Skt.; TB)"). Similar fragmentation affects Portuguese/Portugueuse, Swahili/Swahil, Korean/Modern Korean/Early Modern Korean/Middle Korean, Latvian/Standard Latvian, Norwegian/Norwegian Bokmal/Norwegian Nynorsk, Persian/Iranian Persian, Haitian Creole/Haitian Creole (Latin script), Chinese/Traditional Chinese, and others. Several opaque codes (qpc, qeb, qcu, xur, hlu, etc.) also appear. - **Noisy date metadata.** The `date` field is unreliable. Many records have missing, estimated, or clearly incorrect dates. Not all dates use the Gregorian calendar, e.g., Arabic documents have metadata in the Islamic calendar. The filtering threshold of 1900 is a rough heuristic, - **No deduplication.** No cross-document or cross-language deduplication has been applied. - **Only 8 subcorpora.** This corpus was extracted from the first 8 subcorpora of Common Corpus, not the full dataset.



