遇见数据集

adzcai/babylm-eng-zho-50-50-stratified

收藏
Hugging Face2026-04-27 更新2026-05-03 收录
官方服务:

资源简介:

BabyLM English–Chinese 50/50 Stratified是一个为多语言BabyLM项目设计的双语训练语料库,结合了50%的英语和50%的汉语BabyBabelLM语料库。采样是按类别分层的,以保持每种语言内原始类别比例。语料库包含英语和汉语的令牌计数分别为49,481,353(41.8%)和68,925,921(58.2%),总计118,407,274个令牌。类别细分包括儿童可用语音、儿童书籍、儿童导向语音、儿童维基、教育、填充开放字幕、填充维基百科和字幕等,展示了每种语言在不同类别中的分布情况。

BabyLM English–Chinese 50/50 Stratified is a bilingual training corpus for the Multilingual BabyLM project, combining 50% of the English and Chinese BabyBabelLM corpora. Sampling is stratified by category to preserve the original category proportions within each language. The corpus contains token counts of 49,481,353 (41.8%) for English and 68,925,921 (58.2%) for Chinese, totaling 118,407,274 tokens. The category breakdown includes child-available-speech, child-books, child-directed-speech, child-wiki, educational, padding-opensubtitles, padding-wikipedia, and subtitles, showing the distribution of each language across different categories.

提供机构:
adzcai
二维码
社区交流群
二维码
科研交流群
商业服务