遇见数据集

BabyLM-community/BabyLM-2026-Strict-Small

收藏
Hugging Face2026-04-08 更新2026-04-12 收录
官方服务:

资源简介:

--- license: mit --- # Detoxified 10M Strict-Small BabyLM Training Dataset (BabyLM Turns 4, 2026 BabyLM) BabyLM 2026 Strict-Small training set. Total: **10M tokens**. Please cite the following: ``` @misc{choshen2026babylmturns4papers, title={BabyLM Turns 4: Call for Papers for the 2026 BabyLM Workshop}, author={Leshem Choshen and Ryan Cotterell and Mustafa Omer Gul and Jaap Jumelet and Tal Linzen and Aaron Mueller and Suchir Salhan and Raj Sanjay Shah and Alex Warstadt and Ethan Gotlieb Wilcox}, year={2026}, eprint={2602.20092}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2602.20092}, } ``` ## Token Counts Breakdown | File | Tokens | |------|-------:| | bnc_spoken.train.txt | 762,073 | | childes.train.txt | 2,841,101 | | gutenberg.train.txt | 2,557,721 | | open_subtitles.train.txt | 2,282,877 | | simple_wiki.train.txt | 1,531,437 | | switchboard.train.txt | 24,791 | | **Total** | **10,000,000** | ## Data Decontamination All training data was subjected to precorpus debiasing using pipeline for detecting and removing toxic content from naturalistic language corpora. The pipeline applies hate speech detection, sentiment scoring, and emotion analysis to flag problematic sentences, which are then filtered using gender and race word lists alongside an explicit slur lexicon. This ensures that demographic mentions and identity-related language in the corpus do not carry harmful associations that could be learned by models trained on this data. Training Data Decontamination was motivated by efforts on the Interaction Track (e.g., Salhan et al 2025; Trhlik, Caines & Buttery 2026). ``` @misc{salhan2025teacherdemonstrationsbabylmszone, title={Teacher Demonstrations in a BabyLM's Zone of Proximal Development for Contingent Multi-Turn Interaction}, author={Suchir Salhan and Hongyi Gu and Donya Rooein and Diana Galvan-Sosa and Gabrielle Gaudeau and Andrew Caines and Zheng Yuan and Paula Buttery}, year={2025}, eprint={2510.20411}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2510.20411}, } ``` ``` @misc{trhlik2026biasdynamicsbabylmscomputeefficient, title={Bias Dynamics in BabyLMs: Towards a Compute-Efficient Sandbox for Democratising Pre-Training Debiasing}, author={Filip Trhlik and Andrew Caines and Paula Buttery}, year={2026}, eprint={2601.09421}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2601.09421}, }

许可证:MIT --- # 经过去毒处理的1000万严格小样本BabyLM训练数据集(BabyLM四岁啦,2026年BabyLM研讨会) BabyLM 2026严格小样本训练集。总规模:**1000万Token(Token)**。 请引用以下文献: @misc{choshen2026babylmturns4papers, title={BabyLM四岁啦:2026年BabyLM研讨会征稿启事}, author={Leshem Choshen and Ryan Cotterell and Mustafa Omer Gul and Jaap Jumelet and Tal Linzen and Aaron Mueller and Suchir Salhan and Raj Sanjay Shah and Alex Warstadt and Ethan Gotlieb Wilcox}, year={2026}, eprint={2602.20092}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2602.20092}, } ## Token统计明细 | 文件名 | Token 数量 | |------|-------:| | bnc_spoken.train.txt | 762,073 | | childes.train.txt | 2,841,101 | | gutenberg.train.txt | 2,557,721 | | open_subtitles.train.txt | 2,282,877 | | simple_wiki.train.txt | 1,531,437 | | switchboard.train.txt | 24,791 | | **总计** | **10,000,000** | ## 数据去污染处理 所有训练数据均通过针对自然语言语料库中有害内容的检测与移除流水线,完成了语料前去偏处理。该流水线通过仇恨言论检测、情感评分与情感分析标记存在问题的语句,并结合性别与种族词汇表以及显性辱骂性歧视词汇表完成过滤。此举可确保语料库中的人口群体提及与身份相关语言不会携带有害关联,进而避免基于该数据集训练的模型学习到此类关联。数据去污染处理的设计灵感源自交互赛道(Interaction Track)的相关研究(例如Salhan等人2025年;Trhlik、Caines与Buttery 2026年)。 @misc{salhan2025teacherdemonstrationsbabylmszone, title={BabyLM最近发展区内的教师演示:面向依随式多轮交互的沙箱}, author={Suchir Salhan and Hongyi Gu and Donya Rooein and Diana Galvan-Sosa and Gabrielle Gaudeau and Andrew Caines and Zheng Yuan and Paula Buttery}, year={2025}, eprint={2510.20411}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2510.20411}, } @misc{trhlik2026biasdynamicsbabylmscomputeefficient, title={BabyLM中的偏差动态:构建算力高效沙箱以推动预训练去偏民主化}, author={Filip Trhlik and Andrew Caines and Paula Buttery}, year={2026}, eprint={2601.09421}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2601.09421}, }

提供机构:
BabyLM-community
二维码
社区交流群
二维码
科研交流群
商业服务