poopoobabylm/BabyLM-2026-Strict-Small
收藏资源简介:
Detoxified 10M Strict-Small BabyLM训练数据集(BabyLM Turns 4, 2026 BabyLM)是BabyLM 2026 Strict-Small的训练集,总计包含1000万词元。数据集来源多样,包括bnc_spoken、childes、gutenberg、open_subtitles、simple_wiki和switchboard等文件,各文件的词元数量分别为762,073、2,841,101、2,557,721、2,282,877、1,531,437和24,791。所有训练数据均经过预处理,通过仇恨言论检测、情感评分和情绪分析等方法,结合性别和种族词汇表以及明确的侮辱性词汇词典,过滤了有害内容,以确保数据集中不包含可能被模型学习的有害关联。
The Detoxified 10M Strict-Small BabyLM Training Dataset (BabyLM Turns 4, 2026 BabyLM) is the BabyLM 2026 Strict-Small training set, totaling 10M tokens. The dataset is sourced from various files including bnc_spoken, childes, gutenberg, open_subtitles, simple_wiki, and switchboard, with token counts of 762,073, 2,841,101, 2,557,721, 2,282,877, 1,531,437, and 24,791 respectively. All training data underwent precorpus debiasing using a pipeline for detecting and removing toxic content, applying hate speech detection, sentiment scoring, and emotion analysis to flag problematic sentences, which were then filtered using gender and race word lists alongside an explicit slur lexicon to ensure the corpus does not carry harmful associations that could be learned by models trained on this data.




