kp7742/YALM-pretrain1-20M
收藏资源简介:
YALM预训练数据集-1是一个包含约2000万样本的混合数据集,包括英语、印地语和混合英语印地语(Hinglish)。这些数据来自不同的来源,用于语言建模任务和YALM(Yet Another Language Model)的开发。数据集包含以下子数据集:HuggingFaceFW/fineweb-edu(英语,9.67M样本),prince-canuma/fineweb-CC-MAIN-2024-10-1B-en(英语,1.5M样本),anirudhlakhotia/baarat-batched-hindi-pre-training(印地语,8.78M样本),Abhishekcr448/Hinglish-Everyday-Conversations-1M(混合英语印地语,1M样本)。
The YALM Pretraining Data - 1 is a mix of approximately 20 million samples consisting of English, Hindi, and Hinglish. These data are sourced from various origins and are intended for language modeling tasks and the development of the YALM (Yet Another Language Model). The dataset includes the following subsets: HuggingFaceFW/fineweb-edu (English, 9.67M samples), prince-canuma/fineweb-CC-MAIN-2024-10-1B-en (English, 1.5M samples), anirudhlakhotia/baarat-batched-hindi-pre-training (Hindi, 8.78M samples), Abhishekcr448/Hinglish-Everyday-Conversations-1M (Hinglish, 1M samples).



