tamil-wikipedia-markdown
收藏资源简介:
该数据集包含泰米尔语维基百科文章转换为Markdown格式的内容,专门用于大型语言模型(LLMs)在泰米尔语内容上的预训练和持续预训练。数据集语言为泰米尔语(ta),格式为单列Parquet文件,包含'text'字段。内容为维基百科文章,标题作为H1标题,后跟文章内容。数据预处理包括将MediaWiki wikitext转换为干净的Markdown格式,移除模板、引用和元数据,保留章节结构、列表和基本格式。数据集适用于语言模型预训练、持续预训练、掩码语言建模、文本生成和迁移学习等任务。数据来源于泰米尔维基百科(ta.wikipedia.org),采用Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)许可。数据集可能存在覆盖偏差、贡献者偏差、时间偏差和领域偏差,建议与其他泰米尔语语料库结合使用以获得更全面的语言覆盖。
This dataset contains Tamil Wikipedia articles converted into Markdown format, specifically designed for pre-training and continued pre-training of Large Language Models (LLMs) on Tamil-language content. The dataset is in Tamil (ta), stored as a single-column Parquet file with a 'text' field. The content consists of Wikipedia articles, where the title serves as the H1 heading, followed by the main article content. Data preprocessing involves converting MediaWiki wikitext into clean Markdown format, removing templates, citations, and metadata while retaining section structures, lists, and basic formatting. This dataset is suitable for tasks such as language model pre-training, continued pre-training, masked language modeling, text generation, and transfer learning. The dataset is sourced from Tamil Wikipedia (ta.wikipedia.org) and is licensed under Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0). This dataset may exhibit coverage bias, contributor bias, temporal bias, and domain bias. It is recommended to combine it with other Tamil-language corpora to achieve more comprehensive language coverage.




