遇见数据集

lianghsun/wikipedia-zh-742M

收藏
Hugging Face2024-12-09 更新2024-12-14 收录
官方服务:

资源简介:

本数据集由自行开发的爬虫抓取Wikipedia上标注繁体中文语系(zh-tw)的文本内容,以确保文本语系是繁体中文。目前Hugging Face上标注为繁体中文语系的许多Wikipedia数据集,其实并非真正的Wikipedia数据集,而是来自Wikimedia的内容。本数据集涵盖范围广泛,不仅限于台湾,也包含其他国家的内容,因此可能包含政治不正确或非客观的信息,使用时请谨慎评估。为便于训练,同一个Wikipedia页面的语料已被切分为多个子语料,使用者可依需求进行合并处理。

This dataset is crawled from Wikipedia using a self-developed crawler, focusing on text content labeled as Traditional Chinese (zh-tw) to ensure the text is in Traditional Chinese. The dataset covers a wide range and may contain politically incorrect or non-objective information, requiring careful evaluation when used. The corpus of the dataset has been split into multiple sub-corpora for easier training. The dataset contains 742,565,363 tokens.

提供机构:
lianghsun
二维码
社区交流群
二维码
科研交流群
商业服务