遇见数据集

istqmh23/wikipedia-id

收藏
Hugging Face2026-05-21 更新2026-05-31 收录
官方服务:

资源简介:

这是一个印度尼西亚语文本数据集,包含文档ID、标题和文本内容。数据集分为训练集和测试集,训练集有1,301,683个示例,测试集有144,632个示例。该数据集可用于自然语言处理任务,如文本分类或语言建模。

This is an Indonesian text dataset containing document IDs, titles, and text content. The dataset is split into training and test sets, with 1,301,683 examples in the training set and 144,632 examples in the test set. This dataset can be used for natural language processing tasks such as text classification or language modeling.

提供机构:
istqmh23
二维码
社区交流群
二维码
科研交流群
商业服务