遇见数据集

jpaulpoliquit/ph-pretrain

收藏
Hugging Face2026-05-20 更新2026-05-31 收录
官方服务:

资源简介:

PH Corpus v0.6-ph-unified是一个用于预训练的多语言语料库,主要聚焦菲律宾相关语言。数据集包含多种语言文本,如宿务语(Cebuano)、瓦瑞语(Waray)、菲律宾语(Filipino,包括正式和随意变体)、伊洛卡诺语(Ilocano)、比科尔语(Bicolano)、邦板牙语(Kapampangan)和邦阿西楠语(Pangasinan),同时也包括英语、越南语等其他语言。数据来源于多个公开资源,包括维基媒体项目(Wikimedia)、FineWeb2以及HaloHalo组合数据。总文档数为7,176,685,分为训练集(7,050,119文档)、验证集(63,409文档)和测试集(63,157文档)。数据以扁平Parquet格式存储,包含文本、语言猜测、注册类型、质量评分和元数据等列。该数据集旨在支持自然语言处理任务的预训练模型开发。

PH Corpus v0.6-ph-unified is a multilingual corpus designed for pretraining, with a focus on Philippine-related languages. The dataset includes text in various languages such as Cebuano, Waray, Filipino (both formal and casual variants), Ilocano, Bicolano, Kapampangan, and Pangasinan, along with other languages like English and Vietnamese. Data is sourced from multiple public resources, including Wikimedia projects, FineWeb2, and HaloHalo combined data. It contains a total of 7,176,685 documents, split into training (7,050,119 documents), validation (63,409 documents), and test (63,157 documents) sets. The data is stored in flat Parquet format with columns for text, language guess, register, quality score, and metadata. This corpus is intended to support the development of pretrained models for natural language processing tasks.

提供机构:
jpaulpoliquit
二维码
社区交流群
二维码
科研交流群
商业服务