jpaulpoliquit/ph-pretrain-02
收藏资源简介:
PH Corpus ph-pretrain-02 是一个用于预训练的菲律宾语料库,包含多种菲律宾语言(如他加禄语、希利盖农语、宿务语、伊洛卡诺语、中比科尔语、查瓦卡诺语、马拉瑙语、瓦瑞语等)和英语文本。数据来源于多个网络资源,包括FineWeb2、BantayWika语料库、维基媒体、HaloHalo组合、GMA新闻、Philstar PSN和Rappler新闻等。语料库总共有187,086个文档,分为训练集(183,379个文档)、验证集(1,877个文档)和测试集(1,830个文档)。数据以扁平Parquet格式存储,避免Hub JSONL嵌套字段转换错误。每个文档包含丰富的元数据,如稳定ID、完整文本、截断预览文本、来源URL、日期、检测语言、词数、标题、来源、语言、令牌数、内容哈希和爬取时间戳等。此外,还包含增强的炼油厂列,如注册信息、语言桶、语言置信度、质量分数、质量桶、许可证、来源类型、发布版本、分割和元数据JSON。数据集适用于自然语言处理预训练任务,特别针对菲律宾语言环境。
PH Corpus ph-pretrain-02 is a pretraining corpus for Philippine languages, including multiple Filipino languages (such as Tagalog, Hiligaynon, Cebuano, Ilocano, Central Bikol, Chavacano, Maranao, Waray, etc.) and English. The data is sourced from various web resources, including FineWeb2, BantayWika Corpus, Wikimedia, HaloHalo Combined, GMA News, Philstar PSN, and Rappler News. The corpus contains a total of 187,086 documents, split into training (183,379 documents), validation (1,877 documents), and test (1,830 documents) sets. Data is stored in flat Parquet format to avoid Hub JSONL nested-field cast errors. Each document includes rich metadata, such as a stable ID, full text, truncated preview text, source URL, date, detected language, word count, title, source, language, token count, content hash, and crawl timestamp. Additionally, it includes enriched refinery columns like register, language bucket, language confidence, quality score, quality bucket, license, source type, release version, split, and metadata JSON. The dataset is suitable for natural language processing pretraining tasks, particularly in the Philippine language context.




