Thai-English-Corpus
收藏资源简介:
Thai-English Corpus 是一个大规模泰英双语语料库,通过整合 Hugging Face 上公开可用的多个数据集构建而成。该语料库汇集了教育内容、网络文档、维基百科文章、法律文本、医学文章、金融文档、软件相关文本以及其他通用领域内容,并统一为适用于语言模型预训练和自然语言处理研究的格式。数据集结构包含两个核心字段:text(文档文本)和source(原始 Hugging Face 数据集 URL),每条记录均保留其原始来源以维护可追溯性。数据来源于十个上游数据集,涵盖通用网络内容、教育材料、维基百科、法律文件、医学文章、财务报告、软件与编程内容、技术文档、开源代码库及相关文本、食品与生活方式内容等多个领域。数据集经过流式加载、主文本字段提取、模式统一、空记录与短文档过滤等处理步骤后,以 Parquet 分片形式导出。该数据集旨在用于基础模型预训练、持续预训练、分词器训练、嵌入模型训练、语言建模研究、检索与表示学习研究以及通用 NLP 实验。需注意,由于数据集由多个独立收集的源聚合而成,可能包含重复文档、近似重复文档、OCR 伪影、格式不一致、样板内容、过时信息、不准确信息、有偏见或不平衡内容以及源特定的质量问题。使用者需自行评估其对于目标用途的适用性,并遵守所有上游数据集的原始许可、归属要求和使用限制。
Thai-English Corpus is a large-scale Thai-English bilingual corpus constructed by integrating multiple publicly available datasets from Hugging Face. It aggregates educational content, web documents, Wikipedia articles, legal texts, medical articles, financial documents, software-related texts, and other general-domain materials, unified into a format suitable for language model pre-training and natural language processing research. The dataset structure includes two core fields: text (document text) and source (original Hugging Face dataset URL), with each record retaining its original source for traceability. Data is sourced from ten upstream datasets, covering areas such as general web content, educational materials, Wikipedia, legal documents, medical articles, financial reports, software and programming content, technical documentation, open-source code repositories and related texts, and food and lifestyle content. After processing steps including streaming loading, main text field extraction, pattern unification, and filtering of empty records and short documents, the dataset is exported as Parquet shards. It is intended for use in foundational model pre-training, continued pre-training, tokenizer training, embedding model training, language modeling research, retrieval and representation learning research, and general NLP experiments. Note that as the dataset is aggregated from multiple independently collected sources, it may contain duplicate documents, near-duplicate documents, OCR artifacts, format inconsistencies, boilerplate content, outdated information, inaccuracies, biased or imbalanced content, and source-specific quality issues. Users must independently assess its suitability for their intended purposes and comply with all original licenses, attribution requirements, and usage restrictions of the upstream datasets.




