遇见数据集

almanach/halvest

收藏
Hugging Face2026-05-21 更新2026-06-14 收录
官方服务:

资源简介:

HALvest是一个未经过滤的版本,包含从Hyper Articles en Ligne (HAL) 开放档案中收集的科学论文全文,并附带用于潜在过滤的额外字段。该数据集主要包含英语和法语论文,但涵盖了56种语言和13个领域的论文。构建过程包括三个步骤:从HAL API获取数据、使用GROBID将PDF转换为结构化格式(xml-tei和txt),以及计算每个文档的统计信息。数据集支持多种语言(如英语、法语、西班牙语等)和领域(如人文社会科学、计算机科学、生命科学等),适用于文本生成和掩码语言建模任务。需要注意的是,由于PDF编码问题,原始版本中的令牌数量可能被夸大,导致部分文档/文本为乱码。数据集遵循HAL的许可条款,用户在使用前需考虑版权问题。

This is the unfiltered version of HALvest, comprising of fulltext from open papers found on Hyper Articles en Ligne (HAL) with extra fields for potential filtering. Our dump is mostly english/french but gather papers written in 56 languages across 13 domains. Building the dataset is a three steps process: data fetching from HAL, data merging and data enriching. We first request HALs API to gather open research papers and parse it, then download the PDFs. Using GROBID, we convert each PDF to an xml-tei format for structured data, and convert to txt before concatenation. Finally, we compute some statistics about each document. Note that the number of tokens is highly inflated in the raw version due to badly encoded PDFs, translating to gibberish documents/texts. The dataset supports tasks like text-generation and fill-mask, and follows HALs license terms.

提供机构:
almanach
二维码
社区交流群
二维码
科研交流群
商业服务