遇见数据集

Resources for the Evaluation of Automatic Indexing: Corpora and Gold Standard Indexes in the Agricultural, Educational, and Medical Domains

收藏
Zenodo2025-11-08 更新2026-05-26 收录
官方服务:

资源简介:

Three corpora have been constructed, each with its respective gold standard index, covering the domains of Agriculture, Education, and Medicine, for use in the evaluation of automatic indexing systems.The Agriculture corpus consists of four subcorpora. One contains 8,000 documents in plain text format, intended for training tasks, while the other three subcorpora contain 500 documents each, suitable for evaluation purposes. Each document in the Agriculture corpus includes a title, abstract, author keywords, and bibliographic references, metadata corresponding to scientific articles published in English in nine journals. The gold standard index for this domain was derived from the AGRICOLA database, maintained by the United States National Agricultural Library.The Education corpus is composed of four sets of documents. The main corpus contains 6,000 documents, and the other three sets include 450 documents each, all in plain text format. Each document in the Education corpus includes a title, abstract, author keywords, and bibliographic references, representing metadata from scientific articles published in English across more than 60 journals. The gold standard index for this domain comes from the ERIC database, provided by the Institute of Education Sciences (IES) of the U.S. Department of Education.The Medicine corpus is made up of five sets of documents. The main corpus includes 1,000 full-text scientific articles in XML-JATS format, published in Spanish across eight journals. The gold standard index for the medical domain originates from the LILACS database, maintained by BIREME in Brazil.

提供机构:
Zenodo
创建时间:
2025-07-26
二维码
社区交流群
二维码
科研交流群
商业服务