SINAI/HEP
收藏资源简介:
--- license: cc-by-nc-sa-4.0 pretty_name: HEP configs: - config_name: default data_files: - split: hepth path: Dataset/metadata-hepth.csv - split: hepex path: Dataset/metadata-hepex.csv - split: astroph path: Dataset/metadata-astroph.csv --- --- # HEP - High Energy Physics collection. ## Description: This corpus is oriented to the study of multi-labeled text classifiers. It is composed of scientific articles in the area of High Energy Physics (HEP) obtained from the CDS document server of the European Nuclear Physics Laboratory (CERN). The corpus is divided into three subsets (called partitions), where each partition is composed, in turn, of two files: one containing the records of each article (with information such as abstracts, authors and, of course, classes or keywords) in compressed XML format, and another containing a plain text version of the full article generated from the PDF available in the CERN databases (in tar + gzip format) The classes are delimited by the XML tag KEYWORD. These are the manually assigned DESY thesaurus tags. More information about the DESY thesaurus is available. - hepth split: 18,114 Theoretical Physics documents (metadata - 5.3 Mb) (articles - 226 Mb) - hepex split: 2,599 papers of Experimental Physics (metadata - 1.6 Mb) (articles - 28 Mb) - astroph split: 2,716 Astrophysics documents (metadata - 1.1 Mb) (articles - 29 Mb) ### Licensing Information HEP Collection is released under the [Apache-2.0 License](http://www.apache.org/licenses/LICENSE-2.0). ## Citation: This corpus has been prepared by Arturo Montejo Ráez, with metadata provided by Jens Vigen and the help of the CDS Team. ```bibtex @Article{montejo2004, author = {Montejo-Ráez, A. and Steinberger, R. and Ureña-López, L. A.}, title = {Adaptive selection of base classifiers in one-against-all learning for large multi-labeled collections}, booktitle = {Advances in Natural Language Processing: 4th International Conference, EsTAL 2004}, pages = {1--12}, year = {2004}, editor = {Vicedo J. L. et al.}, location = {Alicante, Spain}, number = {3230}, series = {Lectures notes in artifial intelligence}, publisher = {Springer} } ```
许可证:CC BY-NC-SA 4.0 数据集简称:HEP 配置项: - 配置名称:default 数据文件: - 拆分集:hepth,路径:Dataset/metadata-hepth.csv - 拆分集:hepex,路径:Dataset/metadata-hepex.csv - 拆分集:astroph,路径:Dataset/metadata-astroph.csv --- --- # HEP——高能物理(High Energy Physics, HEP)数据集 ## 数据集说明 本语料库面向多标签文本分类器研究,由欧洲核子研究中心(CERN)CDS文档服务器获取的高能物理领域学术文章组成。该语料库划分为三个子集(又称分区),每个子集又包含两类文件:一类为压缩XML格式的单篇文章记录文件,涵盖摘要、作者、类别与关键词等信息;另一类为从CERN数据库的PDF文件提取的完整文章纯文本版本,采用tar+gzip格式打包。类别信息通过XML标签`KEYWORD`界定,均为人工标注的DESY叙词表标签。更多关于DESY叙词表的相关信息可另行查阅。 - hepth拆分集:收录18114篇理论物理文献(元数据:5.3 Mb;文章文件:226 Mb) - hepex拆分集:收录2599篇实验物理文献(元数据:1.6 Mb;文章文件:28 Mb) - astroph拆分集:收录2716篇天体物理文献(元数据:1.1 Mb;文章文件:29 Mb) ### 授权信息 HEP数据集采用[Apache-2.0许可证](http://www.apache.org/licenses/LICENSE-2.0)发布。 ## 引用信息 本语料库由Arturo Montejo Ráez整理,元数据由Jens Vigen提供,并得到CDS团队的协助。 bibtex @Article{montejo2004, author = {Montejo-Ráez, A. and Steinberger, R. and Ureña-López, L. A.}, title = {Adaptive selection of base classifiers in one-against-all learning for large multi-labeled collections}, booktitle = {Advances in Natural Language Processing: 4th International Conference, EsTAL 2004}, pages = {1--12}, year = {2004}, editor = {Vicedo J. L. et al.}, location = {Alicante, Spain}, number = {3230}, series = {Lectures notes in artifial intelligence}, publisher = {Springer} }
HEP - High Energy Physics Collection
数据集描述
该数据集旨在支持多标签文本分类器的研究,包含从欧洲核物理实验室(CERN)的CDS文档服务器获取的高能物理(HEP)领域的科学文章。数据集分为三个子集(称为分区),每个分区包含两个文件:一个包含每篇文章的记录(如摘要、作者和关键词)的压缩XML格式文件,另一个包含从CERN数据库中的PDF生成的全文纯文本版本(tar + gzip格式)。关键词由XML标签KEYWORD定义,这些标签是手动分配的DESY叙词表标签。
数据集分区
- hepth分区:包含18,114篇理论物理文档,元数据大小为5.3 Mb,文章大小为226 Mb。
- hepex分区:包含2,599篇实验物理论文,元数据大小为1.6 Mb,文章大小为28 Mb。
- astroph分区:包含2,716篇天体物理学文档,元数据大小为1.1 Mb,文章大小为29 Mb。
数据文件配置
- 默认配置:
- hepth分区:路径为
Dataset/metadata-hepth.csv - hepex分区:路径为
Dataset/metadata-hepex.csv - astroph分区:路径为
Dataset/metadata-astroph.csv
- hepth分区:路径为
许可证信息
HEP Collection 根据 Apache-2.0 License 发布。




