ACL OCL Corpus
收藏资源简介:
ACL OCL Corpus是由新加坡国立大学计算机学院创建的一个学术语料库,源自ACL Anthology,旨在支持计算语言学领域的开放科学研究。该数据集整合并增强了之前的ACL Anthology版本,提供了元数据、PDF文件、引用图和附加的结构化全文,包含章节、图表和链接到大型知识资源(如Semantic Scholar)。ACL OCL涵盖了七个十年,包含73,000篇论文和210,000个图表。数据集通过监督神经模型检测论文主题,展示了计算语言学中的趋势,如对“句法:标记、分块和解析”的兴趣减退和对“自然语言生成”的兴趣复苏。该数据集适用于多模态研究,如图表标题生成,并通过链接大型科学知识图谱来丰富外部信息,如从Semantic Scholar获取的引用数据和与其他平台(如arXiv)的链接。
The ACL OCL Corpus is an academic corpus developed by the School of Computing, National University of Singapore, sourced from the ACL Anthology, and purpose-built to support open scientific research in computational linguistics. This dataset integrates and enhances prior versions of the ACL Anthology, offering metadata, PDF files, citation graphs, and additional structured full texts comprising sections, figures, and links to large-scale knowledge resources such as Semantic Scholar. Spanning seven decades, the ACL OCL Corpus encompasses 73,000 papers and 210,000 figures. Leveraging supervised neural models to identify paper topics, the dataset uncovers trends in computational linguistics, including the declining interest in "syntax: tagging, chunking, and parsing" and the resurgent interest in Natural Language Generation. This dataset is applicable to multimodal research scenarios such as figure caption generation, and enriches external information by linking to large-scale scientific knowledge graphs, encompassing citation data obtained from Semantic Scholar and connections to other platforms like arXiv.




