遇见数据集

Methodology data of "Twenty years of research in Digital Humanities: a topic modeling study"

收藏
Zenodo2021-02-19 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

This document contains the datasets created in the thesis "Twenty years of research in Digital Humanities: a topic modeling study". The methodological approach of the work is based on two datasets built by web scraping DH journals’ official web pages and API requests to popular academic databases (Crossref, Datacite). The datasets constitute a corpus of DH research and include research papers abstracts and abstract papers from DH journals and international DH conferences published between 2000 and 2020. Probabilistic topic modeling with latent Dirichlet allocation is then performed on both datasets to identify relevant research subfields. <strong>Data</strong> Folder <strong>"<em>data/</em>" </strong>contains four folders which relate to two datasets: The first dataset, which will be referred to as the journals dataset, contains original research papers published in journals exclusively devoted to digital humanities scholarshipis [1] and is composed of 2,464 articles from 26 journals. The second dataset, the conference dataset, contains abstract papers available in ADHO conference archives and is composed of 2,160 articles from 15 years of ADHO conferences and 4 conferences promoted by journals Both datasets are provided with: URL (if available); identifier and related scheme (if available); abstract or abstract paper; title; authors’ given name, family name; author’s affiliation name, found within the document metadata or text; normalized affiliation name, country of the affiliation, identifiers of the affiliation provided by the Research Organization Registry Community (ROR, https://ror.org); publisher (if available); publishing date (complete date when provided or only the year); keywords (if available); journal title; volume and issue (if available); electronic and/or print ISSN (if available). The two folders <strong>"data/no_abstracts..."</strong> are licensed under a Creative Commons public domain dedication (CC0), while the others keep their original license (the one provided by their publisher) because they contain full abstracts of the papers. These latter datasets are provided in order to favor the reproducibility of the results obtained in our work. <strong>Topic modeling</strong><br> <br> <em><strong>"topic_modeling/"</strong></em> directory contains input and output data used within MITAO, a tool for mashing up automatic text analysis tools, and creating a completely customizable visual workflow [2]. The topic modeling results are divided in two folders, one for each of the datasets. <strong>Note:</strong> It's necessary to unzip the file to get access to all the files and directories listed below. <strong>References</strong> Spinaci, G., Colavizza, G., Peroni, S., <em>Preliminary Results on Mapping Digital Humanities Research</em>, in: Atti del IX Convegno Annuale AIUCD. La svolta inevitabile: sfide e prospettive per l'Informatica Umanistica, Milan, Università Cattolica del Sacro Cuore, 2020, pp. 246 - 252 (atti di: IX Convegno Annuale AIUCD. La svolta inevitabile: sfide e prospettive per l'Informatica Umanistica, Milano, Italy, 15-17 gennaio 2020) Ferri, P., Heibi, I., Pareschi, L., &amp; Peroni, S. (2020). MITAO: A User Friendly and Modular Software for Topic Modelling [JD]. PuntOorg International Journal, 5(2), 135–149. https://doi.org/10.19245/25.05.pij.5.2.3

本文件收录了论文《数字人文研究二十年:主题建模研究》中构建的数据集。本研究的方法论依托两套数据集构建:通过网络爬取数字人文(Digital Humanities)期刊官方网页,以及向主流学术数据库(Crossref、Datacite)发起API请求获取数据而成。该两套数据集构成数字人文研究语料库,收录2000年至2020年间数字人文期刊的研究论文摘要,以及国际数字人文会议的会议论文摘要。随后针对两套数据集开展基于潜在狄利克雷分配(Latent Dirichlet Allocation,LDA)的概率主题建模,以识别相关研究子领域。<strong>数据</strong>文件夹<strong>"<em>data/</em>"</strong>包含四个子文件夹,对应两套数据集:第一套数据集(下文称为期刊数据集)收录仅刊载数字人文研究成果的期刊发表的原创研究论文,共包含来自26种期刊的2464篇文章[1]。第二套数据集为会议数据集,收录ADHO会议档案中可获取的会议论文摘要,包含15届ADHO年会以及4个由期刊主办的会议的2160篇文章。两套数据集均附带以下元数据(如可获取):URL、标识符及关联规范、摘要/会议摘要、标题、作者名与姓、从文档元数据或正文中提取的作者所属机构名称、标准化机构名称、机构所属国家、研究机构注册表社区(Research Organization Registry Community,ROR,https://ror.org)提供的机构标识符、出版者(如可获取)、出版日期(如提供完整日期则使用完整日期,否则仅保留年份)、关键词(如可获取)、期刊名称、卷期(如可获取)、电子版及/或印刷版ISSN(如可获取)。其中两个名为<strong>"data/no_abstracts..."</strong>的文件夹采用知识共享公共领域授权(CC0)许可,其余文件夹保留原出版者提供的原始授权,因其包含论文的完整摘要。提供后一类数据集旨在保障本研究所得结果的可复现性。<strong>主题建模</strong><br><br><em><strong>"topic_modeling/"</strong></em>目录包含MITAO工具所需的输入与输出数据。MITAO是一款可整合自动文本分析工具并构建完全可定制化可视化工作流的软件[2]。主题建模结果分为两个子文件夹,分别对应两套数据集。<strong>注意:</strong>需解压本文件才能访问下文所列的全部文件与目录。<strong>参考文献</strong><br>Spinaci, G., Colavizza, G., Peroni, S.:<em>《数字人文研究映射的初步成果》</em>,载于:《第九届AIUCD年度会议论文集:不可避免的转折:人文信息学的挑战与前景》(Atti del IX Convegno Annuale AIUCD. La svolta inevitabile: sfide e prospettive per l'Informatica Umanistica),米兰,圣心天主教大学,2020年,第246-252页(原会议为:第九届AIUCD年度会议:不可避免的转折:人文信息学的挑战与前景,意大利米兰,2020年1月15-17日)<br>Ferri, P., Heibi, I., Pareschi, L., & Peroni, S.(2020):<em>《MITAO:一款易用且模块化的主题建模软件[JD]》</em>,《PuntOorg国际期刊》,第5卷第2期,第135-149页。https://doi.org/10.19245/25.05.pij.5.2.3

提供机构:
Zenodo
创建时间:
2021-02-19
二维码
社区交流群
二维码
科研交流群
商业服务