Survey data of "Mapping Research output to the SDGs"
收藏资源简介:
<strong>This dataset contains information on what papers and concepts researchers find relevant to map domain specific research output to the 17 Sustainable Development Goals (SDGs).</strong> Sustainable Development Goals are the 17 global challenges set by the United Nations. Within each of the goals specific targets and indicators are mentioned to monitor the progress of reaching those goals by 2030. In an effort to capture how research is contributing to move the needle on those challenges, we earlier have made an initial classification model than enables to quickly identify what research output is related to what SDG. (This Aurora SDG dashboard is the initial outcome as proof of practice.) In order to validate our current classification model (on soundness/precision and completeness/recall), and receive input for improvement, a survey has been conducted to<strong> capture expert knowledge from senior researchers in their research domain related to the SDG</strong>. The survey was open to the world, but mainly distributed to researchers from the Aurora Universities Network. <strong>The survey was open from October 2019 till January 2020, and captured data from 244 respondents in Europe and North America.</strong> 17 surveys were created from a single template, where the content was made specific for each SDG. Content, like a random set of publications, of each survey was ingested by a data provisioning server. That collected research output metadata for each SDG in an earlier stage. It took on average 1 hour for a respondent to complete the survey.<strong> The outcome of the survey data can be used for validating current and optimizing future SDG classification models for mapping research output to the SDGs</strong>. <strong>The survey contains the following questions (see inside dataset for exact wording):</strong> <strong>Are you familiar with this SDG?</strong> Respondents could only proceed if they were familiar with the targets and indicators of this SDG. Goal of this question was to weed out un knowledgeable respondents and to increase the quality of the survey data. <strong>Suggest research papers that are relevant for this SDG (upload list)</strong> This question, to provide a list, was put first to reduce influenced by the other questions. Goal of this question was to measure the completeness/recall of the papers in the result set of our current classification model. (To lower the bar, these lists could be provided by either uploading a file from a reference manager (preferred) in .ris of bibtex format, or by a list of titles. This heterogenous input was processed further on by hand into a uniform format.) <strong>Select research papers that are relevant for this SDG (radio buttons: accept, reject)</strong> A randomly selected set of 100 papers was injected in the survey, out of the full list of thousands of papers in the result set of our current classification model. Goal of this question was to measure the soundness/precision of our current classification model. <strong>Select and Suggest Keywords related to SDG (checkboxes: accept | text field: suggestions)</strong> The survey was injected with the top 100 most frequent keywords that appeared in the metadata of the papers in the result set of the current classification model. respondents could select relevant keywords we found, and add ones in a blank text field. Goal of this question was to get suggestions for keywords we can use to increase the recall of relevant papers in a new classification model. <strong>Suggest SDG related glossaries with relevant keywords (text fields: url)</strong> Open text field to add URL to lists with hundreds of relevant keywords related to this SDG. Goal of this question was to get suggestions for keywords we can use to increase the recall of relevant papers in a new classification model. <strong>Select and Suggest Journals fully related to SDG (checkboxes: accept | text field: suggestions)</strong> The survey was injected with the top 100 most frequent journals that appeared in the metadata of the papers in the result set of the current classification model. Respondents could select relevant journals we found, and add ones in a blank text field. Goal of this question was to get suggestions for complete journals we can use to increase the recall of relevant papers in a new classification model. <strong>Suggest improvements for the current queries (text field: suggestions per target)</strong> We showed respondents the queries we used in our current classification model next to each of the targets within the goal. Open text fields were presented to change, add, re-order, delete something (keywords, boolean operators, etc. ) in the query to improve it in their opinion. Goal of this question was to get suggestions we can use to increase the recall and precision of relevant papers in a new classification model. <strong>In the dataset root you'll find the following folders and files:</strong> <strong>/00-survey-input/</strong> This contains the survey questions for all the individual SDGs. It also contains lists of EIDs categorised to the SDGs we used to make randomized selections from to present to the respondents. <strong>/01-raw-data/</strong> This contains the raw survey output. (Excluding privacy sensitive information for public release.) This data needs to be combined with the data on the provisioning server to make sense. <strong>/02-aggregated-data/</strong> This data is where individual responses are aggregated. Also the survey data is combined with the provisioning server, of all sdg surveys combined, responses are aggregated, and split per question type. <strong>/03-scripts/</strong> This contains scripts to split data, and to add descriptive metadata for text analysis in a later stage. <strong>/04-processed-data/</strong> This is the main final result that can be used for further analysis. Data is split by SDG into subdirectories, in there you'll find files per question type containing the aggregated data of the respondents. <strong>/images/</strong> images of the results used in this README.md. <strong>LICENSE.md</strong> terms and conditions for reusing this data. <strong>README.md</strong> description of the dataset; each subfolders contains a README.md file to futher describe the content of each sub-folder. <strong>In the /04-processed-data/ you'll find in each SDG sub-folder the following files.:</strong> <strong>SDG-survey-questions.pdf</strong> This file contains the survey questions <strong>SDG-survey-questions.doc</strong> This file contains the survey questions <strong>SDG-survey-respondents-per-sdg.csv</strong> Basic information about the survey and responses <strong>SDG-survey-city-heatmap.csv</strong> Origin of the respondents per SDG survey <strong>SDG-survey-suggested-publications.txt</strong> Formatted list of research papers researchers have uploaded or listed they want to see back in the result-set for this SDG. <strong>SDG-survey-suggested-publications-with-eid-match.csv</strong> same as above, only matched with an EID. EIDs are matched my Elsevier's internal fuzzy matching algorithm. Only papers with high confidence are show with a match of an EID, referring to a record in Scopus. <strong>SDG-survey-selected-publications-accepted.csv</strong> Based on our previous result set of papers, researchers were presented random samples, they selected papers they believe represent this SDG. (TRUE=accepted) <strong>SDG-survey-selected-publications-rejected.csv</strong> Based on our previous result set of papers, researchers were presented random samples, they selected papers they believe not to represent this SDG. (FALSE=rejected) <strong>SDG-survey-selected-keywords.csv</strong> Based on our previous result set of papers, we presented researchers the keywords that are in the metadata of those papers, they selected keywords they believe represent this SDG. <strong>SDG-survey-unselected-keywords.csv</strong> As "selected-keywords", this is the list of keywords that respondents have not selected to represent this SDG. <strong>SDG-survey-suggested-keywords.csv</strong> List of keywords researchers suggest to use to find papers related to this SDG <strong>SDG-survey-glossaries.csv</strong> List of glossaries, containing keywords, researchers suggest to use to find papers related to this SDG <strong>SDG-survey-selected-journals.csv</strong> Based on our previous result set of papers, we presented researchers the journals that are in the metadata of those papers, they selected journals they believe represent this SDG. <strong>SDG-survey-unselected-journals.csv</strong> As "selected-journals", this is the list of journals that respondents have not selected to represent this SDG. <strong>SDG-survey-suggested-journals.csv</strong> List of journals researchers suggest to use to find papers related to this SDG <strong>SDG-survey-suggested-query.csv</strong> List of query improvements researchers suggest to use to find papers related to this SDG <strong>Cite as:</strong> <em>Survey data of "Mapping Research output to the SDGs"</em> by Aurora Universities Network (AUR) doi:10.5281/zenodo.3798385 <strong>Attribute as:</strong> <em><strong>Survey data of "Mapping Research output to the SDGs</strong>"</em> by Aurora Universities Network (AUR); Alessandro Arienzo (UNA); Roberto Delle Donne (UNA); Ignasi Salvadó Estivill (URV); José Luis González Ugarte (URV); Didier Vercueil (UGA); Nykohla Strong (UAB); Eike Spielberg (UDE); Felix Schmidt (UDE); Linda Hasse (UDE); Ane Sesma (UEA); Baldvin Zarioh (UIC); Friedrich Gaigg (UIN); René Otten (VUA); Nicolien van der Grijp (VUA); Yasin Gunes (VUA); Peter van den Besselaar (VUA); Joeri Both (VUA); Maurice Vanderfeesten (VUA);<strong> is licensed under a Creative Commons Attribution 4.0 International License.</strong> https://aurora-network.global/project/sdg-analysis-bibliometrics-relevance/
**本数据集收录了研究者认为可将特定领域研究成果映射至17项可持续发展目标(Sustainable Development Goals, SDGs)的相关论文与概念信息。** 可持续发展目标是联合国设定的17项全球性挑战议题,每项目标均配有具体目标与指标,用于监测截至2030年的目标达成进度。为探究研究成果如何推动这些全球性挑战的解决进程,我们此前已开发初始分类模型,可快速识别与各可持续发展目标相关的研究成果(Aurora SDG仪表盘即为该实践的初步产出成果)。为验证当前分类模型的可靠性(稳健性/精确率)与完整性(召回率),并获取改进所需的输入信息,我们面向相关领域的资深研究者开展了专家知识征集调研。本次调研面向全球开放,主要通过Aurora大学联盟(Aurora Universities Network)向研究者分发,调研周期为2019年10月至2020年1月,共收集到来自欧洲与北美的244份有效回复。17份调研问卷基于统一模板生成,每份针对一项SDG定制内容。每份问卷的相关内容(如随机选取的论文集)由数据供应服务器导入,该服务器此前已收集了各SDG对应的研究成果元数据。受访者完成本次调研平均耗时1小时。**本次调研收集的数据可用于验证现有SDG分类模型,并优化未来用于研究成果与SDGs映射的分类模型。** **本次调研包含以下问题(具体措辞详见数据集内文件):** **1. 您是否熟悉本SDG?** 受访者仅在确认熟悉本SDG的目标与指标后方可继续作答,该问题用于筛除不了解相关内容的受访者,提升调研数据质量。 **2. 推荐与本SDG相关的研究论文(上传列表)** 该问题置于问卷首位,以避免受其他问题的影响,用于衡量当前分类模型结果集中论文的完整性(召回率)。为降低作答门槛,受访者可通过引用管理器上传.ris或bibtex格式文件(优先推荐),或直接提供论文标题列表。异构输入将由人工进一步处理为统一格式。 **3. 筛选与本SDG相关的研究论文(单选按钮:接受/拒绝)** 从当前分类模型结果集的数千篇论文中随机抽取100篇纳入调研问卷,用于衡量当前分类模型的稳健性(精确率)。 **4. 筛选并推荐与SDG相关的关键词(复选框:接受 | 文本框:补充建议)** 调研中嵌入了当前分类模型结果集论文元数据中出现频率最高的100个关键词,受访者可选择符合要求的已有关键词,并通过文本框补充新关键词。该问题用于获取可用于提升未来分类模型相关论文召回率的关键词建议。 **5. 推荐包含相关关键词的SDG相关术语表(文本框:URL)** 开放文本框用于提交包含本SDG相关关键词的列表URL,该问题用于获取可用于提升未来分类模型相关论文召回率的术语表建议。 **6. 筛选并推荐与SDG完全相关的期刊(复选框:接受 | 文本框:补充建议)** 调研中嵌入了当前分类模型结果集论文元数据中出现频率最高的100种期刊,受访者可选择符合要求的已有期刊,并通过文本框补充新期刊。该问题用于获取可用于提升未来分类模型相关论文召回率的完整期刊列表建议。 **7. 针对当前查询提出改进建议(按目标分设文本框)** 我们向受访者展示了当前分类模型中针对每项目标所用的查询语句,开放文本框供受访者按照自身观点修改、添加、重排序或删除查询中的内容(如关键词、布尔运算符等)以优化查询效果。该问题用于获取可用于提升未来分类模型相关论文召回率与精确率的查询改进建议。 **在数据集根目录下,您将看到以下文件夹与文件:** **/00-survey-input/** 包含所有单个SDG对应的调研问卷,同时包含按SDG分类的EID列表,用于随机抽取样本供受访者作答。 **/01-raw-data/** 包含调研原始输出数据(为符合公开发布要求,已移除隐私敏感信息)。该数据需与数据供应服务器中的数据结合使用,方可实现完整解读。 **/02-aggregated-data/** 包含经聚合处理的个体回复数据。将调研数据与数据供应服务器的数据结合后,可对所有SDG调研的回复按问题类型进行聚合与拆分。 **/03-scripts/** 包含用于拆分数据,并为后续文本分析添加描述性元数据的脚本。 **/04-processed-data/** 可用于后续分析的核心最终结果。数据按SDG拆分为子目录,每个子目录下包含按问题类型划分的聚合受访者数据文件。 **/images/** 本README.md中所用的结果示意图。 **LICENSE.md** 本数据集的复用条款与条件。 **README.md** 本数据集的说明文档;每个子文件夹均包含单独的README.md文件,用于进一步说明该子文件夹的内容。 **在/04-processed-data/的每个SDG子目录中,您将找到以下文件:** **SDG-survey-questions.pdf** 包含调研问卷的PDF文件 **SDG-survey-questions.doc** 包含调研问卷的Word文档 **SDG-survey-respondents-per-sdg.csv** 调研与回复的基础信息 **SDG-survey-city-heatmap.csv** 各SDG调研受访者的来源地信息 **SDG-survey-suggested-publications.txt** 经格式化的论文列表,包含受访者上传或列出的希望纳入本SDG结果集的研究论文 **SDG-survey-suggested-publications-with-eid-match.csv** 与上述文件内容一致,但匹配了EID。EID通过爱思唯尔(Elsevier)内部模糊匹配算法完成匹配,仅高置信度匹配的论文会显示其Scopus数据库中的对应记录EID。 **SDG-survey-selected-publications-accepted.csv** 基于此前的论文结果集,向受访者展示随机抽取的样本,受访者可选择认为与本SDG相关的论文(TRUE=已接受) **SDG-survey-selected-publications-rejected.csv** 基于此前的论文结果集,向受访者展示随机抽取的样本,受访者可选择认为与本SDG无关的论文(FALSE=已拒绝) **SDG-survey-selected-keywords.csv** 基于此前的论文结果集,向受访者展示其元数据中包含的关键词,受访者可选择认为与本SDG相关的关键词 **SDG-survey-unselected-keywords.csv** 与"selected-keywords"对应,该列表包含受访者未选择的与本SDG相关的关键词 **SDG-survey-suggested-keywords.csv** 受访者推荐的可用于查找本SDG相关论文的关键词列表 **SDG-survey-glossaries.csv** 受访者推荐的可用于查找本SDG相关论文的术语表(包含关键词)列表 **SDG-survey-selected-journals.csv** 基于此前的论文结果集,向受访者展示其元数据中包含的期刊,受访者可选择认为与本SDG相关的期刊 **SDG-survey-unselected-journals.csv** 与"selected-journals"对应,该列表包含受访者未选择的与本SDG相关的期刊 **SDG-survey-suggested-journals.csv** 受访者推荐的可用于查找本SDG相关论文的期刊列表 **SDG-survey-suggested-query.csv** 受访者推荐的可用于查找本SDG相关论文的查询改进方案列表 **引用方式:** > *《研究成果与SDGs映射调研数据》* 由Aurora大学联盟(Aurora Universities Network, AUR)发布,DOI: 10.5281/zenodo.3798385 **署名要求:** > *《研究成果与SDGs映射调研数据》* 由Aurora大学联盟(Aurora Universities Network, AUR);Alessandro Arienzo(UNA);Roberto Delle Donne(UNA);Ignasi Salvadó Estivill(URV);José Luis González Ugarte(URV);Didier Vercueil(UGA);Nykohla Strong(UAB);Eike Spielberg(UDE);Felix Schmidt(UDE);Linda Hasse(UDE);Ane Sesma(UEA);Baldvin Zarioh(UIC);Friedrich Gaigg(UIN);René Otten(VUA);Nicolien van der Grijp(VUA);Yasin Gunes(VUA);Peter van den Besselaar(VUA);Joeri Both(VUA);Maurice Vanderfeesten(VUA)创作,本数据集采用知识共享署名4.0国际许可协议(Creative Commons Attribution 4.0 International License)进行授权。 > 相关链接:https://aurora-network.global/project/sdg-analysis-bibliometrics-relevance/



