遇见数据集

Advancing Open Science Through Research Assessment Reform? Content Analysis of CoARA Action Plans (v2.0.0) [Data set]

收藏
Zenodo2026-04-15 更新2026-05-26 收录
官方服务:

资源简介:

Open Science Term Analysis (CoARA) This repository houses data used in the study 'Advancing Open Science Through Research Assessment Reform? Content Analysis of CoARA Action Plans'. Primary Content: Excel Datasets It consists of four Excel data files in the 'Data_CoARA_Action_Plan_OS_Study' folder: Term_Binary_individual_terms.xlsx – Contains the binary term-level data indicating the presence/absence of individual open science terms. Term_Binary_with_individual_terms+Clusters.xlsx – Contains the binary term-level data for individual open science terms and thematic clusters. Term_Frequencies_Normalized_&_Raw.xlsx – Contains the raw and normalized term frequency data for individual open science terms. Supplementary File_Crosstabulation_Analysis.xlsx – Contains cross-tabulation analysis of binary term-level data for individual open science terms across organization types. Green markings indicate terms with a higher share within a given organization type (i.e., the terms most frequently mentioned in that category), while red indicates that the term is not mentioned (i.e., zero occurrences). Note Two additional columns have been added to these Excel files containing the following metadata: Country and Type of Organisation. Primary source:CoARA Signatories Dashboard Where metadata was missing or unavailable via the dashboard, it was added manually Term_Frequencies_Normalized_&_Raw.xlsx file also has a document wordcount column added. These variables are not generated by the script itself What the Script Does Loads and processes all PDF files from a specified folder Extracts and cleans text using pdftools Removes predefined CoARA boilerplate statements Counts: Exact term frequencies Binary presence (0/1) of each term Saves cleaned text files for each PDF Outputs CSV files with: term frequency dataset binary dataset (individual terms) binary dataset of individual terms and thematic clusters Detects how many PDFs contain CoARA boilerplate language Output Files The script generates the following files in the input folder: *_cleaned.txt – cleaned text for each PDF Term_Frequencies.csv – term frequency dataset Term_Binary.csv – binary dataset (individual terms) Term_Binary_with_Clusters.csv – binary dataset of individual terms and thematic clusters Requirements R packages used: pdftools stringr dplyr purrr tools Install missing packages with: install.packages(c("pdftools", "stringr", "dplyr", "purrr")) Usage Set the path of the CoARA_Original_Action_Plans folder as the value of the folder_path variable in the script. Run the R script in R or RStudio. Review the outputs (cleaned text files, CSVs) in the same folder as the PDFs. Re-use We encourage re-use of the dataset, which is licensed under CC-BY 4.0. Please cite the dataset as: Rushforth, Alexander., Gogadze, Nino., Skhirtladze, Tornike., Pölönen, Janne. (2026) Dataset for "Advancing Open Science Through Research Assessment Reform? Content Analysis of CoARA Action Plans" Contact information: For questions about the dataset extraction procedures: Tornike Skhirtladze — t.skhirtladze@seu.edu.ge For general enquiries about the study design and results: Alexander Rushforth — a.d.rushforth@cwts.leidenuniv.nl License CC-BY 4.0 license

提供机构:
Zenodo
创建时间:
2026-04-14
二维码
社区交流群
二维码
科研交流群
商业服务