遇见数据集

False Authorship: Methods and materials package

收藏
Zenodo2025-03-18 更新2026-05-26 收录
官方服务:

资源简介:

This package contains Python, shell, awk scripts, and data used to obtain the curated table and excerpt associated with the above named article. Data Contents The following data files are included. * README.md: This file * article-details.xlsx: Curated table with details of published articles in Microsoft Excel file format * index.html: HTML document with * links to GIJIR materials saved in the Internet Archive * a list of all the GIJIR articles’ citation data according to Crossref and links to each article’s locally available landing page, full-text PDF, plus links to Crossref metadata and the article via DOI and original journal URL. (Note that non-local, non-archived links may rot over time.) * ybs-works.json: Results of Crossref query to obtain all the publisher’s works made on 2024-09-22 * ChatGPT: Prompts and responses associated with the generation of a fake article in one of the journal’s topics. * global-us/metadata/: Article metadata as HTML files collected on 2024-09-10 * global-us/global-us.mellbaou.com/index.php/global/article/download/: A copy of the journal’s article PDFs as crawled on 2024-09-10 * spinellis business - Google Scholar.pdf: Printout of a Google Search query for the terms spinellis business made on 2025-02-06. Executable Contents The following programs and scripts are used to obtain the above contents. Makefile: Commands that orchestrate the articles’ analysis get-metadata.sh: Obtain article metadata pages from the journal’s web site apply-to-pdfs.sh: Apply the specified Python script to all article PDFs extract-citations-emails.py: Extract number of probable in-text citations and corresponding author email from article PDF extract-doi-affiliations.py: Extract article DOI and affiliations from an article’s metadata extract-all-doi-affiliations.sh: Extract article DOI and affiliations from all articles’ metadata emails-to-csv.awk: Convert emails and article numbers to CSV with URL for sending emails

本套件包含用于获取上述指定文章相关整理表格与节选内容的Python脚本、Shell脚本、Awk脚本及配套数据。 ## 数据内容 本套件包含以下数据文件: * README.md:本说明文件 * article-details.xlsx:采用Microsoft Excel格式存储的已发表文章详情整理表格 * index.html:包含以下内容的HTML文档: 1. 保存于互联网档案馆(Internet Archive)的GIJIR相关材料链接 2. 所有GIJIR文章的引用数据清单(基于交叉参考平台(Crossref)),并附带每篇文章的本地可用着陆页、全文PDF链接,以及通过数字对象标识符(DOI)和原期刊网址指向文章与交叉参考平台元数据的链接。(注:非本地、未归档的链接可能随时间失效。) * ybs-works.json:2024年9月22日通过交叉参考平台查询获取的该出版社所有作品的查询结果 * ChatGPT:与某期刊主题下虚构文章生成相关的提示词与回复内容 * global-us/metadata/:2024年9月10日采集的文章元数据HTML文件 * global-us/global-us.mellbaou.com/index.php/global/article/download/:2024年9月10日爬取的该期刊文章PDF副本 * spinellis business - Google Scholar.pdf:2025年2月6日针对关键词"spinellis business"的Google Scholar搜索结果打印件 ## 可执行内容 以下程序与脚本用于获取上述数据内容: * Makefile:用于统筹文章分析流程的命令集合 * get-metadata.sh:从期刊网站采集文章元数据页面的Shell脚本 * apply-to-pdfs.sh:将指定Python脚本应用于所有文章PDF的Shell脚本 * extract-citations-emails.py:从文章PDF中提取疑似文内引用数量及对应作者邮箱的Python脚本 * extract-doi-affiliations.py:从文章元数据中提取数字对象标识符(DOI)与作者机构信息的Python脚本 * extract-all-doi-affiliations.sh:从所有文章的元数据中提取数字对象标识符与作者机构信息的Shell脚本 * emails-to-csv.awk:将邮箱与文章编号转换为带邮件发送链接的CSV文件的Awk脚本

提供机构:
Zenodo
创建时间:
2024-09-11
二维码
社区交流群
二维码
科研交流群
商业服务