遇见数据集

AI-Impacts-Science-Matters-Arising — data archive (sampling frame, embeddings, edge caches)

收藏
Zenodo2026-05-28 更新2026-06-05 收录
官方服务:

资源简介:

Re-analysis intermediates for a Matters Arising on Xu et al. (2025), Nature. This deposit hosts the generated intermediate and analysis-ready files for an independent re-analysis of: Xu, F. et al. "Artificial intelligence tools expand scientists' impact but contract science's focus." Nature (2025). doi:10.1038/s41586-025-09922-y The raw inputs (paper_info.pkl, venue_info.pkl) are not redistributed here; they are the authors' own published dataset and must be downloaded from doi:10.5281/zenodo.15779090. The complete analysis pipeline (scripts, small result tables, figures) lives on GitHub: github.com/zhichen-hub/AI-Impacts-Science-Matters-Arising. Contents. interim/sampling_frame_2010_2024.pkl — 5.3 GB, post-2010 sampling frame built from the authors' raw inputs. interim/openalex_cache.pkl, interim/ke_openalex_text_cache.pkl — OpenAlex enrichment snapshots dated April 2026. interim/eg_focal_citing_edges.pkl, interim/eg_unique_citing_works.pkl — citing-work edge cache for the ~5,000-focal engagement (EG) sample. interim/ke_embeddings_{specter2,scibert}[_masked]_l2.npy — four KE-side embedding arrays (SPECTER2 / SciBERT × plain / method-masked), all L2-normalised. interim/citing_embeddings_{specter2,scibert}[_masked]_l2.npy — four citing-side embedding arrays, same variants, L2-normalised. processed/ke_sample_pair_complete.pkl — 5,093 AI vs non-AI KE-pair sample. processed/eg_outdeg_per_citing.pkl — per-citing in-family incident degree (column incident_deg_in_family; satisfies 2k = Σ di). processed/engaged_disengaged_pairs.pkl — 1.4 GB engaged / disengaged citing-pair table used by Pillar 3. Total size ~10.5 GB across 16 files. Per-file manifest, MD5 checksums and reproducibility tiers are documented in data/README.md on the companion GitHub repository. Reproducibility tiers. Tier 1 — verify headline numbers from the small CSV / JSON tables already in the GitHub repo (no download needed). Tier 2 — re-run analyses (scripts 10–16): fetch ke_sample_pair_complete.pkl, engaged_disengaged_pairs.pkl, eg_outdeg_per_citing.pkl, and the eight *_l2.npy arrays. ~5–10 GB. Tier 3 — full rebuild from sampling frame: additionally fetch sampling_frame_2010_2024.pkl and the OpenAlex caches; the raw paper_info.pkl / venue_info.pkl come from the authors' Zenodo above. License. Code: MIT (see GitHub). Data files in this deposit are released under CC BY 4.0 for the derived intermediates produced in this re-analysis; the underlying authors' raw dataset retains its original license. Author. Zhichen Guo, The Chinese University of Hong Kong (Shenzhen). Contact: guozhichen@cuhk.edu.cn.

提供机构:
Zenodo
创建时间:
2026-05-28
二维码
社区交流群
二维码
科研交流群
商业服务