遇见数据集

SCLib: Reproducible LLM-Extracted arXiv cond-mat.supr-con Bibliometric Dataset (Freeze 2026-05-28)

收藏
Zenodo2026-05-28 更新2026-05-29 收录
官方服务:

资源简介:

SCLib is a corpus-level bibliometric dataset extracted from the full arXiv cond-mat.supr-con primary-submission record (1991–2026), with LLM-based named-entity recognition for superconducting materials, critical temperatures (Tc), pressure regimes, evidence types, and author geography. This deposit is a standalone data resource: it contains the frozen relational database, all derived analysis outputs, the multi-model NER audit corpus, the production NER prompts (SHA-256 pinned), and all reproducibility code. Publications that use this dataset will appear under "Related identifiers" as they are released. Headline numbers (freeze 2026-05-28) 43,183 active papers (1991–2026); 17 retracted (excluded) 19,028 distinct Tc records (strict filter) 5,880 distinct superconducting materials 7,768 distinct papers with at least one Tc record 99.99% per-paper author-country geographic NER coverage 3,200 LLM extraction calls in the multi-model audit (Claude Opus 4.7, GPT-5.5, GPT-5.4-mini, Gemini 2.5 Flash) with Krippendorff α / Cohen κ / Fleiss κ inter-rater reliability What's included data/ — Frozen PostgreSQL dump (data-only, column-inserts) plus 10 per-table / per-view CSV exports (papers, materials, v_tc_geo_strict, paper_geo, audit_reports, manual_overrides, …) audit/ — Multi-model NER audit SQLite DB, 100-paper cached inputs, full LLM stdout/stderr logs (compressed), and 49 analytical output CSVs (q*.csv + IRR + power-law + valley statistics) schema/ — 37 Alembic database migrations (0001 → 0037_paper_geo) scripts/ — ~40 reproducibility scripts (refresh_corpus_stats, compute_reliability, powerlaw_fit, valley_statistics, timeline_plot, ...) prompts/ — SHA-256 pinned production NER prompts (material_ner_v2_core, author_geo_ner_text, author_geo_ner_pdf, and a PROMPT_MANIFEST.json with provenance) What's intentionally NOT included arXiv full text (chunks table, 1.8 GB) — recoverable via arXiv OAI-PMH using the recipe in REPRODUCE.md User accounts (users, api_keys) and user activity (ask_history, bookmarks) — privacy Manuscripts / paper PDFs — this is a pure data deposit; publications using this dataset will be listed under "Related identifiers" Licensing Data (everything under data/, audit/, schema/) — CC-BY-4.0 Code (everything under scripts/, prompts/) — MIT Reproducibility See REPRODUCE.md for three reproduction levels: Load frozen DB and reproduce analyses (5–30 min) Re-run NER on the same 100-paper audit set (~1 hr, ~$5–20 LLM cost) Rebuild the entire corpus from arXiv (~3 days, ~$200–500 LLM cost) Strict filter convention All bibliometric analyses use the canonical strict filter: tc_kelvin ∈ (0, 300] AND papers.status != 'retracted'. This filter is the source-of-record for the 19,028 record count.

提供机构:
Zenodo
创建时间:
2026-05-28
二维码
社区交流群
二维码
科研交流群
商业服务