遇见数据集

Data and code for: "Low-Resource Topic Modeling for Domain Experts: A Pilot Study in CO₂ Mineralization Literature"

收藏
Zenodo2026-08-07 更新2026-08-13 收录
官方服务:

资源简介:

This repository contains the data and analysis code accompanying the manuscript "Low-Resource Topic Modeling for Domain Experts: A Pilot Study in CO₂ Mineralization Literature" (Sabatino et al.). The study benchmarks four topic-modeling pipelines (MalletLDA, ProdLDA, BERTopic+SciBERT, SciBERT-NVCTM) implemented on consumer hardware (Mac Mini M2, 16 GB RAM) by domain experts without formal NLP training, on a corpus of 286 CO₂-mineralization papers (2013–2025). It also includes a four-level zero-shot ablation of a Large Language Model (Claude Haiku 4.5) for innovation detection, a calibrated outlier-detection ensemble, and a semantic bridge-term network analysis. Contents data/metrics/ — per-seed model metrics, aggregated metrics, ANOVA + Bonferroni results, with regeneration script. data/stability/ — cross-seed stability (ARI, Jaccard, doc- and topic-word-level). data/outlier_ensemble/ — ensemble outlier scores for all documents and the 9 declared outliers. data/llm_ablation/ — per-paper R1–R4 decisions, ablation table, run summaries, pick-level overlaps. data/topic_assignments/ — top-1 topic assignments per model. data/figure2_macroclusters/ — topic→cluster mapping, per-period weights, Jaccard matrices, figure-regeneration scripts. code/ — full pipeline: LLM baseline (with verbatim prompts and raw outputs), outlier detection, topic metrics, topic modeling, and semantic network analysis (see the NOTA_provenienza.md / README.md files in each subfolder for provenance). All Claude Haiku 4.5 calls used the model snapshot claude-haiku-4-5-20251001. The analysis was produced on consumer hardware without fine-tuning, retrieval augmentation, or specialized infrastructure. Related / alternate identifiers To add after acceptance: the journal article DOI, relation "is supplemented by this upload" (or "is supplement to"). Cite the concept DOI (the "all versions" DOI) in the manuscript's Data Availability Statement, so it always resolves to the latest version. Version 2.0.0 (August 2026). This version accompanies the revised manuscript submittedto Analytics after peer review. It adds the materials produced during the revision andcorrects several files deposited in version 1.0.0. Added. Expert evaluation of the 82 LLM-generated topic labels (data/interrater_labels/:the three completed rating sheets, their key files, compute_interrater.py,make_rating_sheets.py and the frozen output). External expert review of the topic-to-macrocluster mapping (data/figure2_macroclusters/expert_validation/: the 82-topic sheet, themajority verdict and the sheet generator). Sensitivity analyses added in reply to thereviewers: cluster_validation.py, temporal_sensitivity.py, outlier_sensitivity.py andeffect_sizes.py, each with its results file. Dependency files for every component(code/environment/). The 24 execution logs of the six-seed benchmark(data/stability/run_logs/), from which the elapsed run times reported in Section 4.3 of therevised manuscript are computed. The topic labels of the canonical relabelling run(data/topic_labels/topic_labels_with_raw_relabel_20260730_150742.csv). The regeneratedFigures 2 and 3 (data/figure2_macroclusters/Figura2_macroclusters.*,Figura3_topic_similarity.*). Corrected. Supplementary_Information.docx is updated to the revised version, which addsTables S10 to S13 and corrects the definition of the metric previously called "DomainExpert", now reported as Technical Term Salience and described as the computed vocabularystatistic it is. topic_cluster_mapping.csv carries the mapping as corrected by the expertreview. data/topic_assignments/BERTopic-SciBERT_topics_table_all_docs.csv replaces a filethat was not an assignment table. extract_topics_for_relabel.py and08_relabel_topics_haiku.py now pin the canonical run explicitly instead of selecting thefirst candidate. make_fig23.py and the Figure 2 outputs are regenerated on the correctedmapping. plot_fig23.py now uses paths relative to its own location, so it runs from thedownloaded archive, and it redraws Figure 3 with hierarchical ordering of the topics.NOTA_provenienza.md documents these corrections and the provenance issues foundwhile making them. Corrections to the outlier analysis and to the LLM ablation. Two defects found during afinal check of every reported figure against the deposited files have been corrected hereand reported in the response to the reviewers. First, co2_outlier_results.csv contained arecord that is not a document: the detection pipeline globs a directory of JSON files andhad picked up doc_id_to_filename.json, the index-to-filename map of the corpus. Theoutlier corpus is therefore 288 documents rather than 289, and the detection rate9/288 = 3.125% rather than 3.11%. The record has been removed, the glob inco2_outlier_analysis.py and co2_outlier_analysis_filtered.py now excludes the map, andanalysis_summary.json and outlier_sensitivity_results.json are regenerated. The ninedeclared outliers and every cell of the threshold-sensitivity grid are unchanged. Second,06_compare_calibration_runs.py identified the declared outliers by matching authorsurnames against a list transcribed from the manuscript, which admitted four papers thatare not ensemble outliers and omitted three that are. It now resolves them through thecorpus manifest, whose source_file field matches the doc_id of the outlier results forall nine, and pick_level_overlap.json, per_paper_outlier_decisions.csv,consensus_outliers.json, the ablation table and the deposited discussion note areregenerated from it.

提供机构:
Zenodo
创建时间:
2026-08-07
二维码
社区交流群
二维码
科研交流群
商业服务