遇见数据集

Doc2LoRA idea genes: reproduction code and archived intermediates for "Doc2LoRA Provides Decodable Representations of Scientific Ideas"

收藏
Zenodo2026-09-29 更新2026-10-01 收录
官方服务:

资源简介:

The reproduction workflow and the archived intermediate results for the paper. Everything needed to rebuild the manuscript's tables is here: the code is a snapshot of the trimmed reproduction repository, and the bundles hold the computed intermediates, so a reader does not have to re-derive them. Reproducing from the raw corpora instead needs roughly 200 GB of intermediates and several GPU-days. doc2lora-embedding-code-<sha>.tar.gz — the workflow itself: a Snakemake pipeline trimmed to the dependency closure of what the manuscript reads (24 rule files; snakemake -n paper_assets resolves 330 jobs from the raw corpora). Includes REPRODUCE.md, which maps every figure, table and quoted number to the rule that produces it, and data/ARTIFACTS.tsv, a SHA-256 per archived file. doc2lora-results.tar.zst (1227 files) — everything the manuscript's tables and figures are computed from that needs a GPU, the licensed APS text, or an LLM judge to produce: the paired per-unit score pools behind the bootstrap confidence intervals, the invertible adapter weights, every method's raw cluster labels and the judge verdicts, the pair-axis decodes (Doc2LoRA and ICAE), the PACS node set, and the manuscript's own tables for comparison. With this bundle alone, snakemake paper_assets --rerun-triggers mtime rebuilds every table and figure the paper reads, running only CPU rules. doc2lora-s2and.tar.zst (50 files) — the gene and text embeddings for the five author-name-disambiguation benchmarks (zbMATH, QIAN, ArnetMiner, PubMed, KISTI), so the disambiguation rows can be re-scored from the vectors up. Unpack the code snapshot, then fetch and verify the bundles against data/ARTIFACTS.tsv (SHA-256 per file): tar xzf doc2lora-embedding-code-<sha>.tar.gz && cd doc2lora-embedding-<sha> cp workflow/config.template.yaml workflow/config.yaml # edit the paths python scripts/fetch_artifacts.py results --record <this record> snakemake paper_assets -j4 --rerun-triggers mtime # rebuilds every table, CPU only The 644k-paper APS gene matrices (~137 GB) are not included: they exceed a Zenodo record and are a deterministic function of the APS corpus plus the published hypernetwork checkpoints (snakemake all_embeddings). The APS corpus itself is licensed and is not redistributed here.

提供机构:
Zenodo
创建时间:
2026-09-29
二维码
社区交流群
二维码
科研交流群
商业服务