遇见数据集

ORCCA: Ontology-Guided Discovery of Multi-Domain Scientific Software — Supplementary Materials and Artifacts

收藏
Zenodo2026-06-18 更新2026-06-18 收录
官方服务:

资源简介:

This deposit contains all artifacts produced by the ORCCA pipeline for the paper "ORCCA: Ontology-Guided Discovery of Multi-Domain Scientific Software at Scale" (ISWC 2026). Files are organized by pipeline stage and type; large binary inputs (the Software Heritage repository dataset and README blob archives) are available separately via a public archive linked from the code repository README. ontologies.zipOWL 2 EL ontologies for MSC 2020 (mathematics), PhySH (physics), and EDAM (bioinformatics), produced by the SKOS-to-OWL conversion pipeline (Stage 1a). Each ontology encodes the taxonomic backbone as SubClassOf axioms with concept labels, descriptions, and synonyms as annotation properties. ont-embeddings.zipOnT encoder embeddings per domain (HDF5 format), metadata pickles, and JSON manifests produced by Stage 1b. Includes fine-tuned embeddings for mathematics and biology, and the pretrained MiniLM-L12-galen checkpoint for physics. box2el-embeddings.zipBox2EL geometric embeddings (NPZ format) mapping ontology classes to axis-aligned boxes in embedding space, encoding SubClassOf as box containment. alignment-scores.zipTopic-to-ontology alignment scores for the full configuration across all three domains (JSONL format). Each record contains a GitHub topic, its top-k aligned ontology concepts, and the fused alignment score from Stage 2 (Equation 1). repositories.zipDiscovery set for the full_readme configuration — the headline results of Table 3. One JSONL file per domain containing repository identifiers (SWHID), GitHub URL, domain ratio score, and README score. audit.zipLLM judge audit data from Section 7: 300-item stratified sample with dual-judge verdicts (Claude Sonnet and Gemini Flash), confidence scores, rationales, Cohen's kappa metrics, and the n=50 expert calibration set. evaluation.zipPipeline 3 evaluation outputs: beta-gamma sensitivity sweep results (Table 5), threshold sweep CSVs, top-100 novel candidate set, ground truth and discoverable subset files, GitHub baseline results, ranking metrics, and component matrices. alignment-scores-ablations.zipTopic-to-ontology alignment scores for all ablation configurations (box2el_struct, specter_flat, full_readme) across all three domains (JSONL format). Required to reproduce the per-configuration rows of Table 3 and the beta-gamma sensitivity sweep of Table 5. repositories-ablations.zipDiscovery sets for all ablation configurations (box2el_struct, ont_only, el_embed, specter_flat, readme_only, codebert, graphcodebert, unixcoder, full) across all three domains (JSONL format). Required to reproduce the full ablation comparison of Table 3. audit-pipeline3.zipEarlier version of the LLM judge audit data produced by Pipeline 3, prior to the final Pipeline 4 run. Included for completeness; the canonical audit results are in audit.zip. registries.zipRegistry snapshots used as ground truth: bio.tools dataset (JSON) and swMATH catalogue (CSV). These define the discoverable subsets used to compute recall in Tables 3-6.

提供机构:
Zenodo
创建时间:
2026-06-17
二维码
社区交流群
二维码
科研交流群
商业服务