遇见数据集

Extraction dataset and analysis script for "Self-Supervised Learning in NLP and Computer Vision: A Quantitative Meta-Analysis of Benchmark Performance, Evaluation Protocol, and Cross-Domain Convergence"

收藏
Zenodo2026-08-16 更新2026-08-20 收录
官方服务:

资源简介:

Extraction dataset, analysis script, and corpus manifest supporting a quantitative meta-analysis of self-supervised learning (SSL) benchmark results across natural language processing and computer vision. Contents. extraction_data.csv: 112 method-benchmark rows drawn from 108 papers (2013-2024), 16 columns, one row per method-benchmark pair as reported by the originating paper. protocol_pairs.csv: the 8 linear-probe / fine-tune pairs behind Table IV, each carrying the source table it was read from. analysis.py: recomputes every corpus-level statistic reported in the paper and verifies that each value appears verbatim in the manuscript. corpus_manifest.md: retrieval record for all 108 source papers, each with a DOI or public URL. Reproducing the results. Install the pinned dependencies (pandas 2.3.3, numpy 2.3.4, scipy 1.16.3) and run "python analysis.py". A clean run exits 0 and reports 40/40 data assertions passed. No statistic in the paper was entered by hand; every one is emitted by this script from these CSVs. The 108 source papers are deliberately NOT included. They are third-party copyrighted works and cannot be redistributed, as PDF or as extracted text. corpus_manifest.md lists every one with a DOI or public URL so any source can be retrieved from its publisher. Licensing. The extraction data is released under CC-BY-4.0; analysis.py is released under the MIT licence.

提供机构:
Zenodo
创建时间:
2026-08-16
二维码
社区交流群
二维码
科研交流群
商业服务