遇见数据集

Reproducibility package for: Benchmark-Task Heterogeneity in Misinformation-Related and Human–AI Text Classification

收藏
Zenodo2026-07-29 更新2026-08-01 收录
官方服务:

资源简介:

Code, derived feature matrices, frozen sentence-BERT embeddings, analysis results and sensitivity analyses accompanying the article "Benchmark-Task Heterogeneity in Misinformation-Related and Human–AI Text Classification" (submitted to Information, MDPI, 2026). Contents: 22-feature stylometric matrices and all-MiniLM-L6-v2 embeddings for fifteen public English-language benchmarks (n = 24,906 documents; derived data only, no raw text); five self-checking reproduction scripts runnable offline and CPU-only from the package root — reproduce_tables.py (Tables 2–3, verified against the published results with bit-identical determinism across two runs), per_feature_fdr.py (330-test per-feature Benjamini–Hochberg stability analysis), truncation_exposure.py (truncation-exposure audit), method_comparison_stats.py (Friedman, benchmark-blocked ANOVA, aggregation scenarios) and paper4_make_tables.py; 15×15 zero-shot transfer matrices; leakage, duplicate and VIF diagnostics; the full sensitivity appendix (docs/Appendix-A4-sensitivity.md, Sections A4.1–A4.7). Pinned dependency versions (Python 3.11; numpy 2.4.4, scipy 1.16.0, scikit-learn 1.6.0). All splits, bootstrap resamples and model training use the fixed per-benchmark random seed 42. Raw benchmark text is not redistributed; the fifteen source corpora remain available under their original licenses (provenance table in README.md). SHA-256 of the archive: 58e5a76da1839f6158bfefc0764d61de2cceea7777f74d22cc90ee1787b36864.

提供机构:
Zenodo
创建时间:
2026-07-29
二维码
社区交流群
二维码
科研交流群
商业服务