Data and analysis outputs for: A systematic evaluation and benchmarking of text summarization methods for biomedical literature
收藏资源简介:
Data and analysis outputs supporting the systematic benchmarking of 62 automatic text summarization methods on 1,000 biomedical abstracts from ScienceDirect and Cell Press, using author-written highlights as reference summaries. Includes the gold-standard dataset (1,000 titles, abstracts, and author-written highlights), all generated summaries with per-paper and per-model metric scores across lexical, semantic, and factual dimensions, raw ratings from eight expert evaluators, LLM-as-a-judge panel scores, the data-leakage analyses, and the aggregated ranking and agreement tables reported in the paper. A snapshot of the analysis code is included as a tarball; the code is archived separately at DOI 10.5281/zenodo.21839980 and maintained at github.com/Delta4AI/LLMTextSummarizationBenchmark.



