HEBTEASESUM
收藏资源简介:
HEBTEASESUM是由耶路撒冷希伯来大学研究团队构建的首个希伯来语多文档摘要数据集,基于历史报纸的前页提要自动提取而成。该数据集包含7,774条高质量摘要-文档对,数据源自数字化报纸档案,通过两阶段流程实现:首先识别前页提要中的关键词短语定位摘要,随后匹配对应版面的完整新闻文档。该资源专门针对低资源语言场景设计,有效解决了希伯来语等语言缺乏高质量摘要训练数据的问题,为跨语言摘要模型评估与优化提供了重要基准。
HEBTEASESUM is the first Hebrew multi-document summarization dataset constructed by a research team from the Hebrew University of Jerusalem, automatically extracted from front-page leads of historical newspapers. This dataset contains 7,774 high-quality summary-document pairs sourced from digitized newspaper archives, and is built via a two-stage pipeline: first, identify key phrases in front-page leads to locate the summaries, then match complete news documents from the corresponding newspaper pages. This resource is specifically designed for low-resource language scenarios, effectively addressing the shortage of high-quality summarization training data for languages such as Hebrew, and providing an important benchmark for the evaluation and optimization of cross-lingual summarization models.



