遇见数据集

Ericu950/classical-swedish-citations-v2

收藏
Hugging Face2026-05-17 更新2026-05-31 收录
官方服务:

资源简介:

Classical-Swedish Citations (v2) 是一个跨语言引用对齐数据集,包含古典希腊语和拉丁语语料库与瑞典文学语料库Litteraturbanken之间的对齐。数据集通过多语言神经检索发现,并使用大型语言模型判断是否为真实引用。这是版本2,取代了之前的版本,具有改进的检索模型、重排序器、多通道复合评分和LLM验证。数据集包含3,090个确认的引用,来自Litteraturbanken的约10%(仅开放EPUBs)。数据集分为两个子集:句子级对齐(每个记录配对一个希腊语/拉丁语句子与一个瑞典语句子)和窗口级对齐(每个记录配对一个源窗口(5个句子的上下文)与一个目标窗口)。数据集支持句子相似性和文本检索任务,适用于古典接受、互文性、跨语言、引用检测等研究领域。

Classical-Swedish Citations (v2) is a cross-lingual citation alignment dataset between classical Greek and Latin corpora and the Swedish literary corpus Litteraturbanken, found by multilingual neural retrieval and judged for genuine citation by a large language model. This is version 2, superseding the previous version, with improved retrieval models, rerankers, multi-channel composite scoring, and LLM verification. It contains 3,090 confirmed citations from ~10% of Litteraturbanken (open EPUBs only). The dataset includes two subsets: sentence-level alignments (pairing one Greek/Latin sentence with one Swedish sentence) and window-level alignments (pairing one source window of 5 sentences with one target window). It supports tasks like sentence similarity and text retrieval, and is relevant for classical reception, intertextuality, cross-lingual, citation detection, and digital humanities research.

提供机构:
Ericu950
二维码
社区交流群
二维码
科研交流群
商业服务