Komi-Yazva–Russian Parallel Corpus
收藏资源简介:
Komi-Yazva–俄语平行语料库是由高等经济大学等研究机构构建的首个专用于极低资源环境下机器翻译评估的数据集,旨在填补科米-亚兹瓦语这一濒危乌拉尔语族语言在计算资源方面的空白。该数据集包含457个对齐的句子对,源自74篇叙事文本,平均句子长度约为5-6个词元,数据来源于语言学家V. I. Lytkin的田野调查专著《Komi-Yazvinsky Dialect》,经数字化和人工对齐处理而成。其创建过程注重文档级结构标识,支持故事级交叉验证,以控制数据泄漏风险。该数据集主要应用于濒危语言机器翻译研究领域,特别针对零样本和少样本大语言模型翻译性能的评估,旨在解决极端数据稀缺场景下翻译模型的可靠性、提示敏感度及评估方法可复现性等核心问题。
ClaimRAG-LAW is a fine-grained, claim-level legal Retrieval-Augmented Generation (RAG) benchmark dataset developed by the research team at the University of Luxembourg, which aims to support fine-grained evaluation and hallucination detection of legal RAG systems. This dataset contains 317 question-answer pairs and 968 manually verified claims, sourced from the English-language General Data Protection Regulation (GDPR) and French domestic civil law (CIVIL), covering diverse question types and user personas for both expert and non-expert users. The dataset construction process involves extracting question-answer pairs from authoritative legal texts, and diversifying the design based on question categories (e.g., general legal research, factual recall, false premise, etc.) and user personas (citizen, civil official, legal expert) to reflect real-world legal scenarios. It is primarily applied to evaluate the retrieval and generation performance of legal RAG systems, especially in hallucination detection, claim-level accuracy analysis and supporting multilingual access to legal information. The dataset aims to address the shortcomings of existing legal RAG benchmarks in fine-grained evaluation, multilingual coverage and meeting non-expert user needs, so as to improve the reliability and transparency of legal artificial intelligence tools.




