NIRVLab/EViRAL
收藏资源简介:
EViRAL(Ede-Vietnamese Retrieval Across Languages)是一个跨语言信息检索基准数据集,专注于低资源的南岛语系语言Ede(ISO 639-3: rad,Glottocode: rade1241),该语言主要在越南中部高地使用。该数据集将Ede查询与来自WebFAQ的越南语段落配对,形成了针对该语言的第一个信息检索基准。查询源自WebFAQ的越南语子集,并使用NIRVLab/ViEde机器翻译模型(越南语到Ede)翻译成Ede语;语料库和相关度判断(qrels)则保留自原始WebFAQ越南语检索子集,未作修改。数据集结构遵循标准BEIR/MTEB格式,包含三个子集:queries(包含Ede和越南语双语查询,分为训练、验证和测试集)、corpus(越南语FAQ段落,共约124,000条)和qrels(查询与相关段落的映射,相关度得分为1表示相关)。数据分割采用分层查询不相交策略,按查询字符长度桶(短/中/长)进行分层,确保每个查询仅出现在一个分割中,分割比例为70/15/15,以防止数据泄漏。该数据集旨在用于跨语言密集和稀疏检索模型的基准测试、多语言文本嵌入模型的训练和评估、研究机器翻译查询在跨语言信息检索评估中的有效性,以及支持越南少数民族语言Ede的NLP研究。
EViRAL is a cross-lingual information retrieval benchmark for Ede (ISO 639-3: `rad`, Glottocode: `rade1241`), a low-resource Austronesian language spoken primarily in the Central Highlands of Vietnam. The dataset pairs Ede queries with Vietnamese passages sourced from WebFAQ, forming the first IR benchmark for this language. Queries are derived from the Vietnamese subset of WebFAQ and translated into Ede using the NIRVLab/ViEde machine translation model (Vietnamese-to-Ede). The corpus and relevance judgments (qrels) are retained from the original WebFAQ Vietnamese retrieval subset without modification. The dataset structure follows the standard BEIR/MTEB format, with three subsets: queries (containing bilingual queries in Ede and Vietnamese, split into train, validation, and test), corpus (Vietnamese FAQ passages, approximately 124,000 entries), and qrels (mapping queries to relevant passages with a relevance score of 1 indicating relevant). The split uses a stratified query-disjoint strategy based on query length buckets (short/medium/long), ensuring each query appears in exactly one split with a ratio of 70/15/15 to prevent data leakage. It is intended for benchmarking cross-lingual dense and sparse retrieval models, training and evaluating multilingual text embedding models, studying the effectiveness of machine-translated queries in cross-lingual IR evaluation, and supporting NLP research for the Ede language.



