遇见数据集

Maskrio/nya-ir-miracl-id-preprocessed

收藏
Hugging Face2026-05-25 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是MIRACL数据集(Zhang等人,2023)的一个派生工作,专注于印尼语(id)子集,通过五种不同的预处理策略来研究印尼语中-nya附着词处理如何影响检索质量。预处理策略包括:保持原样(keep)、无条件剥离后缀(naive_strip)、使用PySastrawi的RemoveInflectionalPossessivePronoun访问器(sastrawi_clitic)、用标记<NYA>替换剥离的后缀(sentinel)以及基于规则的前驱词替换(rule_resolved)。数据集包含开发集查询和训练集语料库的JSONL文件,每个文件有id和contents字段。遵循Apache-2.0许可证,与MIRACL一致。

This dataset is a derivative work of MIRACL (Zhang et al., 2023) restricted to the Indonesian (id) subset, preprocessed five different ways to study how Indonesian `-nya` clitic handling affects retrieval quality. Preprocessing strategies include: keep (baseline pass-through), naive_strip (unconditional stripping of suffix), sastrawi_clitic (using PySastrawis RemoveInflectionalPossessivePronoun visitor), sentinel (replacing stripped suffix with <NYA> token), and rule_resolved (rule-based antecedent replacement). The dataset includes JSONL files for dev queries and train corpus, each with id and contents fields. Licensed under Apache-2.0, matching MIRACL.

提供机构:
Maskrio
二维码
社区交流群
二维码
科研交流群
商业服务