hcm777/query-embedding-mix-word-mix
收藏资源简介:
该数据集伴随ACL 2026论文发布,专注于多语言密集检索中的混合语言查询。它包含了用于附录验证工作流程的词级代码混合查询捆绑包,用于检查嵌入级插值是否遵循类似的比率趋势。数据集覆盖了四种语言对(EN-ZH, EN-VI, ZH-VI, HI-ID),每种语言对都提供了Hugging Face配置支持的标准化Parquet文件、原始TSV捆绑包以及每对的元数据和校验和。数据集主要用于研究目的,不是训练/测试基准包。
This dataset accompanies our ACL 2026 paper on mixed-language queries in multilingual dense retrieval. It packages the word-level code-mixed query bundles used in the appendix validation workflow, where word-mix is used as a probe to check whether embedding-level interpolation follows similar ratio trends. The release covers four validation pairs (EN-ZH, EN-VI, ZH-VI, HI-ID), each available as a Hugging Face configuration backed by a normalized Parquet file, the original TSV bundle, and per-pair metadata and checksums. The dataset is a paper-facing artifact release, not a train/test benchmark package.




