SINAI/ALIA-es-legal-administrative-triplets
收藏资源简介:
ALIA西班牙法律和行政三元组语料库是一个用于密集检索训练的数据集,专门针对西班牙法律和行政语言领域。它包含从原始查询-段落对中自动生成的困难负样本,这些负样本在语义上与查询相似但不包含正确答案,有助于提升检索模型的鲁棒性和排名性能。数据集分为两个配置:训练配置提供多负样本训练数据,包括查询、一个正段落和多个困难负段落,并带有训练阶段和难度标签;评估配置提供用于检索模型评估的三元组数据,包括查询、段落候选和参考答案。数据集支持多种难度级别(高中、大学、博士),以促进课程学习和模型泛化。生成过程基于SentenceTransformers和FAISS相似性搜索,使用Qwen3-Embedding-0.6B嵌入模型自动挖掘困难负样本。该数据集旨在改进西班牙法律编码器的开发,并促进公民对法律信息的访问。
ALIA Spanish Law and Administrative Triples Corpus is a dataset for dense retrieval training, specifically tailored for the Spanish legal and administrative language domain. It contains hard negative samples automatically generated from original query-passage pairs, which are semantically similar to the query but do not contain the correct answer, helping to improve the robustness and ranking performance of retrieval models. The dataset is divided into two configurations: the training configuration provides multi-negative sample training data, including queries, one positive passage, multiple hard negative passages, along with training stages and difficulty labels; the evaluation configuration provides triple data for retrieval model evaluation, including queries, passage candidates and reference answers. The dataset supports multiple difficulty levels (high school, university, doctoral) to facilitate curriculum learning and model generalization. The generation process is based on SentenceTransformers and FAISS similarity search, using the Qwen3-Embedding-0.6B embedding model to automatically mine hard negative samples. This dataset aims to improve the development of Spanish legal encoders and promote public access to legal information.




