SINAI/ALIA-es-legal-triplets
收藏资源简介:
ALIA西班牙法律与行政三元组语料库,源自SINAI/ALIA-es-legal数据集,包含表格实例,旨在通过使用基于段落的查询-答案数据训练和评估检索导向模型(如密集检索器/嵌入编码器)。该数据采用Qwen3风格的提示工作流程生成,保留了原始文档和段落的来源,并提供了问题类型和难度(从高中到博士水平)等控制参数。数据集专注于特定领域的法律行政文本,并兼容跟踪文档/段落来源的文档分割工作流。
The ALIA Spanish Legal and Administrative Triplets Corpus, derived from the SINAI/ALIA-es-legal dataset, contains tabular instances designed to train and evaluate retrieval-oriented models (e.g., dense retrievers / embedding encoders) using passage-grounded query–answer data produced with a Qwen3-style prompting workflow. It preserves provenance to the original document and chunk while exposing controls such as question type and difficulty (ranging from high_school to phd level). The dataset is focused on domain-specific legal-administrative text and compatible with document segmentation workflows that track document/chunk provenance.




