遇见数据集

DAISLab-Unisa/it-rag-bench

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

IT-RAG-Bench是一个合成的意大利语检索基准数据集,旨在评估密集嵌入模型在意大利语的文档检索和检索增强生成(RAG)任务上的性能。数据集包含3,200个意大利语段落和640个自然语言查询,覆盖三种典型的意大利信息检索场景文档风格:百科式(维基百科风格,涉及AI、NLP和意大利监管主题)、常见问题解答(公共管理问答对)以及法律/法规文章摘录(意大利立法风格,采用Art. N格式)。数据集使用固定随机种子(42)生成,基于意大利语主题词汇列表(如人工智能、神经网络、信息检索、GDPR法规等)和参数化模板,确保完全可重复性,不依赖爬取或授权数据。数据集结构包括三个配置:corpus(段落库,3,200行)、queries(查询,640行)和qrels(查询-文档相关性判断,1,246行)。相关性得分为二进制(0或1),每个查询有1到3个相关文档。该数据集专为比较评估设计(例如模型排名),由于相关性标签是随机分配的,绝对检索性能指标可能低于人工标注基准。它适用于意大利语企业检索场景的评估,如行政门户、法律数据库和公共常见问题解答。

IT-RAG-Bench is a synthetic Italian-language retrieval benchmark designed to evaluate dense embedding models on document retrieval and Retrieval-Augmented Generation (RAG) tasks in Italian. The dataset provides 3,200 Italian passages and 640 natural-language queries spanning three document styles representative of real Italian information retrieval scenarios: encyclopedic passages (Wikipedia-style, covering AI, NLP, and Italian regulatory topics), FAQ passages (public-administration question–answer pairs), and legal/regulatory article excerpts (Italian legislative style in Art. N format). It was generated synthetically with a fixed random seed (42) using Italian-language topic vocabularies (e.g., artificial intelligence, neural networks, information retrieval, GDPR regulations) drawn from AI/NLP, legal, and public-administration domains, ensuring full reproducibility without relying on crawled or licensed data. The dataset structure includes three configurations: corpus (3,200 passages), queries (640 queries), and qrels (1,246 query–document relevance judgements). Relevance scores are binary (0 or 1), with each query having between 1 and 3 relevant documents. Designed for comparative evaluation (e.g., ranking models against each other), the dataset is best used for benchmarking in Italian enterprise retrieval scenarios such as administrative portals, legal databases, and public FAQs, though absolute metric scores may be lower than on human-annotated benchmarks due to randomly assigned relevance labels.

提供机构:
DAISLab-Unisa
二维码
社区交流群
二维码
科研交流群
商业服务