遇见数据集

DW-Bench: Benchmarking LLMs on Data Warehouse Graph Topology Reasoning

收藏
Zenodo2026-04-21 更新2026-05-26 收录
官方服务:

资源简介:

DW-Bench is a benchmark suite for evaluating large language models on graph topology reasoning tasks over data warehouse schemas. It contains 1,046 questions across 5 real-world and synthetic datasets (AdventureWorks, TPC-DS, TPC-DI, OMOP CDM, Syn-Logistics), organized into 13 subtypes spanning lineage tracing, FK routing, and silo detection. Each dataset includes original and obfuscated schema variants. The benchmark includes per-question evaluation results for 3 frontier LLMs (Gemini 2.5 Flash, DeepSeek-V3, Qwen2.5-72B) across 6 baselines (Flat Text, Vector-RAG, Graph-Augmented, Tool-Use, ReAct-Code, Oracle). All schema graphs are provided as PyTorch Geometric HeteroData objects.

提供机构:
Zenodo
创建时间:
2026-04-21
二维码
社区交流群
二维码
科研交流群
商业服务