遇见数据集

T2S-Bench-E2E

收藏
魔搭社区2026-04-03 更新2026-08-23 收录
官方服务:

资源简介:

<div align="center"> <h1> T2S-Bench &amp; Structure-of-Thought</h1> <h3><b>Benchmarking Comprehensive Text-to-Structure Reasoning</b></h3> </div> <p align="center"> 🌐 <a href="https://t2s-bench.github.io/T2S-Bench-Page/" target="_blank">Project Page</a> • 📚 <a href="https://arxiv.org/abs/2603.03790" target="_blank">Paper</a> • 🤗 <a href="https://huggingface.co/T2SBench" target="_blank">T2S-Bench Dataset</a> • 📊 <a href="https://t2s-bench.github.io/T2S-Bench-Page/#leaderboard" target="_blank">Leaderboard</a> • 🔮 <a href="https://t2s-bench.github.io/T2S-Bench-Page/#examples" target="_blank">Examples</a> </p> T2S-Bench is a comprehensive benchmark for evaluating models' ability to extract structured representations from scientific text. It includes three curated components: T2S-Train-1.2k for training, T2S-Bench-MR (500 samples) for multi-hop reasoning, and T2S-Bench-E2E (87 samples) for end-to-end structuring. Covering 6 scientific domains, 17 subfields, and 32 structure types, T2S-Bench provides high-quality, structure-grounded samples drawn from peer-reviewed academic papers. Every sample underwent 6K+ model search, 6 rounds of validation, and 3 rounds of human review, ensuring correctness in structure, text, and reasoning logic. T2S-Bench is organized into three subsets: | Subset | Size | Dataset (Location) | Goal | Design | Metrics | | ------------------ | ------------: | ------------------------------------------------------------ | ------------------------------------------------------------ | ------------------------------------------------------------ | ------------------------------------------------------------ | | **T2S-Train-1.2k** | 1,200 samples | [T2SBench/T2S-Train-1.2k](https://huggingface.co/datasets/T2SBench/T2S-Train-1.2k) | Provide **verified text–structure pairs** for training / instruction tuning | Multi-hop QA; supports **single-select** & **multi-select** | **Exact Match (EM)**, **F1** | | **T2S-Bench-MR** | 500 samples | [T2SBench/T2S-Bench-MR](https://huggingface.co/datasets/T2SBench/T2S-Bench-MR) | Answer **multi-choice** questions requiring reasoning over an **implicit/explicit structure** extracted from text | Multi-hop QA; supports **single-select** & **multi-select** | **Exact Match (EM)**, **F1** | | **T2S-Bench-E2E** | 87 samples | [T2SBench/T2S-Bench-E2E](https://huggingface.co/datasets/T2SBench/T2S-Bench-E2E) | Extract a **node-link graph** from text that matches the target **key-structure** | Fixes key nodes/links; partially constrains generation to reduce ambiguity | **Node Similarity (SBERT-based)**, **Link F1 (connection-based)** | ## 📄 Citation If you find **T2S-Bench** useful for your research and applications, please cite: ```bibtex @misc{wang2026t2sbenchstructureofthoughtbenchmarking, title={T2S-Bench & Structure-of-Thought: Benchmarking and Prompting Comprehensive Text-to-Structure Reasoning}, author={Qinsi Wang and Hancheng Ye and Jinhee Kim and Jinghan Ke and Yifei Wang and Martin Kuo and Zishan Shao and Dongting Li and Yueqian Lin and Ting Jiang and Chiyue Wei and Qi Qian and Wei Wen and Helen Li and Yiran Chen}, year={2026}, eprint={2603.03790}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2603.03790}, } ```

提供机构:
maas
创建时间:
2026-03-12
二维码
社区交流群
二维码
科研交流群
商业服务