遇见数据集

Encyclo-K

收藏
魔搭社区2026-04-28 更新2026-07-15 收录
官方服务:

资源简介:

## Encyclo-K Dataset <p align="center"> 🌐 <a href="https://encyclo-k.github.io/">Homepage</a> | 🤗 <a href="https://huggingface.co/datasets/m-a-p/Encyclo-K">Dataset</a> | 📖 <a href="https://arxiv.org/abs/2512.24867">ArXiv</a> | 🏆 <a href="https://encyclo-k.github.io/#:~:text=isolated%20factual%20recall.-,Leaderboard,-We%20evaluate%2050">Leaderboard</a> | 🐱 <a href="https://github.com/multimodal-art-projection/Encyclo-K">GitHub</a> </p> **_Encyclo-K_** is a statement-based benchmark that rethinks benchmark construction from the ground up. Our key observation is that the question itself need not be the atomic unit of curation—individual knowledge statements can be. ### Key Features - **Dynamic Evaluation**: We extract standalone knowledge statements from authoritative textbooks and dynamically compose them into evaluation questions through random sampling at test time. The combinatorial space is too vast to memorize, enabling reliable periodic dataset refresh. - **Multi-Statement Comprehension**: Each question aggregates 8-10 statements for comprehensive multi-knowledge assessment, going beyond what single-statement questions can probe. - **Cost-Effective Annotation**: Annotators only verify formatting compliance without requiring domain expertise, substantially reducing annotation costs. - **Contamination Resistance**: Even if individual statements appear in training data, their compositions form a combinatorial space too vast to memorize. ### Question Distribution The dataset comprises 5,038 questions across 11 disciplines, 44 fields, and 62 subfields. The disciplinary distribution is proportional to statement ratios: Science has the most questions (1,242, 24.7%), while Philosophy has the fewest (61, 1.2%). Each question contains 8–10 statements, 4–8 options, and 2–4 combinations. <p align="center"> <img src="https://cdn-uploads.huggingface.co/production/uploads/6555e8d8a0c34cd61a6b9ce3/0fkmaVaX9i6y902LeRQMu.png" width="75%" alt="Question Distribution"> </p> <!-- ![question_distribution](https://cdn-uploads.huggingface.co/production/uploads/6555e8d8a0c34cd61a6b9ce3/0fkmaVaX9i6y902LeRQMu.png) --> ## 📊 Benchmark Characteristics ### Key Findings <!-- <p align="center"> <img src="https://cdn-uploads.huggingface.co/production/uploads/6555e8d8a0c34cd61a6b9ce3/Ox9rFSCwCn_JBrUrcJZ_h.png" width="42%" alt="Multi-statement Challenge"> <img src="https://cdn-uploads.huggingface.co/production/uploads/6555e8d8a0c34cd61a6b9ce3/5zh06fK4Wn3askhapLjKH.png" width="47%" alt="Random Seed Stability"> </p> --> <!-- Performance comparison between single-statement judgment and multi-statement comprehensive understanding tasks. --> <!-- Model accuracy across five dynamically generated question sets with different random seeds. --> <!-- ![multi_kowledge_statements](https://cdn-uploads.huggingface.co/production/uploads/6555e8d8a0c34cd61a6b9ce3/Ox9rFSCwCn_JBrUrcJZ_h.png) ![random_seed_line_plot](https://cdn-uploads.huggingface.co/production/uploads/6555e8d8a0c34cd61a6b9ce3/5zh06fK4Wn3askhapLjKH.png) --> <table align="center"> <tr> <td><img src="https://cdn-uploads.huggingface.co/production/uploads/6555e8d8a0c34cd61a6b9ce3/Ox9rFSCwCn_JBrUrcJZ_h.png" width="80%"></td> <td><img src="https://cdn-uploads.huggingface.co/production/uploads/6555e8d8a0c34cd61a6b9ce3/5zh06fK4Wn3askhapLjKH.png" width="100%"></td> </tr> </table> #### 1. Multi-Statement Comprehensive Assessment Each question aggregates 8–10 knowledge statements, requiring models to jointly comprehend multiple knowledge points rather than isolated factual recall. This design introduces significant cognitive complexity beyond simple statement-level verification. #### 2. Dynamic Question Generation Encyclo-K supports dynamic question generation by varying random seeds that control statement selection and combination. Model rankings remain highly consistent across different question sets, confirming that the combinatorial design creates a vast question space resistant to memorization-based shortcuts. This enables periodic dataset refresh to prevent overfitting. ## 📈 Experimental Results We evaluate 50+ LLMs on Encyclo-K. The benchmark poses substantial challenges with strong discriminative power: | Model Type | Best Model | Accuracy | Range | |:----------:|:----------:|:--------:|:-----:| | Chat | Qwen3-235B-A22B-Instruct | 50.40% | 9.71% – 50.40% | | Reasoning | OpenAI-GPT-5.1-high | 62.07% | 16.04% – 62.07% | 👉 **For complete leaderboard and more model results, please visit our [Homepage](https://encyclo-k.github.io/#:~:text=isolated%20factual%20recall.-,Leaderboard,-We%20evaluate%2050).** ## 🛠️ Dataset Maintenance Despite multiple rounds of manual review, there may still be a small number of errors in the dataset. If you find any, please paste the `question_id` and `statement` index to the [Issues](https://github.com/multimodal-art-projection/Encyclo-K/issues) page, and we will make the corresponding corrections. Our team is committed to long-term maintenance of this dataset to ensure its quality! ## 📚 Citation If you find Encyclo-K useful in your research, please cite our paper: ```bibtex @article{liang2025encyclo0k0, title = {Encyclo-K: Evaluating LLMs with Dynamically Composed Knowledge Statements}, author = {Yiming Liang and Yizhi Li and Yantao Du and Ge Zhang and Jiayi Zhou and Yuchen Wu and Yinzhu Piao and Denghui Cao and Tong Sun and Ziniu Li and Li Du and Bo Lei and Jiaheng Liu and Chenghua Lin and Zhaoxiang Zhang and Wenhao Huang and Jiajun Zhang}, year = {2025}, journal = {arXiv preprint arXiv: 2512.24867} } ```

提供机构:
maas
创建时间:
2026-01-02
二维码
社区交流群
二维码
科研交流群
商业服务