遇见数据集

RHELM

收藏
魔搭社区2026-06-12 更新2026-07-19 收录
官方服务:

资源简介:

# RHELM: Beyond Static Dialogues **Benchmarking Realistic, Heterogeneous, and Evolving Long-Horizon Memory** [![Paper](https://img.shields.io/badge/Paper-arXiv-red)](https://arxiv.org/pdf/2605.31086) [![HuggingFace](https://img.shields.io/badge/🤗%20HuggingFace-Dataset-yellow)](https://huggingface.co/datasets/microsoft/RHELM) [![GitHub](https://img.shields.io/badge/Code-GitHub-black)](https://github.com/Hanzhang-lang/RHELM_Benchmark) RHELM is a benchmark for evaluating long-horizon memory capabilities in AI assistants. Unlike benchmarks built around static dialogues, RHELM provides **realistic**, **heterogeneous**, and **temporally evolving** memory sources, together with challenging questions that require multi-hop reasoning, temporal synthesis, and hallucination detection. > ⚠️ All characters, events, and personal details in this dataset are **fully synthetic**. > Any resemblance to real individuals is coincidental. ## Dataset Summary | Item | Count | |------|-------| | Characters (personas) | 10 | | QA pairs | 1,305 | | Conversation sessions (`.json`) | 629 | | Emails (`.txt`) | 625 | | Attachments (`.md` / `.html`) | 1,053 | ### Question types | Type | Count | |------|-------| | attachment | 249 | | mixed | 210 | | fact | 207 | | hallucination | 197 | | aggregation | 192 | | temporal | 185 | | misleading | 65 | ## Directory Structure ``` data/ (uploaded to repo root) ├── conversations/<Character>/*.json # dated dialogue sessions ├── emails/<Character>/*.txt # email threads ├── attachments/<Character>/*.md|*.html# documents, notes, reports └── QA_final/low_score_qa_<Character>_all_validated.jsonl ``` ## QA Schema Each line in a `QA_final/*.jsonl` file is a JSON object: | Field | Description | |-------|-------------| | `id` | Unique question identifier | | `question` | The user query | | `answer` | Ground-truth answer | | `question_date` | Date the question is asked from | | `question_type` | One of: fact, temporal, hallucination, aggregation, misleading, attachment, mixed | | `supporting_evidence` | References to source items (e.g. `"2024-10-13:1"` or `"56_report_task_*.md:Section"`) | | `characteristics` | Fine-grained challenge labels (see taxonomy) | ## Usage ```python from datasets import load_dataset qa = load_dataset("microsoft/RHELM", data_files="QA_final/*.jsonl", split="train") print(qa[0]) ``` To work with the full multi-source context (conversations, emails, attachments), download the repository snapshot: ```python from huggingface_hub import snapshot_download local_dir = snapshot_download("microsoft/RHELM", repo_type="dataset") ``` ## Challenge Taxonomy RHELM organizes questions into **7 categories** with **26 challenge characteristics** across three QA domains: Dialogue History QA, External Source QA, and Hybrid Context QA. See the [evaluation code repository](https://github.com/Hanzhang-lang/RHELM_Benchmark) for the full taxonomy and benchmark harness. ## License Released under [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/).

提供机构:
maas
创建时间:
2026-06-03
二维码
社区交流群
二维码
科研交流群
商业服务