SUITE (Selective Unlearning of Isolated Topics and Events)
收藏资源简介:
SUITE是由图宾根人工智能中心构建的细粒度评估协议与训练语料库,专为大型语言模型的机器遗忘任务设计。该数据集涵盖真实世界的事实性知识领域,每个遗忘主题包含25个核心事实,通过直接、反向和间接三种查询模态进行多维度测试,并引入语义分级、句法结构和词汇敏感性的保留集探针。数据构建过程采用人工审核与大语言模型生成相结合的方式,确保训练集与评估集严格分离,并覆盖了多种查询变体和推理路径。该数据集旨在解决机器遗忘中的不对称泛化问题,精准刻画遗忘与保留的边界,为评估模型在移除特定知识时是否产生欠遗忘或过遗忘现象提供标准化测试框架。
SUITE is a fine-grained evaluation protocol and training corpus developed by the Tübingen AI Center, specifically designed for machine forgetting tasks in large language models. This dataset covers real-world factual knowledge domains, with each forgetting topic containing 25 core facts. It conducts multi-dimensional testing via three query modalities: direct, reverse, and indirect, and introduces retention set probes that incorporate semantic hierarchy, syntactic structure, and lexical sensitivity. The data construction process combines manual review and large language model generation, ensuring strict separation between the training and evaluation sets while covering diverse query variants and reasoning paths. This dataset aims to address the asymmetric generalization problem in machine forgetting, accurately characterize the boundary between forgetting and retention, and provide a standardized testing framework for evaluating whether models suffer from under-forgetting or over-forgetting when specific knowledge is removed.

- 1Forget Narrowly, Retain Broadly: Unlearning as an Asymmetric Generalization Problem图宾根大学·图宾根人工智能中心 · 2026年



