遇见数据集

unlearning-cleanslate/eval-09-qwen3-8b-simnpo-gentle-igm-10b-target-100-checkpoint-355

收藏
Hugging Face2026-04-29 更新2026-05-03 收录
官方服务:

资源简介:

该数据集用于评估语言模型在训练数据中的记忆行为。每个样本包含一段文本内容(由content_id、title、creators、year标识),以及对该文本进行窗口划分后,每个窗口被模型记忆的概率(p_z)和相关统计信息。数据集还提供了最佳记忆窗口的种子、目标文本及其位置,以及评估模型、窗口大小、步长等参数。旨在分析模型是否记住了训练数据中的特定片段。

This dataset is designed to evaluate the memorization behavior of language models on training data. Each sample includes a piece of text content (identified by content_id, title, creators, year), along with windowed partitions of the text, the probability of each window being memorized (p_z), and related statistics. It also provides the best memorized windows seed, target text, and its position, as well as evaluation model, window size, stride, and threshold parameters. The dataset aims to analyze whether models have memorized specific fragments of the training data.

提供机构:
unlearning-cleanslate
二维码
社区交流群
二维码
科研交流群
商业服务