unlearning-cleanslate/eval-11-qwen3-8b-simnpo-gentle-bm25-6t-target-100-checkpoint-173
收藏资源简介:
该数据集用于评估语言模型在文本片段上的记忆行为,包含每个文本窗口的记忆状态、概率分数、最佳窗口的种子和目标文本等统计信息,以及相关文本内容的元数据(如标题、创作者、年份)。数据集中每个样本代表一个文本,记录了其字符长度、窗口数量、被记忆的窗口数量、记忆分数、覆盖率、p_z统计量(最大、平均、中位数、最小、标准差),以及评估模型、窗口大小、步长、评估阈值等参数。此外,每个样本包含一个详细的窗口列表,包括每个窗口的起始/结束字符、索引、是否被记忆、对数概率、目标令牌数、概率、种子和目标文本等。
This dataset is used to evaluate the memorization behavior of language models on text segments. It contains statistics for each text window, including memorization status, probability scores, best window seed and target text, as well as metadata of the original content (title, creators, year). Each sample in the dataset represents a text, recording its character length, number of windows, number of memorized windows, memorized fraction, coverage, p_z statistics (max, mean, median, min, std), and evaluation parameters such as model, window size, stride, and threshold. Additionally, each sample includes a detailed list of windows with start/end characters, index, whether memorized, log probability, number of target tokens, probability, seed, and target text.



