遇见数据集

unlearning-cleanslate/eval-10-llama-3_1-8b-simnpo-gentle-bm25-6t-target-100-checkpoint-187

收藏
Hugging Face2026-04-29 更新2026-05-03 收录
官方服务:

资源简介:

该数据集包含文本记忆化评估的统计特征和详细窗口信息。特征包括文本长度(字符数)、窗口数量、被记忆的窗口数量、记忆化分数、覆盖率、概率z值(最大、平均、中位数、最小、标准差)、最佳窗口的索引、概率、种子和目标文本等。每个窗口记录起始字符、索引、是否被记忆、对数概率、目标令牌数、概率z值、种子、目标文本、目标令牌的对数概率和排名。此外还包含内容ID、标题、创作者和年份。数据集共有4663条训练样本。

This dataset contains statistical features and detailed window information for evaluating text memorization. Features include text length in characters, number of windows, number of memorized windows, memorization fraction, coverage, probability z-values (max, mean, median, min, std), best window index, probability, seed, and target text. Each window records start character, index, memorization flag, log probability, number of target tokens, probability z-value, seed, target text, target log probabilities, and target ranks. It also includes content ID, title, creators, and year. The dataset has 4663 training examples.

提供机构:
unlearning-cleanslate
二维码
社区交流群
二维码
科研交流群
商业服务