unlearning-cleanslate/eval-16-debug-llama-3_1-8b-simnpo-gentle-baseline-target-100-localtrain-checkpoint-1
收藏资源简介:
该数据集包含多个特征,主要用于分析文本窗口的记忆情况。特征包括文本长度字符数、窗口数量、记忆窗口数量、记忆分数、覆盖率、各种概率统计值(最大、平均、中位数、最小、标准差)、最佳窗口索引及其相关属性(概率、种子、目标、起始和结束字符位置)、评估模型、窗口大小、步长、评估阈值等。此外,还包含每个窗口的详细信息(如结束字符、索引、是否记忆、对数概率、目标令牌数、概率、种子、起始字符、目标、目标对数概率列表和目标排名列表)以及内容ID、标题、创作者和年份。数据集分为训练集,包含4663个样本,总大小为2665222676字节。
The dataset includes multiple features primarily used for analyzing the memorization of text windows. Features include text length in characters, number of windows, number of memorized windows, memorized fraction, coverage, various probability statistics (max, mean, median, min, std), best window index and its related attributes (probability, seed, target, start and end character positions), evaluation model, window size, stride, evaluation threshold, etc. Additionally, it contains detailed information for each window (such as end character, index, is memorized, log probability, number of target tokens, probability, seed, start character, target, target log probabilities list, and target ranks list) as well as content ID, title, creators, and year. The dataset is split into a training set with 4663 examples and a total size of 2665222676 bytes.




