unlearning-cleanslate/eval-checkpoint-80-w100-s10
收藏资源简介:
该数据集是一个用于分析文本记忆和评估模型性能的数据集,包含4663个训练样本,总大小约2.66 GB。数据特征包括文本长度(字符数)、窗口数量、记忆窗口数量、记忆分数、覆盖率、概率统计(如最大、最小、平均、中位数、标准差概率值)、最佳窗口索引及其概率、种子、目标、开始和结束字符位置。评估模型参数包括窗口大小、步长和评估阈值。每个样本还包含窗口列表,其中详细记录了每个窗口的结束字符、索引、是否被记忆、对数概率、目标令牌数量、概率值、种子、开始字符、目标、目标对数概率列表和目标排名列表。此外,数据集还提供了内容ID、标题、创建者和年份等元信息。该数据集适用于自然语言处理中的模型评估、记忆分析等任务。
This dataset is designed for analyzing text memorization and evaluating model performance, containing 4663 training examples with a total size of approximately 2.66 GB. The features include text length (in characters), number of windows, number of memorized windows, memorized fraction, coverage, probability statistics (such as maximum, minimum, mean, median, and standard deviation probabilities), best window index and its probability, seed, target, start and end character positions. Evaluation model parameters include window size, stride, and evaluation threshold. Each example also includes a list of windows, detailing each windows end character, index, whether it is memorized, log probability, number of target tokens, probability value, seed, start character, target, target log probabilities list, and target ranks list. Additionally, metadata such as content ID, title, creators, and year are provided. This dataset is suitable for tasks in natural language processing, including model evaluation and memorization analysis.




