unlearning-cleanslate/eval-17-debug-qwen3-8b-simnpo-gentle-baseline-target-100-localtrain-checkpoint-1
收藏资源简介:
该数据集包含多个特征,主要涉及文本长度、窗口数量、记忆窗口、记忆比例、覆盖率以及各种概率统计指标(如最大、平均、中位数、最小和标准差的概率值)。此外,还包括最佳窗口的索引、概率、种子、目标、起始和结束字符等信息。数据集还记录了评估模型、窗口大小、步长、评估阈值等参数。每个窗口的详细信息如结束字符、索引、是否记忆、对数概率、目标令牌数量、概率值、种子、起始字符、目标、目标对数概率和目标排名等也被包含。数据集还提供了内容ID、标题、创作者和年份等元数据。数据集分为训练集,包含4663个示例,总大小为2666928705字节。
The dataset includes multiple features related to text length, number of windows, memorized windows, memorized fraction, coverage, and various probability statistics (such as max, mean, median, min, and std probability values). It also contains information about the best windows index, probability, seed, target, start and end characters. The dataset records parameters like the evaluation model, window size, stride, and evaluation threshold. Detailed information for each window, such as end character, index, is_memorized, log probability, number of target tokens, probability value, seed, start character, target, target log probabilities, and target ranks, is included. The dataset also provides metadata like content ID, title, creators, and year. The dataset is split into a training set with 4663 examples and a total size of 2666928705 bytes.




