遇见数据集

unlearning-cleanslate/eval-olmo-3-7b-undial-baseline

收藏
Hugging Face2026-04-29 更新2026-05-03 收录
官方服务:

资源简介:

该数据集用于评估语言模型对文本片段的记忆程度,包含文本长度、窗口数量、记忆窗口数量、记忆比例、覆盖率、p_z统计量(最大、平均、中位数、最小、标准差)、最佳窗口索引及其p_z、种子、目标文本、起始和结束字符位置、评估模型、窗口大小、步长、评估阈值,以及每个窗口的详细信息(结束字符、索引、是否记忆、对数概率、目标token数、p_z、种子、起始字符、目标、目标对数概率、目标排名等)。此外还包括内容ID、标题、创作者和年份。数据集主要用于研究模型记忆行为。

This dataset is used to evaluate the memorization degree of language models on text fragments. It includes features such as text length, number of windows, memorized windows, memorized fraction, coverage, p_z statistics (max, mean, median, min, std), best window index and its p_z, seed, target text, start and end character positions, evaluation model, window size, stride, evaluation threshold, and detailed information for each window (end character, index, is_memorized, log probability, number of target tokens, p_z, seed, start character, target, target log probabilities, target ranks, etc.). It also includes content ID, title, creators, and year. The dataset is primarily used to study model memorization behavior.

提供机构:
unlearning-cleanslate
二维码
社区交流群
二维码
科研交流群
商业服务