遇见数据集

unlearning-cleanslate/eval-13-qwen3-8b-undial-baseline-target-100-checkpoint-1078

收藏
Hugging Face2026-04-29 更新2026-05-03 收录
官方服务:

资源简介:

该数据集用于评估语言模型对文本的记忆化程度。包含文本长度、窗口数量、记忆化窗口数量、记忆化比例、覆盖率、概率统计(最大、平均、中位数、最小、标准差)以及每个窗口的详细信息(如起始字符、索引、是否被记忆化、对数概率、种子、目标文本等)。数据集共有4663个训练样本,总大小约2.6GB。

This dataset is used to evaluate the memorization degree of language models on text. It includes features such as text length, number of windows, number of memorized windows, memorized fraction, coverage, probability statistics (max, mean, median, min, std), and detailed information for each window (e.g., start character, index, whether memorized, log probability, seed, target text). The dataset contains 4663 training examples with a total size of approximately 2.6GB.

提供机构:
unlearning-cleanslate
二维码
社区交流群
二维码
科研交流群
商业服务