unlearning-cleanslate/eval-llama-3_1-8b-post_books_2025_v2
收藏数据链接:
官方服务:
资源简介:
该数据集用于评估语言模型在文本窗口上的记忆性,包含文本长度、窗口数量、记忆分数、覆盖度、概率分布统计等特征,以及每个窗口的详细信息(如起始字符、索引、是否被记忆、对数概率等)。同时包含内容元数据(ID、标题、创建者、年份)和评估参数(模型名称、窗口大小、步长、阈值)。
This dataset is used to evaluate the memorization behavior of language models on text windows. It includes features such as text length, number of windows, memorized fraction, coverage, probability statistics, and detailed window-level information (e.g., start character, index, is_memorized, log probabilities, etc.). It also contains content metadata (ID, title, creators, year) and evaluation parameters (model name, window size, stride, threshold).



