unlearning-cleanslate/eval-nemotron-nano-9b-v2-simnpo-baseline
收藏资源简介:
该数据集是一个用于分析文本记忆或相似性检测的数据集,包含4663个训练示例,总大小约为2.69 GB。数据集的特征包括文本长度(字符数)、窗口数量、记忆窗口数量、记忆比例、覆盖率、概率统计值(如最大、平均、中位数、最小和标准差的p_z值)、最佳窗口索引和相关信息(如种子、目标文本、起始和结束字符位置)、评估模型参数(如窗口大小、步长、评估阈值),以及内容元数据(如内容ID、标题、创建者和年份)。每个示例还包含一个窗口列表,详细记录每个窗口的结束字符、索引、是否被记忆、对数概率、目标令牌数量、p_z值、种子、起始字符、目标文本、目标对数概率和目标排名。数据集旨在支持NLP任务,如文本生成、记忆评估或相似性分析,但具体应用背景未在README中明确说明。
This dataset is designed for analyzing text memorization or similarity detection, containing 4663 training examples with a total size of approximately 2.69 GB. The features include text length (in characters), number of windows, memorized windows count, memorized fraction, coverage, probability statistics (such as max, mean, median, min, and std p_z values), best window index and related information (e.g., seed, target text, start and end character positions), evaluation model parameters (like window size, stride, eval threshold), and content metadata (such as content ID, title, creators, and year). Each example also includes a list of windows detailing each windows end character, index, is_memorized status, log probability, number of target tokens, p_z value, seed, start character, target text, target log probabilities, and target ranks. The dataset is intended to support NLP tasks such as text generation, memorization evaluation, or similarity analysis, but the specific application context is not explicitly stated in the README.




