unlearning-cleanslate/fsid-curated-gemma-12b
收藏资源简介:
该数据集用于研究机器学习模型中的遗忘和记忆行为,特别关注文本生成任务。它包含多个配置:forget配置涉及模型对特定内容的记忆分数分析,包括请求ID、内容ID、标题、窗口索引、前缀、后缀、记忆比例和规则名称等特征;forget_pool配置提供内容池信息,如内容ID、标题、创作者、年份、歌词、记忆比例、最大概率等,用于训练;retain配置关注模型保留的文本内容,包括文本和规则名称;retain_pool配置则包含详细的评估指标,如文本长度、记忆窗口数、记忆比例、ROUGE-L分数、困惑度统计等,用于分析模型在生成任务中的表现。数据集支持多个分片(如baseline、bm25_10B、igm_10B),可能对应不同的检索或生成策略,适用于评估模型记忆机制、遗忘学习算法和文本生成性能。
This dataset is designed for studying forgetting and memorization behaviors in machine learning models, with a focus on text generation tasks. It includes multiple configurations: the forget configuration analyzes memorization scores of models for specific content, featuring request ID, content ID, title, window index, prefix, suffix, memorized fraction, and rule name; the forget_pool configuration provides a content pool with details such as content ID, title, creators, year, lyrics, memorized fraction, and maximum probability, used for training; the retain configuration focuses on text content retained by models, including text and rule name; the retain_pool configuration contains detailed evaluation metrics like text length, number of memorized windows, memorized fraction, ROUGE-L scores, perplexity statistics, etc., for analyzing model performance in generation tasks. The dataset supports multiple splits (e.g., baseline, bm25_10B, igm_10B), likely corresponding to different retrieval or generation strategies, and is suitable for evaluating memory mechanisms, forgetting learning algorithms, and text generation performance.



