unlearning-cleanslate/fsid-curated-llama-8b-target-100
收藏资源简介:
该数据集包含多个配置,主要涉及遗忘(forget)和保留(retain)相关的数据。forget配置包含请求ID、内容ID、内容标题、窗口索引、前缀、后缀、记忆分数和规则名称等特征,分为基线、bm25_10B、bm25_6T和igm_10B等分割。forget_pool配置包含内容ID、标题、创作者、年份、歌词、记忆分数等特征,主要用于训练。retain配置包含文本和规则名称特征,同样分为多个分割。retain_pool配置包含文本长度、窗口数量、记忆窗口、记忆分数、复制窗口、复制分数等特征,主要用于训练,并包含详细的评估指标和窗口信息。
This dataset contains multiple configurations, mainly involving data related to forget and retain. The forget configuration includes features such as request ID, content ID, content title, window index, prefix, suffix, memory score and rule name, and is divided into splits including baseline, bm25_10B, bm25_6T and igm_10B. The forget_pool configuration includes features such as content ID, title, creator, year, lyrics, memory score, etc., and is mainly used for training. The retain configuration includes text and rule name features, and is also divided into multiple splits. The retain_pool configuration includes features such as text length, number of windows, memory window, memory score, copy window and copy score, which is mainly used for training, and contains detailed evaluation metrics and window information.



