onlytojay/memoir-c0-data
收藏资源简介:
MEMOIR C₀数据集是一个预缓存的JSONL数据集,专门用于计算知识编辑(如MEMIT和AlphaEdit方法)中的C₀协方差矩阵。该数据集包含多个子集:共享文件夹中包括来自不同来源的模型无关公共数据集,如维基百科、WikiText-103、OLMo Dolma维基分割、Dolma starcoder代码、Dolma代数堆栈和开放网络数学、Dolma pes2o科学文本以及Dolma DCLM网络数据;此外,还有针对特定模型(如OLMo-2-1124-7B-Instruct、Llama-3.1-8B-Instruct和Qwen3-8B)自生成的数据文件。每个JSONL文件包含100,000个样本,格式为带有text字段的新行分隔JSON,用于提供文档内容,以支持MEMIT知识编辑中的C₀计算流水线。
The MEMOIR C₀ Data is a pre-cached JSONL dataset designed for computing C₀ covariance matrices in knowledge editing, specifically for methods like MEMIT and AlphaEdit. It consists of multiple subsets: the shared folder contains model-agnostic public datasets from various sources, including Wikipedia, WikiText-103, OLMo Dolma wiki split, Dolma starcoder code, Dolma algebraic-stack and open-web-math, Dolma pes2o scientific text, and Dolma DCLM web data; additionally, there are self-generated data files for specific models such as OLMo-2-1124-7B-Instruct, Llama-3.1-8B-Instruct, and Qwen3-8B. Each JSONL file contains 100,000 samples in newline-delimited JSON format with a single text field, providing document content to support the C₀ computation pipeline in MEMIT-based knowledge editing.




