June30916/multimodality-poc-llama31-ruler16k
收藏资源简介:
这是一个用于KV-cache压缩研究的预RoPE查询/隐藏状态语料库,基于Llama-3.1-8B-Instruct模型在RULER-16K数据集上的预填充过程捕获的原始张量。数据集包含65个.npz文件,每个文件约414MB,总大小约26GB。每个文件包含隐藏状态、预RoPE查询、子采样预填充位置、任务名称、提示索引和原始提示长度等信息。数据集用于研究每个(层,kv_head)查询分布是否是单峰高斯分布,这是Expected Attention的MGF闭式假设的基础。
A pre-RoPE query / hidden-state corpus for KV-cache compression research, consisting of raw tensors captured during prefill of the Llama-3.1-8B-Instruct model on the RULER-16K dataset. The dataset includes 65 .npz files, each ~414 MB, totaling ~26 GB. Each file contains hidden states, pre-RoPE queries, sub-sampled prefill positions, task name, prompt index, and original prompt length. The dataset is used to study whether the per-(layer, kv_head) query distribution is unimodal Gaussian, the assumption underpinning Expected Attentions MGF closed-form.




