reve-dataset
收藏资源简介:
REVE预训练数据集(开放子集)是一个用于训练EEG(脑电图)大规模基础模型的开放访问数据集。该子集包含所有允许重新分发和开放研究的记录,总大小约为4.4TB,占整个REVE预训练数据量的一半。数据集以NumPy二进制文件(.npy)格式存储,并针对内存映射(memmap)进行了优化,以便在不将整个多TB数据集加载到RAM的情况下实现高效访问。该数据集旨在用于自监督学习(如掩码自编码器、对比学习)、评估EEG基础模型的迁移能力以及研究数据量对神经信号解码的影响。数据集汇集了多个主要的开源EEG存储库,并强调了引用原始来源的重要性。
The REVE Pre-training Dataset (Open Subset) is an openly accessible dataset for training large-scale foundation models for EEG (Electroencephalography). This subset includes all records that permit redistribution and open research, with a total size of approximately 4.4 TB, accounting for half of the entire REVE pre-training dataset. The dataset is stored in NumPy binary file (.npy) format, and optimized for memory mapping (memmap) to enable efficient access without loading the entire multi-TB dataset into RAM. This dataset is designed for use in self-supervised learning (e.g., masked autoencoders, contrastive learning), evaluating the transfer capabilities of EEG foundation models, and researching the impact of data scale on neural signal decoding. The dataset aggregates multiple major open-source EEG repositories, and emphasizes the importance of citing their original sources.




