遇见数据集

tankalapavankalyan/eeg-corpus-manifest

收藏
Hugging Face2026-05-20 更新2026-05-31 收录
官方服务:

资源简介:

EEG规范语料库索引是一个可查询的索引,覆盖了79,313小时的无损脑电图(EEG)数据,这些数据以统一的.eegz(Zarr v3)格式存储。索引仅包含元数据,实际信号字节存储在亚马逊S3上的规范Zarr存储中,可按需读取。数据集包含185,802条录音记录,涉及43,114名受试者和411个源数据集,规范磁盘大小为2.81 TB,相比源EEG数据实现了3.05倍的压缩。数据以四个Parquet表的形式组织:datasets(源数据集信息,包括许可证、DOI、范式摘要等)、recordings(事实表,每条录音记录包含规范URI、归档URI、持续时间、通道数等)、channels(每条录音的每个通道信息,包括名称、类型、单位、电极状态和3D坐标)和subjects(每个哈希受试者的信息,包括年龄、性别、临床状态等)。数据集支持多种查询工具(如Pandas、DuckDB、Polars),并提供了数据转换的往返合同(包括bit_exact和near_lossless两种类型),确保数据完整性。语料库混合了手动整理的锚数据集(如HBN-EEG、PEERS Memory EEG、TUH-EEG Corpus和OpenNeuro EEG数据集)以及用于基础模型基准测试的数据集(如PhysioNet系列和Mumtaz数据集)。数据许可遵循原始来源的许可证,索引本身使用CC-BY-4.0许可。数据集旨在为EEG基础模型预训练提供统一、可访问的数据资源。

The Standardized EEG Corpus Index is a queryable index covering 79,313 hours of lossless electroencephalography (EEG) data stored in a unified .eegz (Zarr v3) format. The index only contains metadata, while the actual signal bytes are stored in a standardized Zarr repository on Amazon S3 and can be read on demand. The dataset comprises 185,802 recording entries involving 43,114 subjects and 411 source datasets, with a standardized disk size of 2.81 TB and a 3.05-fold compression ratio compared to the raw source EEG data. The dataset is organized into four Parquet tables: 1. "datasets": Contains source dataset information including license, DOI, paradigm summary and other relevant details; 2. "recordings": A fact table where each recording entry includes standardized URI, archive URI, duration, number of channels and other metadata; 3. "channels": Stores information for each channel of every recording, including name, type, unit, electrode status and 3D coordinates; 4. "subjects": Contains information for each hashed subject, including age, gender, clinical status and other demographic or clinical details. The dataset supports multiple query tools such as Pandas, DuckDB and Polars, and provides round-trip data transformation contracts with two types: bit_exact and near_lossless, to ensure full data integrity. The corpus combines manually curated anchor datasets (e.g., HBN-EEG, PEERS Memory EEG, TUH-EEG Corpus and OpenNeuro EEG datasets) as well as datasets intended for foundation model benchmarking (e.g., the PhysioNet series and the Mumtaz dataset). Data licensing follows the original source licenses, while the index itself is licensed under CC-BY-4.0. This dataset aims to provide a unified and accessible data resource for pre-training EEG foundation models.

提供机构:
tankalapavankalyan
二维码
社区交流群
二维码
科研交流群
商业服务