遇见数据集

NMT-4K-EEG: A Curated Clinical EEG Dataset for Normal and Abnormal EEG Detection

收藏
Zenodo2026-06-19 更新2026-06-21 收录
官方服务:

资源简介:

Dataset Overview NMT-4K-EEG is a curated clinical electroencephalography dataset prepared for research on normal and abnormal EEG detection, clinical EEG interpretation, and machine learning based EEG analysis. The dataset contains EEG recordings in EDF format, de-identified clinical reports, abnormal case annotation files, and structured metadata for reproducible dataset use. The dataset is organized into fixed training and evaluation splits. Each recording is linked to an anonymized recording identifier, recording level interpretation, demographic metadata, report file, and file integrity information. The structure is designed to support transparent model development, reproducible benchmarking, and reliable reuse by the research community. Dataset Contents The dataset includes 4,473 EEG recordings divided into training and evaluation sets. Training split: Normal recordings: 2,787 EDF files and 2,787 report files Abnormal recordings: 686 EDF files, 686 annotation CSV files, and 686 report files Evaluation split: Normal recordings: 540 EDF files and 540 report files Abnormal recordings: 460 EDF files, 460 annotation CSV files, and 460 report files The abnormal recordings include annotation CSV files. The normal recordings include EDF recordings and linked report files. All recordings are organized by split and interpretation label. Metadata The metadata folder contains dataset level files that describe the recordings, file linkage, and file integrity information. The recordings.tsv file contains one row per recording and includes the anonymized recording identifier, split assignment, hospital or source code, recording year, date, recorded gender, age information, and recording level interpretation. The report_linkage.tsv file maps each recording identifier to its corresponding de-identified report file and relative report path. The sha256.txt file provides SHA256 checksums for released files so that users can verify file integrity after download or transfer. Recommended Use The training split is intended for model development, training, and internal validation. The evaluation split is intended for final testing and should remain separate from the training process to reduce the risk of data leakage. Users should rely on metadata/recordings.tsv as the official source for recording identifiers, labels, and split assignments. Report files should be linked using metadata/report_linkage.tsv, and file integrity should be checked using metadata/sha256.txt. Data Organization The dataset follows a clear folder structure with separate directories for training and evaluation data. Normal and abnormal recordings are stored separately. EDF recordings, reports, and annotation files are stored in their respective folders. This structure allows users to load the dataset directly for machine learning experiments while also preserving the connection between each recording, its label, its report, and its annotation file where available. Privacy and Responsible Use The dataset is intended for research use in de-identified form. Users must not attempt to re-identify individuals represented in the data. Users are responsible for handling the dataset according to applicable ethical, institutional, and legal requirements. This dataset is provided for scientific research and should not be used as a substitute for clinical decision making without appropriate expert review and validation. Version Dataset name: NMT-4K-EEG Version: 1.0 Maintainer: Hira Masood Institution: National University of Sciences and Technology, NUST

提供机构:
Zenodo
创建时间:
2026-06-19
二维码
社区交流群
二维码
科研交流群
商业服务