遇见数据集

The NMT Scalp EEG Dataset: An Open-Source Annotated Dataset of Healthy and Pathological EEG Recordings for Predictive Modeling

收藏
Zenodo2026-07-27 更新2026-08-02 收录
官方服务:

资源简介:

Dataset Overview The NMT Scalp EEG Dataset is an open source collection of 2,417 anonymised scalp electroencephalogram recordings obtained from unique participants. The collection contains approximately 625 hours of EEG data and focuses on a South Asian clinical population. Each recording has been assigned a record level label of either normal or abnormal based on neurological assessment. The current release contains 2,002 normal EEG recordings and 415 abnormal or pathological EEG recordings. Demographic information is also provided, including age, gender and the location of each record within the training or evaluation sections of the dataset. Purpose of the Dataset The dataset was developed to support research on automated EEG analysis and the pre-diagnostic screening of normal and abnormal EEG recordings. It addresses the limited availability of large, carefully organised and neurologist labelled EEG datasets from South Asian populations. The data can be used to develop, train and evaluate machine learning and deep learning models for EEG abnormality detection. It also supports research on model generalisation, transfer learning, domain shift, dataset diversity and the effect of acquisition sources on predictive performance. Data Collection The recordings were collected at Pak-Emirates Military Hospital in Rawalpindi, Pakistan. EEG acquisition was performed using the KT88-2400 EEG system manufactured by Contec Medical Systems. The standard international 10-20 electrode placement system was used. The original acquisition setup included 19 scalp channels together with the A1 and A2 ear reference channels. All channels were recorded at a sampling frequency of 200 Hz. The average duration of an EEG recording is approximately 15 minutes. Participant ages range from less than one year to 90 years, which provides broad age representation across paediatric, adult and older populations. Annotation and Quality Control Each EEG recording was initially assessed by hospital neurological staff trained in EEG interpretation. The assigned label was then reviewed by two expert neurologists. A recording was included only when both neurologists agreed on its final normal or abnormal label. Records for which agreement could not be reached were excluded. This review procedure was used to improve the reliability and clinical quality of the ground truth labels. Preprocessing and Signal Format The EEG signals were originally acquired using a linked ear reference at 200 Hz. For improved comparability with other established EEG datasets, the recordings were rereferenced offline using an average reference. The processed EEG records contain 21 channels. All recordings are provided in European Data Format, commonly known as EDF. EDF is an open format that stores the physiological signal together with information such as channel names, channel count, sampling frequency and filter settings. The date and time values contained inside the EDF files refer to the time at which the files were saved in their present format. They should not be interpreted as the original clinical recording date and time. Dataset Organisation The dataset is organised according to the clinical label and the intended experimental split. The abnormal directory contains EEG records labelled as abnormal. It includes separate train and eval subdirectories. The normal directory contains EEG records labelled as normal. It follows the same train and eval structure. The Labels.csv file provides the metadata and ground truth information associated with each EEG record. It contains the following fields. recordname contains the filename of the EEG recording. label contains the neurologist assigned class, either normal or abnormal. age contains the participant's age in years. gender contains male, female or not specified. loc shows whether the recording belongs to the training or evaluation section. Demographic Composition The dataset contains 1,608 recordings from male participants, 808 recordings from female participants and one recording for which gender was not specified. The age distribution covers participants from infancy to 90 years of age. This demographic range makes the dataset relevant for research involving EEG variation across different age groups. The dataset preserves the naturally occurring imbalance between normal and abnormal clinical recordings. Researchers should account for this class distribution when designing training procedures, evaluation measures and data sampling strategies. Potential Research Applications The dataset may be used for binary classification of normal and abnormal EEG recordings, EEG representation learning, signal processing research, deep learning model development and model benchmarking. It is also suitable for studying cross-dataset generalisation, transfer learning, fine-tuning, model bias, population differences and the effect of EEG acquisition equipment on predictive performance. The supplied training and evaluation organisation can support reproducible comparisons between newly developed methods and the baseline experiments reported in the associated publication. Ethics, Consent and Privacy The study was reviewed and approved by the Institutional Review Board of Pak-Emirates Military Hospital, Rawalpindi, Pakistan. The reported IRB reference is 51214MH, dated 15 March 2019. Written informed consent was obtained through the approved study procedure. Personal identity information was removed before the EEG recordings were added to the research repository. The shared data therefore contains anonymised EEG signals and limited demographic metadata rather than directly identifying patient information. Dataset Limitations The current release provides record level labels for normal and abnormal EEG recordings. It does not provide detailed diagnostic categories, seizure event boundaries or fine grained clinical annotations. The data was collected at a single clinical institution using a specific acquisition system. Differences in population demographics, hospital procedures, electrode preparation, recording hardware and signal characteristics may affect the performance of models when they are applied to EEG data from other sources. The class distribution is also unbalanced, with considerably more normal recordings than abnormal recordings. Evaluation methods should therefore consider measures in addition to overall accuracy. Responsible Use This dataset is intended for scientific research, education, algorithm development and reproducible benchmarking. Predictive models trained on this resource should be independently validated on additional clinical populations and acquisition systems before being considered for real clinical use. The dataset should not be treated as a replacement for assessment by a qualified neurologist or other trained medical professional. Associated Publication Khan, H. A., Ul Ain, R., Kamboh, A. M., Butt, H. T., Shafait, S., Alamgir, W., Stricker, D., and Shafait, F. The NMT Scalp EEG Dataset: An Open-Source Annotated Dataset of Healthy and Pathological EEG Recordings for Predictive Modeling. Frontiers in Neuroscience, Volume 15, Article 755817, 2022. DOI: 10.3389/fnins.2021.755817 The dataset size, label distribution and demographic composition are reported in the paper. The collection, consent, anonymisation and neurologist agreement process are described in the data collection protocol. The technical format and directory structure are described in the dataset section.

提供机构:
Zenodo
创建时间:
2026-07-27
二维码
社区交流群
二维码
科研交流群
商业服务