Dataset: A multi-label classifier for predicting the most appropriate instrumental method for the analysis of contaminants of emerging concern
收藏资源简介:
NORMAN Suspect List Exchange was used for the generation of the dataset. Datasets with clear label (LC or GC) were used. More specifically, we used S3 NORMANCT15, which contains a list of compounds that were detected in surface water from the Danube River in a pan-European collaborative trial employing both GC-HRMS and LC-HRMS. Moreover, the GC and LC target list were used by the following two institutes: National and Kapodistrian University of Athens (NKUA) and Helmholtz Centre for Environmental Research (UFZ). S21 UATHTARGETS is the LC target list of NKUA, S65 UATHTARGETSGC is the GC target list of NKUA and S53 UFZWANATARG contains the LC and GC target list of UFZ. Finally, two GC target lists (S51 WRIGCHRMS and S70 EISUSGCEIMS) were used. These lists contain GC substance lists and were provided by two Slovak institutes, the Water Research Institute (WRI) and Environmental Institute. The aforementioned compound lists were merged together to form a labelled dataset. The SMILES were used to calculate 1446 molecular descriptors. 1446 descriptors were produced by PaDEL-descriptor, logP was produced by JRgui and boiling point by USEPA ECOSAR. The dataset is used in the publication: "A multi-label classifier for predicting the most appropriate instrumental method for the analysis of contaminants of emerging concern" authored by Nikiforos Alygizakis, Vasileios Konstantakos, Grigoris Bouziotopoulos , Evangelos Kormentzas, Jaroslav Slobodnik and Nikolaos S. Thomaidis Github repository: https://github.com/nalygizakis/LCvsGC
本数据集的构建采用了NORMAN可疑名单交换平台(NORMAN Suspect List Exchange)。本次研究采用了带有明确分类标签(液相色谱(Liquid Chromatography, LC)/气相色谱(Gas Chromatography, GC))的数据集。具体而言,本研究使用了S3 NORMANCT15数据集,该数据集收录了泛欧协作试验中于多瑙河地表水中检出的化合物名单,该项试验同时采用了气相色谱-高分辨质谱(GC-HRMS)与液相色谱-高分辨质谱(LC-HRMS)两种分析技术。此外,下述两家机构分别提供了对应的GC与LC靶向化合物名单:雅典国立与卡波迪斯特里亚大学(National and Kapodistrian University of Athens, NKUA)以及亥姆霍兹环境研究中心(Helmholtz Centre for Environmental Research, UFZ)。其中S21 UATHTARGETS为NKUA的LC靶向化合物名单,S65 UATHTARGETSGC为NKUA的GC靶向化合物名单,而S53 UFZWANATARG则收录了UFZ的LC与GC靶向化合物名单。最后,本次研究还使用了两份GC靶向化合物名单:S51 WRIGCHRMS与S70 EISUSGCEIMS。这两份名单由斯洛伐克的两家机构——水研究所(Water Research Institute, WRI)与环境研究所提供,其收录了各类GC靶向化合物。上述所有化合物名单经合并整合后,构建得到带标签的数据集。本研究通过化合物的简化分子线性输入规范(Simplified Molecular Input Line Entry System, SMILES)计算得到1446项分子描述符,其中1446项描述符由PaDEL-descriptor工具生成,正辛醇-水分配系数(logP)由JRgui工具计算得到,沸点数据则由美国环境保护署ECOSAR(USEPA ECOSAR)模块生成。本数据集已应用于学术论文《用于预测新兴污染物分析最优仪器方法的多标签分类器》("A multi-label classifier for predicting the most appropriate instrumental method for the analysis of contaminants of emerging concern"),作者为Nikiforos Alygizakis、Vasileios Konstantakos、Grigoris Bouziotopoulos、Evangelos Kormentzas、Jaroslav Slobodnik与Nikolaos S. Thomaidis,对应的GitHub代码仓库地址为:https://github.com/nalygizakis/LCvsGC



