遇见数据集

Benchmarking AI Workflows for Hit Detection in High-Content Screening

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

The data set contains Python scripts, raw and processed data to benchmark hit detection in high-content screening. The raw data contains the extracted numerical features from high content images of 641 validated, highly selective pharmaceutically relevant inhibitors for over 123 targets. There were 979 features extracted using Perkin Elmer software Columbus (2.9.1532). In addition, two further raw data sets are available that were used to validate the machine learning models. These contain staurosporines in different concentrations or compounds that effect the cell cycle. The processed data were cleaned using an outlier test and Z-score transformed. The processing of the data is also available as a Python script and can be followed there. The other Python scripts contain the benchmarking of AI workflows for hit detection. Various machine learning models were tested, ranging from classic classifiers from the Scikit-Learn library over a partial least square regression to simple neural networks. To deal with unknown patterns in the hit detection, novelty detection was also tested. In addition, the used packages or environments with which the benchmarking was carried out are included.

本数据集包含用于高内涵筛选(high-content screening)命中检测基准测试的Python脚本、原始数据与处理后数据。 原始数据包含从641个经过验证、具有高度选择性且与药学相关的抑制剂的高内涵筛选图像中提取的数值特征,这些抑制剂覆盖超过123个靶点。本次共提取979项特征,使用的工具为珀金埃尔默(Perkin Elmer)Columbus软件(版本2.9.1532)。此外,另有两份原始数据集可用于验证机器学习模型。 这两份数据集包含不同浓度的星形孢菌素(staurosporines)或影响细胞周期的化合物。 处理后的数据通过异常值检验完成清洗,并经过Z分数(Z-score)标准化转换。数据处理的完整流程已封装为Python脚本,可通过该脚本复现全部处理步骤。 其余Python脚本用于实现命中检测的AI工作流基准测试。本次测试涵盖多种机器学习模型,包括来自scikit-learn库(Scikit-Learn)的经典分类器、偏最小二乘回归模型,以及简单神经网络。为应对命中检测任务中的未知模式,研究同时测试了新颖性检测方法。此外,本数据集还包含本次基准测试所使用的依赖包与运行环境信息。

创建时间:
2022-09-15
二维码
社区交流群
二维码
科研交流群
商业服务