NEWT multimodal embedding datasets and input data
收藏资源简介:
This dataset provides curated and preprocessed inputs for training and evaluating the NEWT (Neural Embeddings for Wide-spectrum Targeting) framework, a multimodal embedding approach for compound–target prediction and single-cell analysis. It includes gene embedding resources derived from multiple biological knowledge sources and integrated multimodal representations, along with processed L1000 perturbation signatures and compound–target mapping files used for model training, validation, and benchmarking. All files are formatted for direct compatibility with the NEWT pipeline and associated scripts, enabling reproducible generation of compound–target predictions, embedding fusion models, and downstream analyses. This dataset accompanies the manuscript: Kidder, B.L. et al. Multimodal gene embeddings enable prediction of drug targets and reconstruction of cellular states. bioRxiv (2026).



