RiemannMol I: Molecular Generation with Learned Latent Metrics
收藏资源简介:
Figure4/ — coupling_vs_old_head.py, pairs.csv: 4,994 random test-set molecule pairs with their Tanimoto similarity and the corresponding raw-z, retired nonlinear-head, and coupling-head projected distances (Pearson r = −0.252 / −0.729 / −0.775). Figure5/ — umap_zprime_substructure.py, molecules.csv: 2,000 test-set molecules with their SMILES, substructure category (triazole / ester / thiazole-thiophene / other), k-means cluster assignment, and UMAP coordinates in raw z and the coupling-projected space z′. Figure6/ and Figure7/ — logp_metric_sampling.py (identical copy in both folders; one training run produces both figures), umap_points.csv and sampled_candidates.csv: the isometry-regularized LogP coupling head's training-set UMAP embedding and the 60 candidates decoded from noise sampled around the top-50 highest-LogP centroid (54/60 valid, mean LogP = 6.23–6.24). Figure8/ and Figure9/ — potency_metric_sampling.py (identical copy in both folders; one training run produces both figures), umap_points.csv and sampled_candidates.csv: the isometry-regularized EGFR-potency coupling head's training-set UMAP embedding and the 60 candidates decoded from noise sampled around the top-50 highest-pIC50 centroid (93% valid, mean predicted pIC50 = 6.09). FigureS1/ (Supporting Information) — raw_z_similarity_merged.py, global_pairs.csv and local_noise_sweep.csv: 4,994 random molecule pairs (global) and a 10-scale Gaussian noise sweep around one seed molecule (local), each with raw-z distance and Tanimoto similarity.



