Literature-derived allelopathy dataset used for machine-learning prediction of Lantana camara–wheat responses
收藏资源简介:
This dataset contains the literature-derived and computationally processed data used to develop and evaluate machine-learning models for predicting concentration-dependent plant growth responses to allelopathic plant extracts. The dataset was constructed through an automated literature-mining pipeline and contains 281 harmonised records spanning 11 donor plant taxa and 14 target plant taxa. The records include experimental information relating to extract concentration, plant species, germination and seedling growth responses. Approximately 54.4% of the retained records originated from Lantana camara studies. The dataset was subsequently expanded through chemoinformatic feature engineering. Major phytochemical compounds were resolved using PubChem and processed with RDKit to generate physicochemical descriptors and molecular fingerprints. Additional engineered variables include phytochemical class ratios, non-linear concentration transformations, concentration–chemical-property interactions and an osmotic-potential proxy. The resulting model feature matrix contains 73 variables. These data were used to train and evaluate Random Forest and XGBoost regression models for germination, root-length and shoot-length responses. The models were subsequently applied to predict the response of an experimentally untested Lantana camara–Triticum aestivum (wheat) interaction. The dataset is provided to support transparency, reproducibility and reuse of the computational analyses described in the associated research article. The repository contains the processed data required to reproduce the reported modelling workflow; accompanying code and analysis scripts are maintained in the associated GitHub repository. Data provenance: The dataset was derived from published scientific literature. Individual source publications are identified within the dataset where applicable. The dataset contains harmonised experimental observations and computationally derived features rather than copies of the source publications. File formats: CSV files containing tabular literature-derived observations, processed variables and model features. Standard spreadsheet or programming software such as Python, R or Microsoft Excel can be used to inspect the data.



