FITTING Data Mining Settings for Ranking Seed Lots
收藏资源简介:
ABSTRACT To enhance speed and agility in interpreting physiological quality tests of seeds, The use of algorithms has emerged. This study aimed to identify suitable machine learning models to assist in the precise management of seed lot quality. Soybean lots from two companies were assessed using the Supplied Test Set, Cross-Validation (with 8, 10, and 12 folds), and Percentage Split (with 66% and 70%) methods. Variables analyzed through Tetrazolium tests included vigor, viability, mechanical damage, moisture damage, bed bug damage, and water content. Method performance was determined by Kappa, Precision, and ROC Area metrics. Classification Via Regression and J48 algorithms were employed. The technique utilizing 66% of data for training achieved 93.55% accuracy, with Precision and ROC Area reaching 94.50% for the J48 algorithm. Applying the cross-validation method with 10 folds resulted in 90.22% of correctly classified instances, with a ROC Area outcome like the previous method. Tetrazolium Vigor was the primary attribute used. However, these results are specific to this study's database, and careful planning is necessary to select the most effective application methods.
摘要:为提升种子生理品质检测解读的速度与灵活性,算法应用应运而生。本研究旨在筛选适配的机器学习模型,以辅助精准管控种子批次品质。研究采用提供的测试集、交叉验证(Cross-Validation,8折、10折及12折)与百分比划分法(Percentage Split,66%与70%划分比例),对两家企业的大豆批次开展评估。通过四唑试验(Tetrazolium test)分析的变量包括种子活力、种子生活力、机械损伤、湿损、臭虫侵害损伤及含水率。模型性能通过Kappa系数、精确率(Precision)与ROC曲线下面积(ROC Area)三类指标进行评估。本研究采用回归分类法(Classification Via Regression)与J48两种算法。其中,采用66%数据作为训练集的方案准确率达93.55%,J48算法的精确率与ROC曲线下面积均达到94.50%。采用10折交叉验证的方案,分类正确实例占比达90.22%,其ROC曲线下面积结果与前述方案相当。四唑活力指标为本次实验的核心特征变量。但本研究结果仅适用于本次实验的数据集,实际应用中需谨慎规划以筛选最优的实施方法。



