Iris flower data set
收藏资源简介:
Iris花数据集,也称为Fisher的Iris数据集,是一个多元数据集,由英国统计学家和生物学家Ronald Fisher于1936年提出,用于解决分类问题。该数据集包含三种Iris花(Iris setosa, Iris virginica和Iris versicolor)的50个样本,每个样本测量了四个特征:萼片和花瓣的长度和宽度,单位为厘米。
The Iris flower dataset, also known as Fisher's Iris dataset, is a multivariate dataset introduced by the British statistician and biologist Ronald Fisher in 1936 for the purpose of solving classification problems. This dataset comprises 50 samples from each of three species of Iris flowers (Iris setosa, Iris virginica, and Iris versicolor). Each sample is characterized by four features: the length and width of the sepals and petals, measured in centimeters.
Iris-Dataset 概述
数据集描述
Iris 花数据集,又称 Fishers Iris 数据集,是由英国统计学家和生物学家 Ronald Fisher 于 1936 年提出的多变量数据集。该数据集用于量化三种相关鸢尾花(Iris setosa, Iris virginica 和 Iris versicolor)的形态变异。数据集包含每种花各 50 个样本,每个样本测量了四个特征:萼片和花瓣的长度及宽度,单位为厘米。
数据集用途
该数据集基于 Fisher 的线性判别模型,已成为机器学习中许多统计分类技术(如支持向量机)的典型测试案例。尽管在聚类分析中不常见,但通过非线性主成分分析的非监督过程,三种鸢尾花种类是可以区分的。
数据集特点
- 包含三种鸢尾花种类的 150 个样本。
- 每个样本具有四个特征:萼片和花瓣的长度及宽度。
- 数据集用于展示监督和非监督技术在数据挖掘中的差异。
数据集应用
- 作为机器学习分类算法的测试案例。
- 用于解释和区分监督与非监督数据挖掘技术。
数据集参考文献
- R. A. Fisher (1936). "The use of multiple measurements in taxonomic problems". Annals of Eugenics.
- Edgar Anderson (1936). "The species problem in Iris". Annals of the Missouri Botanical Garden.
- A. N. Gorban, A. Zinovyev. Principal manifolds and graphs in practice: from molecular biology to dynamical systems, International Journal of Neural Systems.




