Fine Identification of Noxious Weeds Based on Close-Range Hyperspectral Imaging and Spectral Features of Alpine Meadow Plants
收藏资源简介:
This dataset provides a comprehensive collection of hyperspectral data subsets, spectral types, feature descriptors, Mahalanobis distance selected bands, and shapefiles derived from close-range hyperspectral imagery, designed for fine-grained identification of alpine meadow vegetation, including noxious weeds. It supports ecological monitoring, remote sensing research, and machine learning applications in vegetation classification. 1. Datasets: Raw hyperspectral imagery: Uncropped original close-range hyperspectral data containing complete spectral information, suitable for subsequent processing and analysis such as hyperspectral data cropping and reflectance correction. Dataset-A (Reflectance Band1–Band367, 367 bands): All original reflectance bands from the close-range hyperspectral imagery. Dataset-B (R, 6 bands): Reduced reflectance bands selected using Mahalanobis distance to optimize spectral discrimination. Dataset-C (R + F + S + C, 81 bands): Combinations of spectral feature variables extracted from reflectance data. Dataset-D (R + F + S + C + GLCM + HOG, 92 bands): Spectral feature variables combined with texture features (GLCM) and shape features (HOG) for improved classification performance. Dataset-E (R + F + S + C + SI, 91 bands): All spectral feature combinations along with spectral indices (SI) for additional vegetation characteristics. Dataset-F (R + F + S + C + GLCM + HOG + SI, 102 bands): Integrates all spectral feature variables, texture features, shape features, and spectral indices for the most comprehensive feature set. 2. Spectrum Types: Original spectral: Raw reflectance spectrum representing the original spectral signature of each target. First-order derivative spectral: Enhances subtle spectral differences for better feature discrimination. Second-order derivative spectral: Highlights fine spectral variations to distinguish similar vegetation types. Continuum removal spectral: Normalized to emphasize absorption features for detailed spectral analysis. 3. Feature Descriptors: HOG (Histogram of Oriented Gradients): Captures local gradient orientation distributions to describe object shapes and structural patterns. GLCM (Gray Level Co-occurrence Matrix): Texture features that can be extracted using ENVI or MATLAB software. Mahalanobis distance: A statistical method used to select optimal spectral bands by measuring the distance between a point and a distribution, improving spectral discrimination and reducing redundancy. 4. Mahalanobis Distance Selected Bands: MD-select-band-R: Spectral bands selected from the original reflectance spectrum (R) using Mahalanobis distance. MD-select-band-F: Spectral bands selected from first-order derivative spectral data (F) using Mahalanobis distance. MD-select-band-S: Spectral bands selected from second-order derivative spectral data (S) using Mahalanobis distance. MD-select-band-C: Spectral bands selected from continuum removal spectral data (C) using Mahalanobis distance. These selected bands provide optimized spectral inputs for improving vegetation classification performance and reducing redundancy. 5. Shapefiles: train.shp: Contains training sample polygons for supervised classification. Each polygon represents a labeled area corresponding to a specific vegetation class, including noxious weeds, providing ground truth for model training. test.shp: Contains validation (testing) sample polygons used to evaluate classification models. Each polygon corresponds to a labeled area for quantitative assessment of model accuracy and generalization. 6. Machine Learning Models: The following machine learning models can be applied to this dataset and are accessible via EnMapBox at https://www.enmap.org/?q=enmapbox.432: LinearSVC (LSVC) Random Forest (RF) LightGBM (LGBM) Logistic Regression (LR) These datasets, feature descriptors, and shapefiles provide spatially explicit ground truth and comprehensive feature sets, enabling detailed vegetation classification and machine learning research in heterogeneous alpine grasslands. 7. Classification Results on Dataset-F: LSVC-Dataset-F-classification: Highest classification accuracy achieved using LinearSVC on Dataset-F. RF-Dataset-F-classification: Highest classification accuracy achieved using Random Forest on Dataset-F. LGBM-Dataset-F-classification: Highest classification accuracy achieved using LightGBM on Dataset-F. LR-Dataset-F-classification: Highest classification accuracy achieved using Logistic Regression on Dataset-F. These results provide benchmark performance metrics for different machine learning models on the most comprehensive feature set, supporting reproducibility and further model comparison.



