Taxa-resolved particle size distributions derived from an IFCB (4-100 micrometers) at Station ALOHA (2017-2025; HOT294 - HOT361)
收藏资源简介:
We provide surface (~7m depth) particle size distribution data collected with an Imaging FlowCytobot (IFCB; McLane, East Falmouth, MA, USA) as part of the Hawaii Ocean Time Series program (HOT; https://hahana.soest.hawaii.edu/hot/) covering the 2017-2025 period (HOT294 - HOT 361). In brief, we provide cruise-averaged particle size distributions (both particle volume [microliters per liter per micron] and counts [number of particles per liter per micron]) over 50 log-spaced diameter bins ranging between 3-100 micrometers. We also provide particle diameters (estimated from cross-sectional areas), as well as major and minor axis lengths for all individual particles imaged by the IFCB during the above mentioned cruises, along with relevant metadata (time, latitude, longitude of collection of each particle). Detailed information is below: Sampling: Particles were sensed using an IFCB mounted in underway mode aboard the R/V Kilo Moana to collect high-resolution (~3.23 pixels per micron) particle images. The IFCB is a flow cytometer equipped with a high-resolution camera triggered by both side-angle light scattering and chlorophyll a fluorescence (680 ± 30 nm), enabling automated imaging of individual particles ~4-100 µm in size (Olson and Sosik, 2007), with the lower limit of detection dependent on instrument gain settings and optical properties of particles. A pre-filter mesh (150 μm) was installed at the intake. IFCB images were collected during 44 cruises between June 2017 and December 2025 at Station ALOHA (A Long-Term Oligotrophic Habitat Assessment; 22.75°N, 158°W), as part of the Hawaiian Ocean Time-series in the North Pacific Subtropical Gyre (all imagery available here: https://ifcbdb2.soest.hawaii.edu/dashboard). The distribution of cruises per year in the 2017-2025 time periods is variable annually, and presented as metadata in the PSD csv files. On an average cruise, ~4 mL of water samples from ~7m depth are run through the instrument every ~20 minutes. To minimize spatial variability, analysis was restricted to samples collected between 22.6-22.9°N and 157.9 - 158.2°W, around the center of Station ALOHA. Sample times were roughly evenly distributed by time of day (at least 212 samples per each hour bin). Each sample contained hundreds to thousands of imaged particles. Samples containing errors (e.g., <2 mL volume, poor camera alignment, blockages, bubbles, etc.) were manually assessed and removed as necessary. Data processing: IFCB images were processed using the IFCB-analysis MATLAB package (github.com/hsosik/ifcb-analysis), which segments and extracts feature statistics, including particle diameter, major axis lengths, perimeter, cross-sectional area, and others (Moberg and Sosik, 2012). Biovolume estimates used herein were calculated using each particle’s two-dimensional cross-sectional area to increase accuracy of plankton with complex shapes. The diameter used for image depth was approximated by using the diameter of a circle with equivalent area to the observed particle silhouette (i.e., area based diameter). For small, near-spherical particles (e.g., Crocosphaera-like, small coccolithophores), area-based methodologies are equivalent to those of diameter-derived estimates (e.g., equivalent spherical diameter, ESD). However, for elongated or irregularly shaped particles, ESD may underestimate particle size due to irregularities in segmented shape (e.g., chain-forming taxonomic categories like pennate diatoms or Hemiaulus) or overestimate particle size by filling in convex silhouettes (e.g., Ceratium, Radiolarians). In contrast, area-based diameter better captures the visual extent of such particles in the focal plane, offering a more robust and consistent metric and producing smoother, more continuous PSDs. For non-spherical taxa, we also examined major axis length distributions to better characterize shape variability, but chose to report PSDs using area-equivalent diameters to maintain a unified framework across taxa and support downstream modeling applications that require normalization across groups. Minor and Major axis data are presented on the per-particle file. Classification of imagery: A convolutional neural network (CNN, Gonzalez et al., 2019, Orestein et al., 2015) was applied to the IFCB dataset to classify images into one of 104 taxonomic categories. Classification accuracy was assessed using F1 scores (F1 = 2 * recall-1 * precision-1) using validation images that were roughly evenly distributed amongst cruises and were not included within any part of the training process (see Table_CNN_F1scores file provided). For clarity and statistical robustness, categories were consolidated into 14 major particle groups for the primary analyses. These are mostly at the phylum-level, and additional subdivisions are presented in the per particle taxa file. For example, phylum Bacillariophyta is further divided into 6 sub-categories; Haptophyta into 10 sub-categories; Cyanobacteria into 7 sub-categories; Dinoflagellate into 4 sub-categories. A specific super-category called “Unidentifiable” was designated for near-spherical particles that were too small to be confidently assigned a phylum label. Images for which the CNN is not able to identify a specific category with a probability of at least 50% across multiple iterations of the model is assigned an “Unknown” label (the 14th major category). The overall performance of the CNN performed was high, with mean F1 scores of 0.929 for major taxonomic categories and 0.852 for minor categories based on validation HOT images. Construction of PSDs: Size distributions (major taxonomic category) were binned into 50 logarithmically spaced bins centered between 3 and 98 µm. These limits were imposed because particles < 3 µm are rarely detected by the instrument, whereas particles > 100 µm in diameter were rare. Cruise-to-cruise variability in IFCB settings affected the detection limit, particularly the smallest particles that were able to be imaged by the instrument. The estimated detection limit for each cruise was calculated by selecting the bin diameter below which average cruise particle abundances (in units of particles per liter per micron) started to decline. This value was ~ 7 micron in size, so data below this threshold should not be over-emphasized. Cruise-averaged volume concentrations (microliter per liter per micron) and particle size distributions (particles per liter per micron), both provided as csv fles above, were used to minimize hour-to-hour and day-to-day variability and cruise-to-cruise variations due to changes in optimal detection limit. Integrated volume or abundances for each taxonomic group per cruise can be calculated by summing particle volume or abundances across all size bins within each cruise, providing a measure of total particle volume or numbers within the observable size range. References: González, P. et al. Automatic plankton quantification using deep features. J. Plankton Res. 41, 449–463 (2019). Moberg, E. A. & Sosik, H. M. Distance maps to estimate cell volume from two-dimensional plankton images. Limnol. Oceanogr. Methods 10, 278–288 (2012). Olson, R. J. & Heidi M. Sosik, H. M. A Submersible Imaging-in-Flow Instrument to Analyze Nano- and Microplankton: Imaging FlowCytobot. Limnol. Oceanogr.: Methods 5, 2007, 195–203. (2007). Orenstein, E., Beijbom, O., Peacock, E. & Sosik, H. WHOI-Plankton- A Large Scale Fine Grained Visual Recognition Benchmark Dataset for Plankton Classification. https://doi.org/10.48550/arXiv.1510.00745 (2015) doi:10.48550/arXiv.1510.00745. 27.



