Revisiting aquatic biodiversity–carbon burial relationships in freshwater lakes: A state-dependent perspective
收藏资源简介:
his dataset includes 12 CSV files, corresponding to Figures 2–5 and Supplementary Figures S3_S4_S5_S8_S9_S11_S12_S13 in the associated manuscript. Each file is described below with its structure, column definitions, units, and organization. Missing data are represented by blank cells (empty strings between commas). 1. Fig.2.csv – Data for Figure 2 File purpose: Geographic coordinates and identification information of 73 study lakes used for regional synthesis analysis. This dataset provides the spatial distribution of sampling sites across China, linking each lake to its corresponding identification codes for mapping and spatial analysis. The dataset also supports the quantification of dataset overlap relationships among algal diversity data, submerged macrophyte coverage data, and organic carbon burial rate data. File structure: CSV (comma‑separated values). 74 rows (header + 73 data rows). 12 columns. UTF‑8 with BOM encoding. Missing values are indicated by blank cells (no NA). Column definitions: Column 1 is Lake ID1, the primary lake identifier in numeric code. Column 2 is NAME1, the primary lake name. Column 3 is LON1, the longitude of the primary lake in decimal degrees. Column 4 is LAT1, the latitude of the primary lake in decimal degrees. Column 5 is Lake ID2, a secondary lake identifier in numeric code. Column 6 is NAME2, the secondary lake name. Column 7 is LON2, the longitude of the secondary lake in decimal degrees. Column 8 is LAT2, the latitude of the secondary lake in decimal degrees. Column 9 is Lake ID3, a tertiary lake identifier in numeric code. Column 10 is NAME3, the tertiary lake name. Column 11 is LON3, the longitude of the tertiary lake in decimal degrees. Column 12 is LAT3, the latitude of the tertiary lake in decimal degrees. Data organization: Each row represents one study lake, containing up to three identification codes and their corresponding coordinates. Some lakes have only one or two IDs depending on data sources and historical nomenclature. Different lakes are grouped by three datasets: Group 1 includes lakes with sedimentary algal diversity data (Lake ID1/NAME1/LON1/LAT1; n = 19), Group 2 includes lakes with submerged macrophyte coverage data derived from remote sensing (Lake ID2/NAME2/LON2/LAT2; n = 54), and Group 3 includes lakes with total organic carbon data for organic carbon burial rate (Lake ID3/NAME3/LON3/LAT3; n = 55). Rows are ordered by Lake ID1, with blank cells in ID1 indicating lakes not included in that specific dataset. The file includes 73 unique lakes in total, with the following overlaps among the three datasets: 2 lakes overlap between the submerged macrophyte coverage data and the algal data; 12 lakes overlap between the algal data and the organic carbon burial rate data; and 6 lakes overlap between the submerged macrophyte coverage data and the organic carbon burial rate data. Correspondence to manuscript figures: This file is used to generate Figure 2, the distribution map of study sites. Lake ID1/NAME1/LON1/LAT1 are used to plot the green circles representing lakes with sedimentary algal diversity data (n = 19). Lake ID2/NAME2/LON2/LAT2 are used to plot the blue hexagons representing lakes with submerged macrophyte coverage data from remote sensing (n = 54). Lake ID3/NAME3/LON3/LAT3 are used to plot the purple triangles representing lakes with total organic carbon data for organic carbon burial rate (n = 55). 2. Fig.3.csv – Data for Figure 3 File purpose: Time‑series data of aquatic biodiversity (algal diversity and macrophyte diversity) and organic carbon burial rate (OCBR) in Lake Liangzi over the past 120 years based on sedimentary ancient DNA (sedaDNA) analysis. The dataset supports the analysis of nonlinear relationships between biodiversity and carbon burial using Generalized Additive Models (GAM), as well as stagewise changes across three ecological stages. File structure: CSV (comma‑separated values). 27 rows (header + 26 data rows). 8 columns. UTF‑8 with BOM encoding. Missing values are indicated by blank cells (no NA). Column definitions: Column 1 is Year1, the age for algal diversity data in year AD. Column 2 is Richness_alage, the species richness of algae (unitless). Column 3 is Year2, the age for OCBR data corresponding to algal diversity in year AD. Column 4 is LZHZ_OCBR1, the organic carbon burial rate in g C m⁻² yr⁻¹ associated with algal diversity. Column 5 is Year3, the age for macrophyte diversity data in year AD. Column 6 is Richness_macrophye, the species richness of aquatic macrophytes (unitless). Column 7 is Year4, the age for OCBR data corresponding to macrophyte diversity in year AD. Column 8 is LZHZ_OCBR2, the organic carbon burial rate in g C m⁻² yr⁻¹ associated with macrophyte diversity. Data organization: Each row represents a time point with paired biodiversity and OCBR measurements. Different proxies may have slightly different age resolutions; blanks indicate no data for that variable at that time point. The time series spans approximately 1900 to 2021, with rows ordered from youngest (top) to oldest (bottom). The data are organized to support GAM fitting of biodiversity–OCBR relationships, with algal and macrophyte diversity paired against their respective OCBR estimates at comparable ages. Correspondence to manuscript figures: This file is used to generate Figure 3. Richness_alage and LZHZ_OCBR1 are used for panel (a), showing the nonlinear relationship between algal diversity and OCBR fitted by GAM. Richness_macrophye and LZHZ_OCBR2 are used for panel (b), showing the nonlinear relationship between macrophyte diversity and OCBR fitted by GAM. The time‑series data of algal richness (Year1 vs. Richness_alage) and macrophyte richness (Year3 vs. Richness_macrophye) are used for panels (c) and (d), respectively, showing stagewise changes in biodiversity. LZHZ_OCBR1 and LZHZ_OCBR2 (Year2 vs. OCBR) are used for panel (e), showing stagewise changes in organic carbon burial. The vertical dashed lines at ca. 1960 and 2010 in the figure are based on the temporal divisions inferred from the sedimentary record and are used to demarcate the three ecological stages. 3. Fig.4.csv – Data for Figure 4 File purpose: Regional‑scale time‑series data of lake aquatic biodiversity (algal diversity and submerged aquatic vegetation coverage) and organic carbon burial rate (OCBR) across Chinese lakes. The dataset compiles three independent time‑series records: (1) 19 algal diversity records (1900–2016) inferred from sedimentary DNA; (2) 54 aquatic vegetation diversity records (1989–2021) derived from remote sensing; and (3) 55 organic carbon burial rate records (1900–2016) calculated from sediment core data. This dataset supports the analysis of regional‑scale biodiversity–carbon burial relationships and their temporal dynamics. File structure: CSV (comma‑separated values). 127 rows (header + 126 data rows). 207 columns. UTF‑8 with BOM encoding. Missing values are indicated by blank cells (no NA). Column definitions: The file is organized into three sequential data blocks, each with a year column followed by multiple lake‑specific columns, plus average columns summarizing central tendencies. Block 1 (algal diversity, 1900–2016): Column 1 is Year1, the age for algal diversity data in year AD. Columns 2–20 (Lake_Hongjiannao1 through Average1) contain algal richness values (unitless) for 19 lakes. The final column of this block, Average1, is the mean algal richness across all 19 lakes. Block 2 (aquatic vegetation coverage, 1989–2021): Column 21 is Year2, the age for vegetation coverage data in year AD. Columns 22–75 (Hei_Citan through Average2) contain submerged aquatic vegetation coverage values (percentage, %). The final column of this block, Average2, is the mean coverage across all 54 lakes. Block 3 (organic carbon burial rate, 1900–2016): Column 76 is Year3, the age for OCBR data in year AD. Columns 77–131 (Lake_Hongjiannao through Average3) contain OCBR values (g C m⁻² yr⁻¹) for 55 lakes. The final column of this block, Average3, is the mean OCBR across all 55 lakes. Data organization: Each row represents a time point with up to three sets of paired observations (year + values). Different variables have different temporal resolutions and coverage periods: algal diversity spans ~1900–2016 with irregular temporal spacing; aquatic vegetation coverage spans ~1989–2021 with annual to sub‑annual resolution; OCBR spans ~1900–2016 with irregular spacing. Blanks indicate no data for that lake at that time point. Each block concludes with an average column summarizing the central tendency across all lakes in that block. Correspondence to manuscript figures: This file is used to generate Figure 4. The algal diversity time‑series (Block 1: Year1 and algal richness columns) are used for panel (c), showing 19 algal richness records, and panel (f), showing stagewise changes at 30‑year intervals. The aquatic vegetation coverage time‑series (Block 2: Year2 and coverage columns) are used for panel (d), showing 54 coverage records, and panel (g), showing stagewise changes at 30‑year intervals. The OCBR time‑series (Block 3: Year3 and OCBR columns) are used for panel (e), showing 55 OCBR records, and panel (h), showing stagewise changes at 30‑year intervals. Panel (a) is generated from the site distribution of OCBR records (lakes in Block 3) combined with Lake Liangzi data. Panel (b) is generated from the site distribution of algal diversity (Block 1) and submerged aquatic vegetation coverage (Block 2) records combined with Lake Liangzi data. 4. Fig.5.csv – Data for Figure 5 File purpose: Integrated time‑series dataset for correlation heatmap and Mantel test analyses examining the dynamic relationships among lake biodiversity (micro‑eukaryotes and aquatic macrophytes), organic carbon burial rate (OCBR), and multiple environmental drivers in Lake Liangzi across three distinct ecological phases: (1) macrophyte‑dominated state (pre‑1960s), (2) macrophyte–algae coexisting state (1960–2010s), and (3) algae‑dominated state (post‑2010s). The dataset includes biodiversity indices for five biotic groups (algae, fungi, invertebrates, protozoa, and macrophytes), OCBR, and environmental factors spanning nutrient levels (Chl‑a, TP), catchment soil erosion (magnetic susceptibility, χlf), climate (mean annual temperature, MAT), hydrology (average water level, AWL), and land‑use change (cultivated land area, CLA). File structure: CSV (comma‑separated values). 13 rows (header + 12 data rows). 81 columns. UTF‑8 with BOM encoding. Missing values are indicated by blank cells (no NA). Column definitions: The file is organized into six sequential temporal blocks corresponding to different sampling periods, followed by a consolidated summary block. Each block contains paired year columns with biodiversity, OCBR, and environmental data. Block 1 (Phase 1: pre‑1960s, ~1911–1959): Columns 1–12 contain early‑period data. Column 1 is Year1. Columns 2–6 contain diversity (richness) data for five biotic groups (Alage, Fungi, Invertebrates, Protozoa, Macrophytes). Columns 7–8 contain NMDS1 ordination scores for algae and macrophytes. Columns 9–10 contain OCBR and TOC data. Columns 11–12 contain environmental data (Chl‑a, TP, χlf, Temperature) for this period. Block 2 (Phase 2: 1960s–2010s, ~1961–2009): Columns 13–20 contain middle‑period data. Column 13 is Year2. Columns 14–18 contain the same five biodiversity richness variables. Columns 19–20 contain OCBR and TOC. Columns 21–24 contain environmental data (Chl‑a, TP, χlf, Temperature, Cultivated_land_area, Average_water_level). Block 3 (Phase 3: post‑2010s, ~2013–2021): Columns 25–36 contain recent‑period data. Column 25 is Year3. Columns 26–30 contain biodiversity richness variables. Columns 31–32 contain NMDS1 scores for algae and macrophytes. Columns 33–34 contain OCBR and TOC. Columns 35–36 contain environmental data. Block 4 (Phase 3 continued, post‑2010s): Columns 37–42 contain additional recent‑period data. Column 37 is Year4. Columns 38–43 contain environmental data (Chl‑a, TP, χlf, Temperature, Cultivated_land_area, Average_water_level) for this period. Block 5 (Phase 3 continued, post‑2010s): Columns 43–50 contain additional recent‑period data. Columns 43–44 contain biodiversity data (Alage_richness4, Mac_diversity4). Columns 45–50 contain OCBR, Chl‑a, TP, χlf, and Temperature data. Block 6 (Phase 3 continued, post‑2010s): Columns 51–55 contain additional recent‑period data. Columns 51–52 contain biodiversity data (Alage_richness5, Mac_diversity5). Columns 53–55 contain OCBR, Chl‑a, TP, χlf, Cultivated_land_area, Temperature, and Average_water_level data. Block 7 (Consolidated summary block): Columns 56–64 contain integrated data for all variables across time. Column 56 is Alage_richness. Column 57 is Mac_diversity. Column 58 is OCBR. Column 59 is Chl‑a. Column 60 is TP. Column 61 is χlf. Column 62 is Cultivated_land_area. Column 63 is Temperature. Column 64 is Average_water_level. Data organization: Each row represents a time point with measurements for biodiversity and environmental variables. Different variables have different temporal resolutions and coverage periods. The data are organized to support three‑phase comparative analyses, with columns grouped by phase. For Phases 1 and 2, data are relatively complete across all variables. For Phase 3, data are distributed across multiple column blocks (Year3 through Year5) due to differences in measurement timing. Blanks indicate no data for that variable at that time point. The final summary block (columns 56–64) provides a consolidated view of all core variables. Correspondence to manuscript figures: This file is used to generate Figure 5. The three temporal blocks (Phase 1: columns 1–12; Phase 2: columns 13–24; Phase 3: columns 25–55) provide the data for correlation heatmaps and Mantel tests for panels (a), (b), and (c), respectively. Biodiversity variables (richness and NMDS1 for algae, fungi, invertebrates, protozoa, and macrophytes) are used to compute correlation matrices against environmental factors (Chl‑a, TP, χlf, Temperature, Cultivated_land_area, Average_water_level) and OCBR for each phase. 5. Fig.S3.csv – Data for Figure S3 File purpose: Geochemical and physical proxy data from sediment core LZHZ‑23‑1 in Lake Liangzi, spanning approximately 1902–2022. This dataset provides multi‑proxy records for reconstructing long‑term environmental changes, including organic matter sources (TOC, TN, C/N), primary productivity (sedimentary and monitored Chl‑a), nutrient levels (sedimentary and monitored TP), catchment erosion indicators (median grain size, magnetic susceptibility), and redox conditions (Fe/Mn ratio). These proxies collectively support the reconstruction of lake trophic state evolution and catchment disturbance history. File structure: CSV (comma‑separated values). 27 rows (header + 26 data rows). 21 columns. UTF‑8 with BOM encoding. Missing values are indicated by blank cells (no NA). Column definitions: The file is organized into six sequential blocks based on measurement timing and proxy types, each with its own age column. Block 1 (Sedimentary geochemistry – early period): Column 1 is Year1, the age for TOC, TN, and Chla1 in year AD. Column 2 is TOC (%), total organic carbon content. Column 3 is TN (%), total nitrogen content. Column 4 is Chla1 (μg g⁻¹), sedimentary chlorophyll‑a concentration. Block 2 (Monitored water Chl‑a): Column 5 is Year2, the age for monitored water Chl‑a data in year AD. Column 6 is Chla (μg L⁻¹), chlorophyll‑a concentration from instrumental water quality monitoring. Block 3 (Sedimentary C/N and TP – early period): Column 7 is CN, the carbon‑to‑nitrogen ratio (unitless). Column 8 is TP1 (mg kg⁻¹), sedimentary total phosphorus concentration. Block 4 (Monitored water TP): Column 9 is Year3, the age for monitored water TP data in year AD. Column 10 is TP (μg L⁻¹), total phosphorus concentration from instrumental water quality monitoring. Block 5 (Physical sediment properties): Column 11 is Median_grain_size (μm), the median grain size of sediment particles. Column 12 is PC1_elements (unitless), the first principal component of elemental composition. Column 13 is Ms (10⁻⁸ m³ kg⁻¹), mass‑specific magnetic susceptibility. Column 14 is FeMn (unitless), the iron‑to‑manganese ratio (mass ratio, %). Data organization: Each row represents one sediment sample or water monitoring time point. Different proxies have different temporal resolutions and coverage periods: sedimentary proxies (Block 1 and Block 3) span ~1902–2021 with irregular temporal spacing; monitored water Chl‑a (Block 2) spans ~2005–2022 with near‑annual resolution; monitored water TP (Block 4) spans ~2005–2022; physical properties (Block 5) are paired with the Year1 age column and are available for most sedimentary samples. Blanks indicate no data for that proxy at that depth/time point. Rows are ordered from youngest (top, ~2022) to oldest (bottom, ~1902). Correspondence to manuscript figures: This file is used to generate Figure S3. TOC, TN, Chla1, CN, TP1, Median_grain_size, PC1_elements, Ms, and FeMn are plotted against Year1 to show long‑term sedimentary geochemical trends. Monitored water Chl‑a and TP are plotted against Year2 and Year3, respectively, to show instrumental water quality records. Combined, these profiles provide the multi‑proxy basis for reconstructing lake trophic evolution, catchment erosion history, and redox conditions over the past century. 6. Fig.S4.csv – Data for Figure S4 File purpose: Major and trace element concentrations from sediment core LZHZ‑23‑1 in Lake Liangzi, spanning approximately 1902–2022. This dataset provides high‑resolution geochemical records for reconstructing changes in sediment source, catchment weathering intensity, and anthropogenic pollution history. The elements include: Fe and Mn (redox‑sensitive elements), P (nutrient), Ti (lithogenic indicator), Cu, Zn, Cd, and Pb (trace metals indicative of human activities, such as industrial emissions, mining, and agricultural runoff). File structure: CSV (comma‑separated values). 27 rows (header + 26 data rows). 9 columns. UTF‑8 with BOM encoding. Missing values: none. Column definitions: Column 1 is Year, the age of the sediment sample in year AD. Column 2 is Fe (mg/kg or μg/g), the concentration of iron. Column 3 is Mn (mg/kg or μg/g), the concentration of manganese. Column 4 is P (mg/kg or μg/g), the concentration of total phosphorus. Column 5 is Ti (mg/kg or μg/g), the concentration of titanium. Column 6 is Cu (mg/kg or μg/g), the concentration of copper. Column 7 is Zn (mg/kg or μg/g), the concentration of zinc. Column 8 is Cd (mg/kg or μg/g), the concentration of cadmium. Column 9 is Pb (mg/kg or μg/g), the concentration of lead. Data organization: Each row represents one sediment sample from the LZHZ‑23‑1 core, with multi‑element concentrations. The core spans approximately 1902–2021, with varying temporal resolution. Rows are ordered from youngest (top, ~2021) to oldest (bottom, ~1902). The Year column is aligned with the main core chronology and corresponds to the same age model used in Fig. S3. The dataset provides a continuous record of sedimentary elemental concentrations, enabling the reconstruction of changes in sediment provenance, weathering intensity, and anthropogenic metal inputs over the past century. Correspondence to manuscript figures: This file is used to generate Figure S4. All element concentrations (Fe, Mn, P, Ti, Cu, Zn, Cd, Pb) are plotted against Year to show long‑term elemental trends in the sedimentary record. These profiles provide essential geochemical evidence for reconstructing: (1) natural versus anthropogenic metal sources (Cu, Zn, Cd, Pb), (2) catchment weathering and erosion intensity (Ti, Fe, Mn), and (3) nutrient enrichment history (P). The trace metal records (especially Pb, Cu, Zn, and Cd) serve as proxies for industrial development and catchment disturbance, complementing the physical and biological proxies (Fig. S3) to provide a comprehensive multi‑proxy reconstruction of lake environmental evolution. 7. Fig.S5.csv – Data for Figure S5 File purpose: Historical socioeconomic and environmental time‑series data from the Liangzi Lake Basin (1901–2022). This dataset compiles long‑term records of climate variables (precipitation and temperature), flood frequency, land‑use change (cultivated land area), population dynamics (total population of Ezhou City), agricultural intensification (fertilizer usage), and economic activity (fishery catches). These data support the analysis of anthropogenic drivers of lake environmental change, linking catchment‑scale human activities to sedimentary records of eutrophication, erosion, and ecological regime shifts in Lake Liangzi over the past century. File structure: CSV (comma‑separated values). 123 rows (header + 122 data rows). 21 columns. UTF‑8 with BOM encoding. Missing values are indicated by blank cells (no NA). Column definitions: The file is organized into six independent time‑series blocks, each with its own year column and corresponding data columns. Block 1 (Climate data – precipitation and temperature): Column 1 is Year1, the age for precipitation and temperature records in year AD. Column 2 is Precipitation (mm), annual total precipitation. Column 3 is Temperature (°C), mean annual temperature. Block 2 (Flood records): Column 4 is Year2, the age for flood grade data in year AD. Column 5 is Flood_grade (unitless, 0–5 scale), a historical flood severity classification (higher values indicate more severe flooding). Block 3 (Land use – cultivated land area): Column 6 is Year3, the age for cultivated land area data in year AD. Column 7 is Cultivated_land_area (unit not specified, likely 10³ ha or 10⁴ mu), the total cultivated land area in the basin. Block 4 (Demographic data – population): Column 8 is Year4, the age for population data in year AD. Column 9 is Total_population_of_Ezhou (unit not specified, likely 10⁴ persons), the total population of Ezhou City. Block 5 (Agricultural intensification – fertilizer usage): Column 10 is Year5, the age for fertilizer usage data in year AD. Column 11 is Fertilizer_usage (unit not specified, likely 10⁴ tons), total chemical fertilizer consumption. Block 6 (Economic activity – fishery catches): Column 12 is Year6, the age for fishery catch data in year AD. Column 13 is Catches_of_fishery_products (unit not specified, likely 10⁴ tons), total annual fishery catches from the lake. Data organization: Each row represents a time point for multiple socioeconomic and environmental variables. Different variables have different temporal resolutions and coverage periods: precipitation and temperature span ~1901–2022 with annual resolution; flood grades span ~1906–1950; cultivated land area spans ~1940–2021 with sub‑decadal resolution; population spans ~1920–2021 with variable resolution; fertilizer usage spans ~1953–2021; fishery catches span ~1950–2021. Blanks indicate no data for that variable at that time point. Rows are ordered by Year1 (climate data), with other variables aligned to their respective years in separate columns. This structure allows independent temporal analysis of each driver while maintaining compatibility with the sediment core chronology. Correspondence to manuscript figures: This file is used to generate Figure S5. Precipitation and temperature (Block 1) are used to reconstruct climate trends over the past century; flood grades (Block 2) indicate extreme hydrological events; cultivated land area (Block 3) and population (Block 4) track land‑use change and demographic pressure; fertilizer usage (Block 5) indicates agricultural intensification; and fishery catches (Block 6) reflect economic exploitation of the lake. These historical data are used to contextualize the sedimentary geochemical and biological records (Figs. S3–S4), enabling attribution of ecological regime shifts to specific anthropogenic drivers (e.g., land‑use change, nutrient loading, and hydrological alteration). 8. Fig.S8.csv – Data for Figure S8 File purpose: Multi‑proxy evidence supporting minimal DNA degradation in sediment samples from core LZHZ‑23‑1. This dataset integrates geochemical indicators (Fe/Mn ratio, TOC), pigment‑based primary productivity (Chl‑a), and sedimentary ancient DNA (sedaDNA) data for six biotic groups (aquatic macrophytes, algae, fungi, invertebrates, and protozoa). The combination of these proxies demonstrates that sedimentary DNA is well preserved and reliably reflects historical community composition, with no evidence of significant post‑depositional degradation that would bias ecological reconstructions. File structure: CSV (comma‑separated values). 27 rows (header + 26 data rows). 22 columns. UTF‑8 with BOM encoding. Missing values are indicated by blank cells (no NA). Column definitions: The file is organized into three sequential data blocks, each with its own year column followed by corresponding proxy data. Block 1 (Geochemical and pigment proxies): Column 1 is Year1, the age for Fe/Mn, TOC, and Chla data in year AD. Column 2 is FeMn, the iron‑to‑manganese ratio (unitless), indicating redox conditions. Column 3 is Year2, the age for TOC and Chla data (identical to Year1 in this dataset). Column 4 is TOC (%), total organic carbon content. Column 5 is Chla (μg g⁻¹), sedimentary chlorophyll‑a concentration. Block 2 (sedaDNA read counts and OTU data): Column 6 is Year3, the age for sedaDNA data in year AD. Columns 7–11 contain read count data for five biotic groups: Macrophyte_OTU (aquatic macrophyte OTU reads), Alage_OTU (algal OTU reads), Fungi_OTU (fungal OTU reads), Invertebrate_OTU (invertebrate OTU reads), and Protozoa_OTU (protozoan OTU reads). Column 12 is Eukaryotic_microorganisms_reads, the total sequencing reads for eukaryotic microorganisms. Block 3 (Richness indices for five biotic groups): Columns 13–17 contain species richness values (number of taxa) for five biotic groups: Richness_alage (algal richness), Richness_fungi (fungal richness), Richness_Protozoa (protozoan richness), and Richness_inverbrate (invertebrate richness). Data organization: Each row represents one sediment sample (depth/age) with paired geochemical, pigment, and sedaDNA data. The time series spans approximately 1902–2021, with rows ordered from youngest (top, ~2021) to oldest (bottom, ~1902). Different proxies have different temporal resolutions, with gaps in the sedaDNA record for some samples (notably the earliest samples). The Fe/Mn ratio provides a key line of evidence for stable redox conditions conducive to DNA preservation. TOC and Chl‑a help assess whether organic matter degradation has biased the sedimentary record. The consistency of OTU reads and richness across different periods (as shown in the figure panels) supports the interpretation that DNA degradation is minimal. Correspondence to manuscript figures: This file is used to generate Figure S8. Fe/Mn ratio (Column 2) is used for panel (a), showing stable redox conditions favorable for DNA preservation. TOC (Column 4) is used for panel (b), showing no significant shifts attributable to DNA degradation. Chl‑a (Column 5) is used for panel (c), supporting changes in community composition. Macrophyte_OTU (Column 7) is used for panel (d), showing consistency across different periods. Richness_alage (Column 13) is used for panel (e); Richness_fungi (Column 14) for panel (f); Richness_Protozoa (Column 15) for panel (g); and Richness_inverbrate (Column 16) for panel (h). Collectively, these panels provide multiple independent lines of evidence that sedimentary DNA in the LZHZ‑23‑1 core is well preserved and suitable for reconstructing historical biodiversity dynamics. 9. Fig.S9.csv – Data for Figure S9 File purpose: Time‑series relative abundance data for five biotic groups (eukaryotic algae, invertebrates, fungi, protozoa, and aquatic macrophytes) from sedimentary ancient DNA (sedaDNA) analysis of Lake Liangzi sediments. This dataset provides the taxonomic composition of each group across approximately 1902–2021, enabling the reconstruction of community dynamics through different ecological regimes. The data are used for CONISS (constrained incremental sums of squares) cluster analysis to identify distinct stratigraphic zones in community composition. File structure: CSV (comma‑separated values). 24 rows (header + 23 data rows). 71 columns. UTF‑8 with BOM encoding. Missing values are indicated by blank cells (no NA). Column definitions: The file is organized into five sequential data blocks, each with a year column followed by relative abundance columns for taxa within that biotic group. Block 1 (Eukaryotic algae, panel a): Column 1 is Year1 (1902–2021), followed by relative abundance data for 9 algal taxonomic groups: Compsopogonales, Diatomea, Dinoflagellata, Euglenozoa1, Pavlovophyceae, Phragmoplastophyta, Prymnesiophyceae, and Rhodellophyceae. Values are read counts (relative abundance). Block 2 (Invertebrates, panel b): Column 10 is Year2 (1902–2021), followed by relative abundance data for 9 invertebrate groups: Annelida, Arthropoda, Cnidaria, Gastrotricha, Mollusca, Nematozoa, Platyhelminthes, Porifera, and Rotifera. Block 3 (Fungi, panel c): Column 19 is Year3 (1902–2021), followed by relative abundance data for 13 fungal groups: Aphelidea, Ascomycota, Basidiomycota, Blastocladiomycota, Chytridiomycota, Cryptomycota, Hyphochytriomycetes, Labyrinthulomycetes, Microsporidia, Mucoromycota, Neocallimastigomycota, Peronosporomycetes, and Zoopagomycota. Block 4 (Protozoa, panel d): Column 32 is Year4 (1902–2021), followed by relative abundance data for 15 protozoan groups: Apicomplexa, Bicosoecida, Breviatea, Cavosteliida, Centrohelida, Dictyostelia, Euglenozoa2, Gracilipodida, Heterolobosea, Preaxostyla, Protosporangiida, Protosteliida, Retaria, Rigifilida, and Schizoplasmodiida. Block 5 (Aquatic macrophytes, panel e): Column 47 is Year5 (1902–2021), followed by relative abundance data for 20 macrophyte taxa: Ceratophyllum_demersum, Hydrilla_verticillata, Myriophyllum_spicatum, Nitellopsis_obtusa, Stuckenia_pectinata, Potamogeton_perfoliatus, Potamogeton_natans, Potamogeton_maackianus, Potamogeton_lucens, Potamogeton_gramineus, Potamogeton_distinctus, Potamogeton_crispus, Potamogeton_berchtoldii, Nelumbo_nucifera, Nymphaea_macrosperma, Phragmites_australis, Zizania_latifolia, Nymphoides_peltata, Pontederia_crassipes, and Trapa_natans. Data organization: Each row represents a time point (approximately 1902–2021) with relative abundance data for all five biotic groups. Different groups may have slightly different temporal resolutions; blanks indicate no data for that taxon at that time point. Rows are ordered from youngest (top, ~2021) to oldest (bottom, ~1902). For each group, relative abundances are expressed as read counts that sum to total reads for that sample. The CONISS cluster analysis (right panel of Figure S9) is performed on this dataset using constrained incremental sums of squares to identify statistically significant stratigraphic zones, with the CONISS dendrogram showing zones of similarity among samples (Zones 1–3 or more). CONISS results are typically generated using Tilia or R (vegan, rioja packages). Correspondence to manuscript figures: This file is used to generate Figure S9. Eukaryotic algae data (Block 1) are used for panel (a); invertebrate data (Block 2) for panel (b); fungi data (Block 3) for panel (c); protozoa data (Block 4) for panel (d); and aquatic macrophyte data (Block 5) for panel (e). The CONISS cluster analysis is performed on the combined dataset and is displayed on the right side of the figure, identifying major stratigraphic zones in community composition across time. 10. Fig.S11.csv – Data for Figure S11 File purpose: Time-series relative abundance data for four major microbial eukaryotic groups (eukaryotic algae, fungi, invertebrates, and protozoa) at the phylum/class level from sedimentary ancient DNA (sedaDNA) analysis of Lake Liangzi sediments. This dataset spans approximately 1902–2021 and provides the taxonomic composition of microbial eukaryotic communities at a broad phylogenetic level, enabling the reconstruction of long-term shifts in community structure at the phylum level through different ecological regimes. File structure: CSV (comma-separated values). 24 rows (header + 23 data rows). 70 columns. UTF-8 with BOM encoding. Missing values are indicated by blank cells (no NA). Values are expressed as percentages (relative abundance). Column definitions: The file is organized into four sequential data blocks, each with a year column followed by relative abundance columns for taxa within that microbial group. Block 1 (Eukaryotic algae, panel a): Column 1 is Year1 (1902–2021), followed by relative abundance data (in %) for 8 algal taxonomic groups: Compsopogonales, Diatomea, Dinoflagellata, Kathablepharidae, Pavlovophyceae, Phragmoplastophyta, Prymnesiophyceae, and Rhodellophyceae. Block 2 (Fungi, panel b): Column 9 is Year2 (1902–2021), followed by relative abundance data (in %) for 13 fungal groups: Aphelidea, Ascomycota, Basidiomycota, Blastocladiomycota, Chytridiomycota, Cryptomycota, Hyphochytriomycetes, Labyrinthulomycetes, Microsporidia, Mucoromycota, Neocallimastigomycota, Peronosporomycetes, and Zoopagomycota. Block 3 (Invertebrates, panel c): Column 22 is Year3 (1902–2021), followed by relative abundance data (in %) for 9 invertebrate groups: Annelida, Arthropoda, Cnidaria, Gastrotricha, Mollusca, Nematozoa, Platyhelminthes, Porifera, and Rotifera. Block 4 (Protozoa, panel d): Column 31 is Year4 (1902–2021), followed by relative abundance data (in %) for 16 protozoan groups: Apicomplexa, Bicosoecida, Breviatea, Cavosteliida, Centrohelida, Dictyostelia, Euglenozoa, Gracilipodida, Heterolobosea, Parabasalia, Preaxostyla, Protosporangiida, Protosteliida, Retaria, Rigifilida, and Schizoplasmodiida. Data organization: Each row represents a time point (approximately 1902–2021) with relative abundance data for all four microbial eukaryotic groups. Different groups have the same temporal resolution for most periods. Rows are ordered from youngest (top, ~2021) to oldest (bottom, ~1902). For each group, relative abundances sum to 100% per time point. This phylum-level dataset complements the genus/species-level data in Fig. S9, allowing examination of ecological shifts at different phylogenetic resolutions. Correspondence to manuscript figures: This file is used to generate Figure S11. Eukaryotic algae phylum/class-level data (Block 1) are used for panel (a); fungi phylum-level data (Block 2) for panel (b); invertebrate phylum-level data (Block 3) for panel (c); and protozoa phylum-level data (Block 4) for panel (d). The stacked bar charts display the relative abundance of major phyla over the past 120 years, revealing long-term shifts in microbial eukaryotic community composition in Lake Liangzi. 11. Fig.S12.csv – Data for Figure S12 File purpose: Time‑series dataset of alpha‑diversity (species richness and Shannon diversity index) and beta‑diversity (turnover and CAP ordination) indices for five multi‑trophic biotic groups (eukaryotic algae, fungi, invertebrates, protozoa, and aquatic macrophytes) from sedimentary ancient DNA (sedaDNA) analysis of Lake Liangzi sediments. This dataset spans approximately 1902–2021 and provides the diversity metrics used to examine temporal variations in community structure across different ecological regimes. The ASV (Amplicon Sequence Variant) data are used for β‑diversity analyses including turnover calculations and CAP (Canonical Analysis of Principal Coordinates). File structure: CSV (comma‑separated values). 24 rows (header + 23 data rows). 75 columns. UTF‑8 with BOM encoding. Missing values are indicated by blank cells (no NA). Column definitions: The file is organized into eight sequential data blocks, each with a year column followed by diversity metrics or ASV abundance data for specific biotic groups. Block 1 (Algae alpha‑diversity): Column 1 is Year1 (1902–2021), followed by Column 2 Richness1 (algal species richness), Column 3 Shannon1 (algal Shannon diversity index), and Column 4 Turnover1 (algal community turnover rate). Block 2 (Fungi alpha‑diversity): Column 5 is Year2 (1902–2021), followed by Column 6 Richness2 (fungal species richness), Column 7 Shannon2 (fungal Shannon diversity index), and Column 8 Turnover2 (fungal community turnover rate). Block 3 (Invertebrates alpha‑diversity): Column 9 is Year3 (1902–2021), followed by Column 10 Richness3 (invertebrate species richness), Column 11 Shannon3 (invertebrate Shannon diversity index), and Column 12 Turnover3 (invertebrate community turnover rate). Block 4 (Protozoa alpha‑diversity): Column 13 is Year4 (1902–2021), followed by Column 14 Richness4 (protozoan species richness), Column 15 Shannon4 (protozoan Shannon diversity index), and Column 16 Turnover4 (protozoan community turnover rate). Block 5 (Algae ASV data for β‑diversity): Column 17 is Year5 (1902–2021), followed by Columns 18–25 (ASV1_1 to ASV8_1) containing read abundances for 8 algal ASVs. Block 6 (Fungi ASV data for β‑diversity): Column 26 is Year6 (1902–2021), followed by Columns 27–39 (ASV1_2 to ASV13_2) containing read abundances for 13 fungal ASVs. Block 7 (Invertebrates ASV data for β‑diversity): Column 40 is Year7 (1902–2021), followed by Columns 41–49 (ASV1_3 to ASV9_3) containing read abundances for 9 invertebrate ASVs. Block 8 (Protozoa ASV data for β‑diversity): Column 50 is Year8 (1902–2021), followed by Columns 51–65 (ASV1 to ASV15) containing read abundances for 15 protozoan ASVs. Data organization: Each row represents a time point (approximately 1902–2021) with paired alpha‑diversity metrics and ASV abundance data for all five biotic groups. Different groups may have slightly different temporal resolutions. Rows are ordered from youngest (top, ~2021) to oldest (bottom, ~1902). The alpha‑diversity metrics (richness, Shannon index) provide measures of within‑sample diversity, while turnover rates quantify the rate of species composition change between consecutive time points. The ASV abundance data are used to calculate β‑diversity (community dissimilarity) through CAP ordination and other multivariate analyses. Correspondence to manuscript figures: This file is used to generate Figure 12. Alpha‑diversity data (Richness and Shannon index) for algae, fungi, invertebrates, protozoa, and macrophytes are used to plot temporal variations in α‑diversity. Turnover rates quantify β‑diversity through community turnover, and ASV abundance data are used for CAP analysis to visualize community composition changes over time. The figure presents multi‑trophic community diversity dynamics, showing how species richness, Shannon index, turnover, and CAP ordination vary across the 120‑year record. 12. Fig.S13.csv – Data for Figure S13 File purpose: Regional integration dataset of total organic carbon (TOC) records reconstructed from 59 lake sediment cores across China. This dataset compiles TOC measurements (% or g/kg) from published and unpublished sediment records, providing a comprehensive regional synthesis of lake carbon burial dynamics across diverse climatic and environmental gradients. The dataset supports the analysis of regional patterns, temporal trends, and drivers of sedimentary organic carbon accumulation in Chinese lakes. File structure: CSV (comma‑separated values). 76 rows (header + 75 data rows). 61 columns. UTF‑8 with BOM encoding. Missing values are indicated by blank cells (no NA). Column definitions: Column 1 is Year, the age of the sediment sample in year AD. Columns 2–60 contain TOC data (%) for individual lakes: Lake_Hongjiannao, Lake_Chenghai, Lake_Hulun, Lake_Changdang, Lake_Chao, Lake_Erhai, Lake_Lugu, Lake_Gonghai, Lake_Yuelianghu, Lake_Sihailongwan, Lake_Taihu, Lake_Gyaring, Lake_Fuxian, Lake_Yangzong, Lake_Guozaco, Lake_Gaoyou, Lake_Luoma, Lake_chongping, Lake_Mulong, Shade_Co, Lake_Xingyun, Lake_Beilianchi, Lake_Cuoqia, Lake_Weishan, Lake_Shitang, Lake_Jili, Lake_Kanas, Lake_Ailike, Lake_Taibai, Lake_Zhangdu, Lake_Moon, Lake_Haixi, Lake_Chahei, Lake_Yuxian, Lake_Tiancai, Lake_Xi, Lake_Jirencuo, Lake_Taiji, Woducuo, Yunlongtianchi, Taipingshuiku, Lake_Qilu, Lake_Xinyi, Lake_Julong, Lake_Baiyangdian, Lake_Wulanpao, Lake_Shengjin, Lake_Dianshan, Lake_Wang1, Lake_Longgan, Lake_Hongze, Lake_Xiliang, Lake_Nanyi, Lake_Liangzi, Lake_Qianghai, Lake_Poyang, Lake_Wang, Lake_Honghu, and Lake_Shijiu. Column 61 is Average, the mean TOC value across all 59 lakes at each time point. Data organization: Each row represents a time point (1900–2016) with TOC measurements from 59 lakes. Different lakes may have varying temporal resolutions and data coverage; blanks indicate no data for that lake at that time point. Rows are ordered from youngest (top, 2016) to oldest (bottom, 1900). The dataset includes lakes from diverse geographic regions across China, spanning different climate zones (from temperate to tropical), trophic states, and catchment characteristics. The average column provides a regional composite record, smoothing site‑specific variability to reveal common regional trends in lake carbon dynamics over the past century. Correspondence to manuscript figures: This file is used to generate Figure 13. TOC data for each lake are plotted as time series to show individual lake trends, while the Average column is used to highlight the regional mean trend. The figure presents the regional integration of TOC records, revealing coherent patterns of organic carbon accumulation across Chinese lakes and their relationship to regional environmental changes.



