Data from: Phenotypic variation across the first steps of experimentally evolved multicellularity.
收藏资源简介:
his BBC_2026__README.txt file was generated on 2026-08-14 by Beatriz Baselga Cervera GENERAL INFORMATION 1. Title of Dataset and code: Data from: Phenotypic variation across the first steps of experimentally evolved multicellularity. 2. Author Information Corresponding Investigator Name: Dr. Beatriz Baselga-Cervera Institution: University of Minnesota Twin cities, Minnesota, US. Email: bbaselga@umn.edu; beabaselga@gmail.com Co-investigator 1 Name: Dr. Noah Gettle Institution: University of Denver, Denver, Colorado, US. Email: Noah.Gettle@du.edu Co-investigator 2 Name: Dr Michael Travisano Institution: University of Minnesota Twin cities, Minnesota, US. Email:travisan@umn.edu 4.Data collector: Dr Beatriz Baselga-Cervera 5. Date of data collection: 2021-2022 6. Geographic location of data collection: Saint Paul, US 5. Funding sources that supported the collection of the data: Fundación Alfonso Martín Escudero, Madrid, Spain and PPFP fellowship of the UMN both to Beatriz Baselga-Cervera 6. Recommended citation for this dataset: Baselga-Cervera et al. (2026), Data from: Phenotypic variation across the first steps of experimentally evolved multicellularity. Zenodo. Data set. DATASET DESCRIPTION This repository contains raw experimental raw data, simulation outputs, derived datasets, and paper figures data code associated with the paper entitled: Phenotypic variation across the first steps of experimentally evolved multicellularity. This study examines the genotype-to-phenotype map underlying the emergence of multicellularity in Saccharomyces cerevisiae. Phenotypic characterization was performed using: Coulter Counter Multisizer 4 FlowCam® 3.0 Fluid Imaging Technologies NetLogo agent-based simulations The study includes: Ancestral diploid wild-type strain Y55 Evolved multicellular strains C1W8.1 and C1W8.2 ACE2 knockout strains ACE2 missense mutant strains (ACE2 c.1934 A>T) For each strain, six independent isolates were characterized in two growth media: YPD (rich medium) SD (minimal medium) REPOSITORY ORGANIZATION archive/ │ ├── raw_data/ │ ├── Snowflake3d.nlogo3d │ ├── Counter_Counter_Raw_data.xlsx │ ├── Coulter_Counter_Raw_data_heterozygous_contructions.xlsx │ ├── Flow_Cam_Raw_data.xlsx │ └── FlowCamPicturesData.zip │ ├── derived_data/ │ ├── bootstrapped_means.csv │ ├── bootstrapped_vars.csv │ ├── Netlogo_model_data.xlsx │ ├── KDE_overlap_strain_media_isolate.csv │ ├── KDE_overlap_strain_media.xlsx │ ├── mean_values_from_coulter_counter.xlsx │ └── heatmap.csv │ ├── scripts/ │ ├── run_all.R │ ├── 01_bootstrap_statistics.R │ ├── 02_overlap_kde.R │ ├── 03_figure2B_2C.R │ ├── 04_figure3Ato 3E.R │ ├── 05_figure4.R │ ├── 06_figureS1.R │ ├── 07_figureS2.R │ └── 08_figureS3.R │ ├── software_versions/ │ └── sessionInfo.txt │ ├── Data_Dictionary.xlsx │ └── README.txt REPRODUCIBILITY WORKFLOW All analyses can be reproduced from a fresh R session. Step 1 Open R or RStudio. Step 2 Set the working directory to the top-level archive folder. Step 3 Run: source("scripts/run_all.R") The script run_all.R executes all analyses in the correct order and recreates all derived datasets and manuscript figures. ANALYSIS PIPELINE Scripts should be executed in the following order. Order Script Purpose Outputs 1 01_bootstrap_statistics.R Bootstrap mean and variance estimates bootstrapped_means.csv, bootstrapped_vars.csv 2 02_overlap_kde.R KDE overlap calculations KDE_overlap_strain_media_isolate.csv KDE_overlap_strain_media.xlsx 3 03_figure2.R NetLogo analyses Figure 2B–C 4 04_figure3.R Coulter Counter analyses Figure 3A–E 5 05_figure4.R Phenotypic variation analyses Figure 4 6 06_figureS1.R Heterozygous construct analyses Figure S1 7 07_figureS2.R Supplementary analyses Figure S2 8 08_figureS3.R Supplementary analyses Figure S3 FIGURE AND OUTPUT MAPPING Output Script Input files Figure 2B–C 03_figure2.R Netlogo_model_data.xlsx Figure 3A–E 04_figure3.R Counter_Counter_Raw_data.xlsx Figure 4 05_figure4.R Counter_Counter_Raw_data.xlsx, bootstrapped_means.csv, bootstrapped_vars.csv Figure 5 Figure_5_Graphs.xlsx Figure generated in Excel workbook Figure S1 06_figureS1.R Coulter_Counter_Raw_data_heterozygous_contructions.xlsx Figure S2 07_figureS2.R Counter_Counter_Raw_data.xlsx Figure S3 08_figureS3.R Counter_Counter_Raw_data.xlsx Bootstrap datasets 01_bootstrap_statistics.R Counter_Counter_Raw_data.xlsx KDE overlap datasets 02_overlap_kde.R Counter_Counter_Raw_data.xlsx SOFTWARE REQUIREMENTS Analyses were conducted using: R version 4.4.0 or later Required CRAN packages include: | Package | Version | | pheatmap | 1.0.13 | | gplots | 3.3.0 | | ggcorrplot | 0.2.0.9000 | | lubridate | 1.9.5 | | forcats | 1.0.1 | | readr | 2.2.0 | | tibble | 3.3.1 | | tidyverse | 2.0.0 | | cowplot | 1.2.0 | | gganimate | 1.0.11 | | ggpubr | 0.6.3 | | devtools | 2.5.2 | | usethis | 3.2.1 | | fitdistrplus | 1.2-6 | | survival | 3.8-6 | | ggplot.multistats | 1.0.1 | | htmltools | 0.5.9 | | gridExtra | 2.3 | | magrittr | 2.0.5 | | MASS | 7.3-65 | | plotly | 4.12.0 | | plyr | 1.8.9 | | reshape2 | 1.4.5 | | ggridges | 0.5.7 | | overlapping | 2.4 | | testthat | 3.3.2 | | overlap | 0.3.9 | | suntools | 1.1.0 | | ggplot2 | 4.0.3.9000 | | stringr | 1.6.0 | | tidyr | 1.3.2 | | dplyr | 1.2.1 | | purrr | 1.2.2 | | readxl | 1.5.0 | No packages are installed automatically by the analysis scripts. Users should install all required packages prior to running the workflow. Package versions used for the archived analysis are recorded in: software_versions.txt which contains the output of: sessionInfo() FILE DESCRIPTIONS Please see: Data_Dictionary.xlsx Raw Data o Counter_Counter_Raw_data.xlsx Coulter Counter measurements of population size distributions for all strains and media conditions. Measurements include: Volume (µm³) Diameter (µm) Columns correspond to individual strain isolates in YPD and SD media. o Coulter_Counter_Raw_data_heterozygous_contructions.xlsx Coulter Counter measurements for heterozygous ACE2 knockout and missense constructs. Measurements include: Volume (µm³) Diameter (µm) Columns correspond to individual strain isolates in YPD and SD media. o Flow_Cam_Raw_data.xlsx FlowCam morphological measurements for all strains. Variables include: Area-based diameter Aspect ratio Circle fit Equivalent spherical diameter Elongation Perimeter Roughness Volume Width o FlowCamPicturesData.zip FlowCam image files associated with particle measurements. o Snowflake3d.nlogo3d Netlogo code and session. Derived Data o bootstrapped_means.csv Bootstrap estimates of mean cell-cluster diameter used to calculate relative contributions to phenotypic variation. Variables: sample diameter_um o bootstrapped_vars.csv Bootstrap estimates of variance in cell-cluster diameter used to quantify phenotypic noise. Variables: sample diameter_um o KDE_overlap_strain_media.xlsx Kernel density overlap statistics calculated using the R package overlapping. Variables: var1 var2 value media o KDE_overlap_strain_media_isolate.csv Kernel density overlap statistics calculated using the R package overlapping. Variables: Strains values o mean_values_from_coulter_counter.xlsx Coulter Counter reads mean value per strain, media and isolate. Variables: sample media mean_diameter_um strain clone o heatmap.csv Overlap values arrange in a matrix. o Netlogo_model_data.xlsx Output from NetLogo simulations of multicellular snowflake yeast growth. Key variables include: run number size-of-turtle (cells) count cells division cycles (ticks) cluster radius maximum cluster distance bounding radius bounding volume bounding surface area



