Marthasterias glacialis ddRADseq datasets
收藏资源简介:
This repository contains the scripts and processed datasets used in the manuscript: Martín-Huete et al. (2026). Hidden genomic structure and widespread structural polymorphism across environmental gradients in the spiny sea star Marthasterias glacialis. Proceedings of the Royal Society B. (under review) The repository includes the scripts used to identify and characterize putative chromosomal inversions, infer neutral population structure and perform genotype–environment association analyses. Data availability Raw ddRAD reads are available at NCBI SRA: BioProject PRJNA1271299. Datasets Three datasets were generated from the original STACKS output and used throughout the analyses. Dataset Description Complete_dataset.vcf.gz 247 individuals and 31,766 SNPs after clone removal and quality filtering. Used for PCA, heterozygosity analyses and haplotype assignment of candidate inversions. High_density.vcf.gz 83 individuals and 339,639 SNPs. Used exclusively for local PCA (LOSTRUCT) to identify candidate chromosomal inversions. Neutral_dataset.vcf.gz 24,346 SNPs genotyped in 247 individuals, derived from the complete dataset after removing candidate inversion regions, loci deviating from Hardy–Weinberg equilibrium, linked SNPs and FST outlier loci. Used to infer neutral population structure (pairwise FST, PCA and STRUCTURE) and as the neutral background in genotype–environment association analyses (pRDA). Scripts Script Description 01_localPCA_lostruct.R Detection of candidate haploblocks using local PCA and multidimensional scaling (lostruct). 02_characterize_candidate_inversions.R Characterization of candidate inversions using PCA, heterozygosity, linkage disequilibrium, FST, and haplotype assignment. 03_prepare_neutral_dataset.txt Workflow and command-line instructions to generate the neutral dataset by removing inversion regions, Hardy–Weinberg outliers, linked loci and FST outliers. 04_population_structure.R Population structure analyses using pairwise FST, PCA and STRUCTURE. 05_environmental_association.R Environmental association analyses using pRDA with SNP genotypes and haplotype assignments. 06_SD_and_centromere_detection.txt Workflow used to identify segmental duplications associated with candidate inversion breakpoints and to identify putative centromeric regions across the genome. Additional files File Description popmap_247ind.txt Population map containing the IDs and population assignments of the 247 individuals included in the complete and neutral datasets. popmap_83ind.txt Population map containing the IDs and population assignments of the 83 individuals included in the high-density dataset. samples_env.csv Environmental data used for the pRDA analyses. all_inv_NEW.bed BED file containing the genomic coordinates of the 16 candidate chromosomal inversions. Files beginning with `structure_` correspond to the input files required to perform the Bayesian clustering analyses in STRUCTURE v2.3.4. The uploaded files correspond to the K = 1 analysis. To run analyses for other values of K (K = 2–6), modify the `NUMPOPS` parameter in `structure_command_MG1.sh` and the corresponding `NUMPOPS` value in `structure_mainparams` before running the analysis. Analyses were performed using: - 10 independent replicates per K - Burn-in: 50,000 iterations - MCMC: 250,000 iterations The optimal number of genetic clusters was determined using the ΔK method of Evanno et al. (2005).



