遇见数据集

Skeletal completeness and phylogenetic information of the dinosaur fossil record in Mexico

收藏
Zenodo2026-08-03 更新2026-08-13 收录
官方服务:

资源简介:

1.1 Data compilation The record of dinosaur specimens was aggregated from those publications that reported Mexican fossils accessioned into a collection and identified to either as Ornithischia, Sauropodomorpha or Theropoda, including Mesozoic birds. Reports that suggested, for instance, “fragments of dinosaur bones” or “indeterminate dinosaur remains” were excluded. Contentiously described material was not excluded, for instance, the vertebrae found in the Campanian of Chihuahua initially described as titanosaurid by Montellano-Ballesteros (2003) also resemble hadrosaurid vertebrae, but this is not conclusive (D’Emic et al. 2010). The analysis only considers osteological remains, and as such, egg fragments and footprints were excluded from the aggregation. Google Scholar (https://scholar.google.com) was primarily used for the search because, although it is independent from academic publishers and other metadata aggregators like Scopus or CrossRef, relying heavily on user curation, it is commonly used by academics due to its intuitiveness and readily access (see Smith et al. 2023). Clarivate and Web of Science require creating an account and several journals routinely used in palaeontology, such as Revue de Paléobiologie, are not indexed. The second source of information was the Paleobiology Database (https://paleobiodb.org), particularly due to its emphasis on temporal and spatial information that is validated by editors. A third source of information was the Léxico Estratigráfico Mexicano (LEM: https://sgm.gob.mx/Lexico_Es/), a list of all the lithostratigraphic units in Mexico on a government-supported free access website (Arriaga 2025). Each geological unit has an information sheet as downloadable PDF and includes both fossiliferous and non-fossiliferous outcrops. A survey through the available files allowed finding reports of dinosaur material documented in passing in other publications. Only the material that had an identification as a specimen, even if uncatalogued, was aggregated. The searches were performed between January and March 2026. For Google Scholar, the search included the terms “Mexico dinosaur diversity”, “Mexico Mexican dinosaur dinosaurs” relying on publications that would include an English abstract, and “dinosaurios mexicanos” to fetch results in Spanish. The results from the search in English mostly led to uploaded documents in ResearchGate (researchgate.net), and publishers like mdpi.com and other aggregators like sciencedirect.com, scielo.com and wiley.com. Although MDPI has faced recent scrutiny over its publishing practices (see Abalkina 2023), it is becoming commonplace in Mexican paleontology, thus the entries were not discarded by its provenance. The search in Spanish led to aggregators like redalyc.org and several science communication pieces published in university journals. In PBDB, the filters were set to “Triassic”, “Jurassic” and “Cretaceous” and “Dinosauria”. The LEM has a search engine that filters out by age and fetches exact matches. For instance, typing “Mesozoico” leads to only two units, namely Santa Eulalia (Baja California) and Rancho Viejo (Durango), whose age has not been constrained further into the Mesozoic. The word “Triásico” led to 31 results, “Jurásico” to 136, and “Cretácico” to 322. To navigate through the LEM, the geological framework described in 2.2 in this study (Figures 1 and 2) on the evolution of the Gulf of Mexico through the Mesozoic was used to select those units that had continental influence. This was complemented by going through the different units in the “Columna Estratigráfica” del LEM that works as an index of all the units mapped for the different states. The detailed geological framework produced for this study can be found in the Supplementary Material. To include a study, a collection number or specimen number needed to be clearly stated. A couple of reviews (Ramírez-Velasco and Hernández-Rivera 2015; Ramírez Velasco 2021) have compiled a catalogue with collection data for dinosaur fossils reported in other sources without accession information, and the data reported therein is aggregated instead. Conference abstracts were excluded even if they mentioned the specimen number. Several studies were retrieved through ResearchGate since they were behind a paywall or otherwise access-controlled and were excluded if the available file was not a final print copy. Some papers were retrieved through shadow libraries. This process was done manually, and no automated tools were developed to find the papers. Taxonomic and collection information tends to be incorporated as part of larger studies and is often neglected in academic literature due to publication trends, so each paper retrieved was individually assessed looking for collection numbers and a description of the nature of remains, even if brief and vague. Along with the specimen number or collection voucher, the following data was extracted from the studies: taxonomic status (if it is a type specimen, if it represents integumentary impressions, if the identification was done through open nomenclature, if there is an informal name given to the material), taxonomic determination, from more inclusive (namely Theropoda, Ornithischia, Sauropodomorpha) to species (going through family, subfamily and tribe), and the locality it was collected from (if mentioned, but if not, at least the Formation needed to be available). The age of the locality was obtained from the most recent sources available, using geochronological methods, i.e. radiometry, magnetostratigraphy or biostratigraphy. The age reported in the publication or the review was only aggregated to the database if the locality was mentioned but the formation is unknown, or if the formation has not been dated through geochronological methods. This was also not automated since this data is non-standardly reported. When reviews mentioned specimens in collections that have not been reported elsewhere, were only incorporated if the authors confirmed they saw the specimens first-hand. The list of sources analyzed are in the Supplementary Material. 1.2 Skeletal completeness metric (SCM) The skeletal completeness is an estimator to report how much a skeleton is preserved and is used both in current and fossil vertebrate remains. One approach, like the one developed for human forensics, relies on using average volume of human bones and assign each bone a rough percentage (Rowbotham et al. 2017, tbl. 1), with the advantage of having a better estimator for repetitive units with different sizes, for instance, metacarpals 2 (both elements averaging 5.6 cm3) and 3 (both totaling 8.8 cm3) getting each a skeletal completeness metric (SCM) of 3%, whereas metacarpal 4 (5.8 cm3) has a SCM=0.2%, and metacarpals 1 (5.6 cm3) and 5 (5.4 cm3) both get a SCM=0.18%. This granularity in an estimator is easier to do when having the same species, but becomes convoluted when dealing with dinosaurs, where the different clades have drastically different volumetric proportions between the skull and the bodies. For Sauropodomorpha, Mannion and Upchurch (2010, tbl. 1 and Supp.) assigned skeletal completeness metrics of 10% for the skull, 55% for the axial skeleton, and 35% for the appendicular skeleton. The skeletal volume can vary greatly between species, as allometric patterns can change through ontogeny, and across clades: for instance, whereas sauropod skulls can be very reduced and would represent less than 10% of the skeletal volume, in ornithischians like Triceratops, it can represent up to 33% of the skeletal volume. As a database management technique, this granularity would require a forced recomputation if, for instance, a titanosaurid (e.g. Montellano-Ballesteros 2003) is reinterpreted as a hadrosaur (e.g. D’Emic et al. 2010). Therefore, the skeletal completeness metric (SCM) was standardized to be applicable across the different dinosaur clades despite the widely variable skull, axial skeleton, and limbs proportions. The total score of 100 is distributed along the different regions of the body: a score of 10 is assigned to the skull, 51 to the axial skeleton, 35 to the appendicular skeleton, and 4 to non-skeletal bone, such as osteoderms, ossified tendons and ossified cartilague. The skeletal completeness metric (SCM) in the skull is distributed along the cranial regions: the upper arcade lateral series (SCM=1, comprising premaxillae and maxillae), the circumorbital series (SCM=2, lacrimal, prefrontal, postfrontal, jugal and postorbital), the median series (SCM=2, nasal, frontal, parietal, post-parietal and the bones of the palate), the cheek series (SCM=1, squamosal, quadrate, and quadratojugal), the braincase (SCM=0.5), the mandibular series (SCM=2.5, predentary, dentary, angular, surangular, articular, and splenial), and the teeth (SCM=1). The skeletal completeness of the axial skeleton is then subdivided into four regions (cervical vertebrae, SCM=15; dorsal vertebrae, SCM=15; sacral vertebrae, SCM=5; and caudal vertebrae, SCM=16), each considering vertebral bodies and ribs or chevrons as subdivisions. The scores for the appendicular skeleton are distributed along the pectoral girdle (SCM=5), forelimbs (SCM=12), pelvic girdle (SCM=6), and hindlimb (SCM=12). When compiling the data, the skeletal completeness metric was assigned according to what was reported accounting penalization for “incomplete” bones, and for the number of bones. For instance, two complete femora account for a SCM=4; if a paper reports one distal half of a femur, then it is scored as SCM=1. In some detailed descriptions, e.g., the monographic description of a new species, the assessment is easy, but when a paper mentions in passing the completeness of a bone, this is done to best capture the nature of specimen, even if overestimated, for instance, “limb remains” are scored as SCM=24 for both fore and hindlimbs. However, in other cases, such as compilations or lists in reviews, the description is given as “fragmentary femur” or “large bones”; in these cases, for instance, the penalization is to give values like “0.5” to account for the presence of one femur. 1.3 Character coverage metric (CCM) Here, the character coverage metric (CCM) is modified to account for the total number of elements needed to assess the hypothesis of homoplasy formulated by each character statement. If a character statement establishes the assessment of the connection between two bones, then that character statement is given a higher weight than one requiring only one bone to assess the hypothesis of homoplasy. For example, the character list by Baron et al. (2017) with 456 character statements has a character coverage of 531, meaning that some characters require multiple elements to be confidently evaluated. Some characters refer to overall regions like “skull” or “supratemporal fenestra”; these are given a weight of “1” because there are different ways to assess the overall shape of these regions when they are incomplete. The tally is then categorized into the same regions defined for the skeletal completeness metric (SCM), allowing for character statements to be represented in several regions. If a character requires assessing the maxilla and the jugal, then the tally gives “1” to each the lateral bearing tooth series and the cheek series; the rationale of this flexibility is that an anatomist may be able to infer the shape of the jugal when absent based solely on the morphology of the maxilla. Because the skeletal completeness metric has been optimized, e.g. a complete skull always scored as 10 regardless of the actual volumetric proportions, the character coverage metric should complement this to give an idea of the phylogenetic character richness of any specimen. For example, in ornithischians, the character list by Dieudonné et al. (2020), the lateral tooth-bearing series (rostral, premaxilla, maxilla) and the mandibular series (predentary, dentary, angular, surangular, articular and splenial) have 29 characters each, the cheek series (squamosal, quadrate and quadratojugal) has 33 and the circumorbital series (lateral, prefrontal, postfrontal, jugal and postorbital) has 37, meaning that five bones in a skeleton (circumorbital series) are required to assess 37 hypotheses of homoplasy, whereas only 3 bones in the cheek series are needed to assess a similar number. Each region was defined a priori for all the skeletons (see Supplementary material) and for each character list, the primary locators and qualifiers (sensu Sereno 2007) were extracted and counted, so that some characters are applied to entire regions (skull, orbit, infra- and supratemporal fenestra) whereas others are applied to several bones. For Dieudonné et al. (2020), 286 characters need 342 evaluations on the skeletal elements (Supplementary material, table 2) Initially, it seems that this is a very dynamic metric, with each character list having a different coverage, but within clades, the number of characters differs little. For instance, in sauropodomorphs, matrices often differ in the taxa included, few characters added, and modified scorings, with the character list growing little by little in each iteration (Regalado Fernández and Werneburg 2022). A statistical study on phylogenetic matrices to include all dinosaurs found that the matrices do not have statistically significant differences in the models they produce (Černý and Simonoff 2023), suggesting that within a clade, the character coverage would not differ drastically. However, the character coverage will be different at different hierarchical levels, for instance, between sauropodomorphs and ornithischians and theropods, because hypotheses of homoplasy are differentially distributed for each group. For this study, only character lists that were published as either part of the paper, an appendix or as associated supplementary material were analyzed. It is quite common that papers with phylogenetic analysis only show the final matrix with the taxa used but no character statements are appended to the file (Godoy et al. 2026, fig. 3). For the selected character lists, each character statement was decomposed in the osteological elements needed to assess it and the counts were categorized using the same regions defined for the SCM (Supplementary material, tables 1-9). The matrices chosen have the following hierarchical focuses: Dinosauria (Baron et al. 2017), Ornithischia (Dieudonné et al. 2020), Theropoda (Rauhut 2003), Sauropodomorpha (Pol et al. 2021), Ornithopoda (Madzia et al. 2020), Marginocephalia (Butler et al. 2011), Thyreophora (Sánchez-Fenollosa and Cobos 2025), and Titanosauria (Gorscak and O’Connor 2019). 1.4 Phylogenetic estimator A phylogenetic estimator should be a function that differentiates between specimens, estimates the usefulness of the occurrence as a datapoint and filters them according to the assumptions needed for analysis. The phylogenetic estimator, p̂, should reflect the spread of completeness and the phylogenetic character richness. Here, this estimator p̂ is computed as a cross product of the skeletal completeness metric and the character coverage (SCM × CCM). For instance, having two ilia from the same individual is more informative to score a character than only one ilium, because it is possible to differentiate preservation or deformation by evaluating the character twice, so the proportional value of skeletal completeness for one ilium is 3.45%, multiplied by the character coverage of the respective character list. To test the robustness and sensitivity of the estimator, four experiments were conducted to see how it is affected by the character list and its phylogenetic hierarchy scope and what the value of the estimator communicates using the compiled dinosaur fossil record from Mexico. The phylogenetic estimator then was used to test the intuition that the more incomplete the material is, the less phylogenetically relevant information it yields. The first experiment calculates the phylogenetic estimator p̂ as a constant applied to all values of skeletal completeness, using the character list by Baron et al. (2017) that compiled characters pertaining to the three lineages of dinosaurs. To plot the correlation between the estimator and the skeletal completeness metric, the values were ranked through a minimum-maximum normalization (p̂* to differentiate from non-standardized p̂), so that the most complete specimen in the dataset gets a value of 100. The second experiment considers the character coverage of the Dinosauria-focused dataset plus the character coverage of three matrices with a scope in Ornithischia (Dieudonné et al. 2020), Theropoda (Rauhut 2003) and Sauropodomorpha (Pol et al. 2021), and the normalized p̂* is calculated for the new values. The third experiment considers what happens when only the characters focused only on one of the intra-Dinosauria lineages. And the fourth experiment evaluates when a mixed-scope is approached, by calculating the character coverage from the character lists for Ornithopoda (Madzia et al. 2020), Marginocephalia (Butler et al. 2011), Thyreophora (Sánchez-Fenollosa and Cobos 2025), Neotheropoda (Dalman et al. 2024) and Titanosauria (Gorscak and O’Connor 2019), as well as the three higher-rank character lists from Ornithischia (Dieudonné et al. 2020), Theropoda (Rauhut 2003) and Sauropodomorpha (Pol et al. 2021) for partially determined specimens. Finally, two experiments were conducted to test the usefulness of the phylogenetic estimator p̂. In one experiment, the normalized values of the estimator (p̂*) were plotted against time to compare their trends with those of the skeletal character completeness. To visualize the correlation of SCM and the estimator p̂ over time, these values were divided into time bins of 4.9 My from 184 Mya (Toarcian, Early Jurassic) to 66 Mya (Maastrichtian, Late Cretaceous). For each specimen, the locality was associated with a geological formation. The age for the formation was taken from the most updated estimation performed, even if it did not match the age suggested in the original paper (see Supplementary material). Then, the tally counted the number of specimens that would be included in each time bin based on the category they fell in, for example, “late Maastrichtian” specimens tallied for the time-bin 73 to 66 Mya. A second experiment was to see if there were any similarities between the depositional environments with the lowest skeletal completeness metrics and phylogenetic estimator in the Mexican outcrops was consistent with a global Mesozoic-wide database of archosaurs in terms of assigned taxonomic hierarchy. All occurrences of Archosauria (fossils, ootaxa, and ichnites) for the time span between the Ladinian (240 Mya) and the Late Maastrichtian (66 Mya) were selected. From the 27,070 records, 23,936 (around 88%) are registered in collections from terrestrial environments, of which 8,166 (nearly 60% of the occurrences in terrestrial environments) are associated with fluvial deposits. The number is likely higher, since most records (11,745 entries) are classified as “indeterminate terrestrial environments” or simply “terrestrial” (full data set as Supplementary material).

提供机构:
Zenodo
创建时间:
2026-08-03
二维码
社区交流群
二维码
科研交流群
商业服务