Relative Age Effect in South American professional football (2026 season)
收藏资源简介:
Supporting data for the article "Persistence of the Relative Age Effect (RAE) in South American professional football: A comprehensive CONMEBOL analysis". Birth quarter of 9,429 footballers from 319 clubs, across the first and second divisions of all ten CONMEBOL countries. The data were captured in March 2026. A subsample of 3,268 players from six first divisions mirrors the scope of Yagüe et al. (2020), and it is the one that makes the temporal comparison possible. The deposit holds a CSV with one row per player, frequency tables by country, division, quarter and month, the article tables, two sensitivity analyses, a field dictionary and a methods note. It also holds code/reproduce_tables.py, which recomputes Tables 3 and 4 reading only the CSV shipped here, with no internet and no access to the original source, and checks every cell against the published values. Where the data come from. From club squad listings published on Transfermarkt.es, extracted in March 2026 with a program written for this study. That commercial database is not redistributed here: what is published is a minimal derivative — country, competition, division, club, and month, quarter, semester and year of birth — leaving out the name, profile URL, exact day of birth, playing position, height and market value. None of those fields enters any calculation in the article. How the data were captured and derived is described in docs/methods_note.md, in enough detail to reimplement it; the extraction tool is not included and the author provides it to anyone who asks with an academic reason. Transfermarkt has no connection to this work. The data are pseudonymised, not anonymous. Player and club identifiers are truncated HMAC-SHA256 hashes, computed with a key that is not published. That does not make the rows anonymous: the subjects are professional footballers with public squad listings, and anyone with access to the source could recognise someone by combining club, division and month of birth. That is why the exact day of birth and every individual attribute were dropped. Two limits, each documented with its sensitivity analysis. 82 of the 9,429 rows repeat a player already counted, because three clubs appear registered in two divisions at once; there are 9,347 distinct players and dropping those rows leaves the aggregate odds ratio at 1.96. And the comparison against Yagüe et al. (2020) uses the estimator those authors published, which is an odds ratio from a 2×2 table rather than the Q1/Q4 ratio of counts this study reports. See docs/methods_note.md. Licences. The data under CC BY 4.0, the code under MIT. One exception: data/external/table1_yague2020_published.csv transcribes values from Yagüe et al. (2020), is not original work and CC BY 4.0 does not cover it; it is included with attribution.



