Measuring Judicial Formalism at Scale: Three Pilot Studies (U.S. Supreme Court, German Federal Constitutional Court, Court of Justice of the EU) - Preliminary Hypothesis Testing
收藏资源简介:
Data and reports from three pilot studies testing a comparative, multidimensional measure of legal formalism, run with the LawAnnotation formalism pipeline (schema v3.1) — an LLM-based instrument that scores judicial opinions on 32 rhetorical variables, aggregated into five operational tenet scores (Legal Autonomy, Textual Determinacy, Judicial Impersonality, Institutional Restraint, Doctrine over Consequences) and an overall formalism index ranging from −1 (maximally anti-formalist) to +1 (maximally formalist). Corpora: the U.S. Supreme Court (50 cases → 128 per-justice segments, CourtListener, 2016–2025), the German Federal Constitutional Court (50 decisions, stratified 5 per year, 2016–2025), and the Court of Justice of the EU (45 judgments plus 34 paired Advocate-General opinions, CELLAR/EUR-Lex, 2016–2025), with a three-way cross-jurisdiction comparison. All runs used identical prompts, coding schema, temperature 0, and double-run validation. Each corpus was also scored under a second model routing; the two tiers agree at r = 0.979–0.995 per court, so no finding depends on model choice. The deposit contains the main report and standalone figures (PDF), per-corpus source reports (Markdown), per-segment and per-decision score and classification datasets (CSV), comparison tables, stability diagnostics, and full methodology documentation. Developed as proof of concept for PostCritLaw — The Life of the Law after Critique (ERC Advanced Grant 2026 proposal, PI Michal Alberstein, Bar-Ilan University), building on Alberstein, Gabay-Egozi & Bogoch, Between Formalism and Discretion, 47 Hofstra L. Rev. 1103 (2019), and its manually coded corpus (doi:10.5281/zenodo.2591114). Preliminary research output, not peer-reviewed; all reliability figures are LLM double-run self-agreement, not human-validated accuracy.



