遇见数据集

Reproducibility Materials for "The Semiotics of Family-Law Subject Construction: Gender and Actor Visibility in Polish, Swedish, and Argentine Legislation"

收藏
Zenodo2026-08-03 更新2026-08-13 收录
官方服务:

资源简介:

This repository contains the datasets, annotation documentation, quality-control records, statistical analysis code, and supplementary results accompanying the article The Semiotics of Family-Law Subject Construction: Gender and Actor Visibility in Polish, Swedish, and Argentine Legislation. The study examines how family-law legislation constructs legally recognisable subject positions through gender-specific reference, generic masculine extension, lexical and relational neutralisation, binary pairing or doubling, explicit inclusion, and actor-suppressing constructions. The analysis covers Polish, Swedish, and Argentine legislation concerning marriage, parentage, parental responsibility and contact, and adoption. The final analytical corpus comprises 493 statutory provisions and 2,533 researcher-validated actor-reference expressions. The reference-level dataset records the linguistic form, actor role, referential scope, reference strategy, and legal relevance of gender or reproductive distinctions. The provision-level dataset records provision length, actor-reference count and density, absence of overt actor reference, reference avoidance, and aggregated category counts. Candidate extraction was supported by a multilingual lexicon, regular-expression patterns, and automatically generated linguistic annotations. A codebook-constrained generative-AI procedure was used for independent pilot calibration and, after the codebook had been frozen, for provisional first-pass classification of the remaining candidate records. These outputs were not treated as final decisions or as independent reliability evidence. The researcher manually reviewed all 467 QA-flagged records and all 493 provisions, adjudicated all revisions, and retained sole responsibility for the final analytical datasets and interpretation. The repository includes: the final reference-level and provision-level analytical datasets; corpus metadata and a data dictionary; the frozen annotation codebook and extraction lexicon; pilot coding, agreement, adjudication, and gold-standard materials; documentation of the AI-assisted first-pass classification and subsequent researcher-led quality control; scripts for corpus construction and statistical analysis; the software requirements needed to reproduce the analyses; Supplementary Tables S1–S7; the final figure showing reference-strategy distributions by jurisdiction. The statistical workflow includes descriptive analysis, stratified cluster-preserving permutation tests, category-specific generalised estimating equations, negative-binomial regression, and provision-level logistic regression. The permutation analyses use 9,999 permutations and the fixed random seed 20260726. Complete statutory texts are not redistributed. The repository instead provides official citations, provision identifiers, source URLs, access dates, version information, corpus-construction documentation, and checksums supporting reconstruction and verification of the legislative corpus. The materials are intended to support transparency, inspection of the annotation and quality-control procedures, and reproduction of the reported quantitative analyses.

提供机构:
Zenodo
创建时间:
2026-08-03
二维码
社区交流群
二维码
科研交流群
商业服务