DelibAI pilot: replication package (derived data and analysis scripts)
收藏资源简介:
# DelibAI pilot — replication package (derived data and analysis scripts) Companion to: *Contestable AI Feedback in Digital Public Participation: Design and Pilot Evaluation of DelibAI* (Vagnoni & Palmirani) and the CERCA 2026 workshop paper *DelibAI: A Pilot Platform for AI-Augmented Deliberation with Contestable Fallacy Feedback*. Pilot: 7–30 May 2026; 34 sessions, 23 active participants, 193 contributions, 451 raw flag records, 1,990 logged events, 121 DelibAI review calls with three parallel model runs each. ## Contents - `data/` — derived tables used for every reported result (latest-valid votes, evaluator–comment judgments, participant condition means, review sequences, rewrite similarity, ideology dyads, per-review model runs). - `results/` — JSON summaries reproducing all statistics reported in the papers. - `scripts/` — the analysis and figure-generation pipeline (`requirements.txt` pins the environment; the numeric analyses use the Python standard library only). ## Privacy and what is not included - Participant codes are **re-pseudonymised** (`P01`–`P34`) consistently across all files; they do not match the codes used on the platform or visible in any figure. - **Not included:** the raw database export (it contains authentication material and participants' free-text comments, drafts, and AI rewrites) and the table of rewrite texts. Free text is released only after screening for personal data. Researchers may request restricted access from the corresponding author. - Consequently, the scripts document the exact computations but require the restricted raw export to run end to end; every intermediate and final output they produce is provided here. ## Key definitions See `results/primary_within_subject_summary.json` (`definitions`) and the supplementary material of the journal article. In brief: the primary outcome is the share of non-self peer evaluators judging a peer-authored comment problematic (latest valid state per evaluator and channel; seeds and self-votes excluded), averaged per author and condition; paired contrast with exact sign-flip test and 50,000-resample bootstrap (seed 20260901). The historical list-render rate (`comment_level_metrics.csv`, `problem_flag_rate`) is retained as an instrumentation check only and is not a reader-normalised outcome. ## Licence Data and documentation: CC BY 4.0. Scripts: MIT. ## Funding ERC project HyperModeLex (grant agreement No. 101055185).



