遇见数据集

Boundary Transparency Index (BTI) Rater Pack v1.1: Scoring Rubric, Rater Materials, and Inter-Rater Reliability Instruments for Decision-Event Record Evaluation

收藏
Zenodo2026-03-16 更新2026-05-26 收录
官方服务:

资源简介:

This archive contains the complete rater materials package for the Boundary Transparency Index (BTI), a minimal inspection-based scoring rubric (0–12) for assessing the independent reconstructability of decision artefacts in AI-mediated decision systems. BTI scores six elements of a decision record (F1 Context, F2 Boundary, F3 Claim, F4 Evidence/Assumptions, F5 Uncertainty, F6 Oversight (H)), each rated 0/1/2 by blinded raters using only the text present in the record. Rater packets use operational shorthand (Options, Claims/Reason, Evidence Locator, Gaps, Oversight Gate); canonical field definitions are as above. Pack contents: — BTI v0.2 Technical Specification (canonical rubric definition, calibration anchors, IRR study plan, mapping appendix, limitations wall) — IRR Rater Packet (Rater 1) with embedded rubric and six stimulus records (3 baseline, 3 DER-structured) — IRR Rater Packet (Rater 2) with embedded rubric and six stimulus records (counterbalanced assignment) — IRR Scoring Template (blank .xlsx with rubric tab) — IRR Internal Key (record-to-event mapping and counterbalancing scheme) v1.1 update from v1.0: (1) Clarifying note added to F4 (Evidence/Assumptions): "A locator is precise if it allows a reviewer to find the exact source row, report, or document without additional search. File name alone is insufficient; row index or report ID is required for score = 2." (2) Clarifying note added to F5 (Uncertainty): "U-codes (U1–U6) are acceptable shorthand if a codebook is provided or if the gap category is self-evident from context. Listing 'uncertainty exists' without specification scores 0." (3) Clarifying note added to F6 (Oversight (H)): "Both actual gate action (what happened) and recommended gate action (what should happen) should be present for score = 2. A record stating only 'REVIEW' without distinguishing actual from recommended scores 1." This pack is designed for use in controlled evaluation studies comparing baseline documentation with structured Decision-Event Records (DERs). It supports the empirical programme described in the Two-Study Architecture (Study 1: MEU v1.0 synthetic validation; Study 2: naturalistic replication with real participants). No formally supervised empirical validation study has been undertaken using this instrument at the time of archival. Preliminary pilot reliability signal: ICC(2,1) = 0.977 (N = 6 records, 2 raters), reported in DSS manuscript (v3.8) for transparency only. Not a formal validation study. BTI measures reconstructability. It does not measure correctness, fairness, bias, safety, or compliance. Extended pilot calibration (Study B, N=12 matched pairs, 2 naive raters) reported in BTI Pilot Calibration Study v1.0 (DOI: 10.5281/zenodo.19052254). Indexed in the project Root Index (https://doi.org/10.5281/zenodo.17992916). Author: Hon Bor So ORCID: 0009-0008-2768-7494 Author note: The author also publishes under the name Sing So (Chöndrel Dorje).

提供机构:
Zenodo
创建时间:
2026-03-12
二维码
社区交流群
二维码
科研交流群
商业服务