Replication Materials for "When the Agent Is the User: Agent-Driven Evolution of AI-Native Tools"
收藏资源简介:
Replication materials (with the paper PDF) for "When the Agent Is the User: Agent-Driven Evolution of AI-Native Tools" (Younmi Park, July 2026). The paper reports an 11-month longitudinal study (26 projects, 5,699 commits, corpus frozen 2026-07-05) in which AI agents, operating a knowledge management system as its primary users, filed structured usability feedback about the tool itself — a pattern the paper names agent-driven tool evolution. Under dual-rater classification (Cohen's κ = 0.90), 236 of 2,293 knowledge entries (10.3%) are tool-directed; the conservative floor is 3.2% in the fourteen projects outside the tool's own development. The paper formalizes AI-facing user interfaces (AFUI), redefines AI-native tools as the agent-primary class via three falsifiable criteria (C1–C3), and identifies articulation closure as the enabling mechanism. This archive contains the paper (PDF) and the materials behind every quantitative claim: - corpus_manifest.jsonl — one line per corpus entry (id, project, group, timestamps, status, label, SHA-256 body hash)- final_labels.json, coding_rules.md, raters/, disagreements.json, adjudications.json — the dual-rater Y/M/N classification: written rule set, raw rater outputs, 63 adjudicated disagreements- calib/ — calibration runs on three frozen projects (incl. the Mohadus 18/18 exact match)- subcode/ — friction-vs-backlog sub-coding of the 189 in-tool tool-directed entries (92% friction-grounded)- solicit/ — provenance probe on a random sample of 50 (seed 20260706): solicited vs. unsolicited composition- batch_flags.json, extract_stats.py — migration-sensitivity flags and the script reproducing all monthly rates, χ² tests, and confidence intervals- cited_entries.jsonl, post_freeze_cited_entries.md, standing_instructions_ko.md — full records of every entry cited by ID in the paper, and the Korean original of the standing agent instructions (Appendix A) Masking rule: full bodies of the 2,293 corpus entries are withheld (private business context from commercial projects); each entry's label, dates, status, project/group, and a SHA-256 hash of its body are published, so any later disclosure is verifiable against this freeze. For the two commercial tax-domain projects (haedream, homtax_scraping), titles in the manifest and rater rationales are additionally masked; the paper cites a few entries from these projects by ID only, and those are masked in the same spirit. No quantitative claim depends on a masked field. A claim-to-file verification map is included in the README. Bodies are available to reviewers on request. License: CC BY 4.0 (data and documents), MIT (scripts).



