Research Artefacts for Auditing Computational Qualitative Analysis: Confirmability in LLM-Automated Grounded Theory Pipelines
收藏资源简介:
Preserved artefacts of two LLM-automated grounded theory pipelines run over the same corpus of AI-conducted professional interviews (Handa et al. 2025): a codebook pipeline (CB) that aggregates open codes into a frequency-ranked codebook before axial coding, and a record pipeline (RC) that passes all 2,883 coded records forward. The release contains, for each pipeline, its source code and prompts, every intermediate file, and its axial and theory outputs. These files are the evidence base for the paper's two analyses: what each category-forming stage received, and whether the quotations each output presents occur in the corpus. Two modifications were made for release: API credentials were removed from the code, and reviewer names in RC's record table were replaced by pseudonyms (R1–R7). No other file was altered. The interview corpus itself is not redistributed; it is available at https://huggingface.co/datasets/Anthropic/AnthropicInterviewer.



