遇见数据集

Comparing LLM and Human Interpretations of Oral History Interviews

收藏
Zenodo2026-07-30 更新2026-08-01 收录
官方服务:

资源简介:

This dataset documents a controlled comparison between a previously published human-led interpretation and an independent large language model (LLM) analysis of the same corpus of 100 oral history interviews with men who completed compulsory military service in Czechoslovakia or the Czech Republic between approximately 1968 and 2004. The experiment focused on the meanings narrators retrospectively attributed to compulsory military service within their life stories. The deposit contains the prompts used during the successive analytical stages, inductive and deductive pilot matrices, a case-level analytical matrix covering all 100 interviews, an AI-only corpus synthesis, a critical methodological audit, a revised LLM-generated typology, transcript-level adjudication findings, and the final Human–AI comparison. The materials document the development of the analytical workflow from individual interview analysis through corpus-level synthesis, critical revision, and comparison with the previously established human interpretation. The analysis followed a controlled no-retrieval design and was conducted through the Anthropic API. During the independent AI-only stages, the model had no access to the previously published interpretation, external sources, web search, retrieval systems, project knowledge, or outputs from other interviews. The initial pilot prompt was written by the author; subsequent prompts were generated by the model in response to the evolving analytical requirements and were reviewed, approved, and implemented by the author. All model outputs were treated as provisional analytical artefacts requiring human verification. The original interview transcripts and audio recordings are not included. Before LLM processing and public deposition, the research materials underwent extensive data minimisation and anonymisation. Direct identifiers and potentially identifying contextual information were removed or generalised, including names, years of birth, specific military units, service locations, functions, ranks, institutional positions, and exact years of service. The publicly deposited matrices contain only a level of descriptive detail consistent with the previously published human interpretation. The dataset is intended to support methodological research and teaching concerning LLM-assisted qualitative analysis, oral history, human–AI interpretive comparison, analytical auditing, research transparency, and the epistemic limits of computationally assisted humanities research.

提供机构:
Zenodo
创建时间:
2026-07-30
二维码
社区交流群
二维码
科研交流群
商业服务