Supplementary materials for "Streamlining learner corpus development with LLMs and NLP"
收藏资源简介:
This dataset contains the supplementary materials for the conference paper: Brooks, G. (2026). Streamlining learner corpus development with LLMs and NLP. JALTCALL Trends. The deposit provides the data and resources required to reproduce the reliability statistics and disagreement analyses reported in the paper. Contents: validation_subset.csv: A 93-token validation subset (3 essays) with per-model classifications, suggested corrections, consensus decisions, and human-in-the-loop flags. raw_classifications.zip: Per-essay JSON outputs containing full token-level classifications from three large language models. prompts.md: Full system prompts, few-shot examples, and API parameters used in the multi-LLM classification stage. compute_kappa.py: Standalone Python script to reproduce Fleiss’ κ and agreement statistics reported in the paper. README.md: Documentation describing the dataset structure and reproduction steps. LICENSE: CC BY 4.0 license. These materials support transparency and reproducibility in learner corpus development using large language models and natural language processing techniques. If you use this dataset, please cite both the accompanying paper and this Zenodo record.



