Anonymized Dataset and Scripts for LLM Intervention in Chinese EFL: Reducing FLA and Enhancing Proficiency
收藏资源简介:
This dataset supports the study titled "The Role of Large Language Models in Reducing Foreign Language Anxiety and Enhancing English Proficiency for Chinese EFL Learners." The research employed a sequential explanatory mixed-methods design to evaluate an 8-week asynchronous speaking intervention using large language models (LLMs) such as DeepSeek and Kimi Chat) delivered via WeChat assignments to 52 intermediate-level Chinese EFL students at one of the Chinese university.The intervention aimed to provide low-stakes, non-judgmental speaking practice to address foreign language anxiety (FLA), increase speaking self-efficacy, and improve oral proficiency in a high-anxiety sociocultural context characterized by Confucian heritage values (e.g., fear of losing face) and exam-oriented education.Quantitative findings showed significant reductions in FLA (FLCAS: t(44) = -3.964, p < .001, d = 0.59), increases in self-efficacy (adapted ESSES: Wilcoxon signed-rank test, p < .001, r = 0.63), and gains in proficiency (VRT rubric: t(11) = 4.011, p = .002, d = 1.16), with engagement (OSE) positively correlated with proficiency gains (r = .682, p < .05). Qualitative thematic analysis of 470 interview turns from 12 participants identified key mechanisms, including non-judgmental AI feedback, cultural relevance of prompts, and risks of over-reliance.Contents of the dataset Quantitative_Data: Anonymized pre- and post-intervention scores for FLCAS, ESSES, OSE, and VRT; processed CSV files; master merged dataset.Qualitative_Data: Anonymized interview transcripts (full and individual speaker files using C-XX codes); initial coding and final refined themes (CSV).Code_and_Scripts: Python scripts for data cleaning, processing, statistical analysis (paired t-tests, Wilcoxon, correlations), thematic coding support, and VRT scoring.Documentation: Intervention summary PDF, anonymized consent form template, README with variable descriptions and replication guide, ethics and anonymization notes. How to replicateAll participant data is fully anonymized: no names, IDs, locations, or identifiable information remain. Transcripts and qualitative files use speaker codes (e.g., C-01 to C-30). Quantitative files use pseudonymized participant codes where applicable.To replicate quantitative results: Load Master_Quantitative_Data_Anonymized.csv and run quant_analysis_final.py or vrt_real_analysis.py for descriptive statistics, paired tests, effect sizes, and correlations.To replicate qualitative analysis: Use Final_Refined_Themes.csv and Initial_Human_Coding_Anonymized.csv together with read_transcripts.py or initial_coding.py for theme verification and LDA support.See README.txt for detailed variable descriptions, folder structure, and step-by-step instructions.This dataset enables full reproduction of the reported statistical results and thematic findings while strictly protecting participant privacy in accordance with informed consent and ethical guidelines.



