Feedback text dataset with template-repeated free-text values: An adversarial diversity stress case
收藏资源简介:
This is an archival deposit of a CSV file named english_teaching.csv, downloaded from a public Kaggle dataset in March 2026. The dataset contains 3,580 rows with the following columns: Student_ID, Feedback_Text, Teaching_Clarity, Engagement, Effectiveness, Responsiveness, Overall_Score, and Label. Inspection of the file reveals that the Feedback_Text field contains only a small number of unique sentences repeated across the rows, with score vectors that vary independently of the text. This pattern is characteristic of template-based data generation rather than authentic student feedback, and suggests the original dataset was LLM-generated or otherwise synthetically constructed. The original Kaggle upload was no longer accessible at the time of archival, and its authorship could not be independently verified. This deposit is provided to enable reproducibility of the evaluation in the NIFTSy paper (Diaz Ramos et al., 2026, in submission), where english_teaching serves as an adversarial diversity-stress case, documenting the increasing presence of AI-generated content in publicly available datasets. We claim no authorship of the underlying data. This deposit is an archival copy provided solely for scientific reproducibility.



