Dataset of the paper "Understanding Students' Interactions with Large Language Models for Software Engineering Tasks"
收藏资源简介:
This dataset contains: Transcriptions of interviews conducted with 20 undergraduate students from the Computer Science course at [anonymized]; The answers from each participant to the proposed survey, which included questions about their profile, use of LLMs, and impressions of LLM responses. Each transcription is provided in both Portuguese and an English version to enable access to the original content and broaden global accessibility/readership. The English versions were translated using Gemini 2.5 Pro. All content has been anonymized to protect the participants' identities. It's important to highlight that participants provided explicit consent for the anonymized publication of their data. The file names in this dataset follow this pattern: interview_pX_pt.txt: The Portuguese transcription for a participant. "X" represents the participant's numerical ID (ranging from 1 to 20); interview_pX_en.txt: The English translation for a participant. "X" represents the participant's numerical ID (ranging from 1 to 20); form_answers.xls: The file containing the survey answers for all participants. The numerical IDs in this spreadsheet correspond to the IDs used in the transcription filenames. The qualitative coding file, generated using MAXQDA as part of the research process, is not provided here. This exclusion is necessary to protect participants' identities, as the file contains names and references to the participants.



