KneechatData:Conversational Dataset from Surgical Patient Interviews
收藏资源简介:
This dataset contains transcribed and annotated excerpts from 41 semi-structured preoperative interviews with patients on the waiting list for total knee arthroplasty (TKA). The corpus comprises 635 utterances (segments), annotated with utterance type, thematic categories and an valence score. Origin and data collectionInterviews were conducted by trained research staff at Hospital Costa del Sol (Málaga, Spain). Conversations were recorded, transcribed verbatim, and segmented into utterances. Of the 635 segments, 226 were labeled as Spontaneous Comment and 55 as Explicit Question; the remainder were marked Clinically Irrelevant. Processing and annotations Thematic annotations were produced via a hybrid pipeline: automated phrase clustering and zero-shot classification followed by manual review by clinical researchers. Sentiment scores were computed automatically for each utterance using the Spanish-language BERT model pysentimiento/robertuito (Hugging Face). Valence is provided as a continuous score in the range [−1,+1][-1, +1][−1,+1] and as a discretized category (Very Negative / Negative / Neutral / Positive / Very Positive). Variables (main columns) Patient ID — pseudonymized patient identifier (e.g., P_1). Utterance (Spanish) — transcribed text of the segment in Spanish. Utterance Type — Spontaneous Comment / Explicit Question / Clinically Irrelevant. Categories "Explicit Question" — thematic label(s) for questions (e.g., Logistics & waiting times, Recovery process). Categories "Spontaneous Comment" — thematic label(s) for comments (e.g., Pain & complications, Daily routine, Fear / worry / anxiety). Valence Score — continuous saentiment valence (float, −1.0 to +1.0). Valence Category — discretized valence label derived from Valence Score. Valence computationValence = P(positive) −P(negative), where probabilities are produced by the RoBERTuito model for each utterance. Higher values indicate more positive affect; lower values indicate more negative affect. Ethics and privacyAll data have been pseudonymized: direct identifiers (names, addresses, phone numbers, etc.) were removed prior to release. The original study received ethics approval (Protocol No. 003_feb24) and all participants provided written informed consent for the use of de-identified data in research. Although anonymized, interview text may contain sensitive information; users should handle the data responsibly and comply with applicable data protection rules. Contact & citationFor questions, contact: mzapaher@gmail.com Please cite the dataset and the related manuscript when using these data.



