Data for: Language Models, Surprisal and Fantasy in Slavic Intercomprehension
收藏资源简介:
The file webresults_cloze_publication.xlsx contains two types of data: a) transcripts of think-aloud protocols and b) respones collected in a web-based intercomprehension experiment for the same stimuli respectively. Part a) Three Polish stimuli sentences were presented to pairs of Czech native speakers in an experimental setting where both participants saw the stimulus sentence on their computer screens. Placed in different rooms, they were asked to communicate over skype and work together in order to come up with a good Czech translation of the sentence. Hence, the experiment output are audio recordings of the two participants trying to decode the stimuli and the written translations they have entered during the experiment. The transcripts are in sheet 1, 3, and 5 of the .xlsx file. Part b) Czech readers (n=23) were asked to translate certain words or phrases within Polish sentences (those that turned out problematic in part a) into Czech in a web-based translation experiment in cloze task design over the website http://intercomprehension.coli.uni-saarland.de/en/. The responses of part b) and corresponding sociodemographic data are in sheet 2, 4, and 6 of the .xlsx file. The responses were checked manually for correctness. Responses with typos were counted as correct, for the main interest was to find out if respondents had understood the stimuli. The column "Total Time Spent (ms)" is the time respondents have spent on entering their response into the gaps in the cloze test until pressing enter. The file surprisal_scores_CS_LM.txt contains surprisal scores obtained from a statistical trigram language model with Kneser-Ney smoothing trained on a Czech corpus (Czech part of InterCorp merged with the Czech part of the Russian National Corpus, size: 175,190 words).
文件webresults_cloze_publication.xlsx 包含两类数据:a) 有声思维实验转录文本,以及b) 针对同一刺激材料开展的基于网页的互理解实验采集的响应数据。 a) 有声思维实验部分:实验中向成对的捷克语母语使用者呈现3条波兰语刺激语句,两名被试均可在各自电脑屏幕上查看该刺激语句。他们被安排在不同房间,需通过Skype进行沟通协作,共同产出该语句的优质捷克语译文。因此本实验的产出内容为两名被试尝试解读刺激语句的音频录音,以及实验过程中他们录入的书面译文。该部分转录文本存放在该xlsx文件的第1、3、5工作表中。 b) 完形填空实验部分:招募23名捷克语读者,在基于网页的完形填空式翻译实验中,要求其将波兰语语句中(即a)部分中被证实存在理解难度的词汇或短语)译为捷克语,实验平台为http://intercomprehension.coli.uni-saarland.de/en/。该部分的响应数据与对应的社会人口统计学数据存放在该xlsx文件的第2、4、6工作表中。所有响应数据均经过人工正确性校验。考虑到本实验核心关注的是被试是否理解了刺激材料,因此存在拼写错误的响应仍被计为正确。"Total Time Spent (ms)"列记录的是被试从在完形填空测试的空缺处录入答案直至按下回车键所花费的时长。 文件surprisal_scores_CS_LM.txt 包含通过统计三元语法语言模型计算得到的surprisal得分,该模型采用Kneser-Ney平滑算法,以捷克语语料库(InterCorp的捷克语部分与俄罗斯国家语料库的捷克语部分合并而成,总词量:175,190)作为训练数据。




