<b>Interactional Variation Online: </b>harnessing emerging technologies in the digital humanities to analyse online discourse in different workplace contexts
收藏资源简介:
The IVO corpus is a collection of approx. 170,000 transcribed words of recorded virtual meetings held between July 2021 and July 2022, itemised in the 'IVO_corpus_file details' file. Recordings vary in length, number of participants, and meeting type.The ‘IVO core meetings corpus’ comprises meetings 1-15. They include 15 recordings from four different institutional contexts, ranging from municipal council meetings (DCC), a non-governmental organisation promoting arts (NCoL), an academic conference organising committee (TaLC) and a state-of-the-art software development company (GitLab). Some of these meetings are hybrid (i.e. some participants are in the same location). The meetings are agenda-driven and can be defined as workplace interaction. There are four remaining meetings (16-19) which are more representative of interviews, training sessions or presentations than meetings and so are not included in the IVO core meetings corpus. The IVO project was co-led by Anne O’Keeffe (anne.keeffe@mic.ul.ie) at Mary Immaculate College (MIC), Limerick and Dawn Knight (KnightD5@cardiff.ac.uk), at the Centre for Language and Communication Research, Cardiff University. The full project team comprised: 2 Principal Investigators (PI – Anne O’Keeffe, Dawn Knight), 2 Co-Investigators (CIs – Svenja Adolphs, Benjamin Cowan, Tania Fahey-Palma, Fiona Farr, Sandrine Peraldi), 1 Postdoctoral Researcher and 2 Research Associates over the course of the project. In addition, there were 9 academic advisors https://ivohub.com/gallery/. The project was co-funded by AHRC and IRC.This data in this corpus has been anonymised using a combination of manual and automated techniques. In addition to transcriptions of speech, the IVO core meetings corpus is tagged for selected nonverbal features. These include annotations for backchannels (head nods and spoken) in the first and last five minutes, emblematic gestures and meaningful gestures for each visible participant - saved as .eaf files (which can be opened in ELAN - see: https://archive.mpi.nl/tla/elan). The extent to which each recording was annotated for these features is detailed in the IVO_corpus_file_details (i.e. this varies from one file to the next). Where more than one feature was annotated, these were assembled into a single combined .eaf file. All .eaf files of the IVO core meetings corpus can be opened/reused in ELAN.The following files are included in this dataset:IVO_corpus_file_details: contains all information about the corpus file recordings (i.e. the meeting sessions captured within)Transcription conventions: guide to the conventions used in the corpus transcriptsYoutube Links to GitLab Videos: links to the source files that were transcribed in the corpus (users can locate these files and recreate the full multimodal corpus)ivo_meetings: containing all of the transcripts (without timestamps) from the entire corpus in a single file (.txt). This can be uploaded to a digital concordancing tool for further exploration (e.g. Sketch Engine)ivo_meetingscore: containing all of the transcripts (without timestamps) from the sample, 'core, corpus in a single file (.txt). This can be uploaded to a digital concordancing tool for further exploration (e.g. Sketch Engine)DCC1_emblems.eaf: DCC1 file annotated in ELAN - contains annotations for emblems onlyDCC2_combined.eaf: DCC3 file annotated in ELAN - contains annotations for backchannels at the start and end of the video (5 minutes each) and emblemsDCC3_combined.eaf: DCC3 file annotated in ELAN - contains annotations for meaningful gestures and emblemsDCC4_combined.eaf: DCC4 file annotated in ELAN - contains annotations for emblems onlyGit1_combined.eaf: Git1 file annotated in ELAN - contains annotations for meaningful gestures and emblemsGit2_combined.eaf: Git2 file annotated in ELAN - contains annotations for meaningful gestures and emblemsGit3_emblems.eaf: Git3 file annotated in ELAN - contains annotations for emblems onlyGit4_emblems.eaf: Git4 file annotated in ELAN - contains annotations for emblems onlyNCoL1_combined.eaf: NCoL1 file annotated in ELAN - contains annotations for meaningful gestures, emblems and backchannels at the start and end of the video (5 minutes each)NCoL2_emblems.eaf: NCoL2 file annotated in ELAN - contains annotations for emblems onlyNCoL2_emblems.eaf: NCoL2 file annotated in ELAN - contains annotations for emblems onlyNCoL4_combined.eaf: NCoL4 file annotated in ELAN - contains annotations for meaningful gestures, emblems and backchannels at the start and end of the video (5 minutes each)TaLC1_combined.eaf: TaLC1 file annotated in ELAN - contains annotations for meaningful gestures, emblems and backchannels at the start and end of the video (5 minutes each)TaLC2_emblems.eaf: TaLC2 file annotated in ELAN - contains annotations for emblems onlyTaLC3_emblems.eaf: TaLC3 file annotated in ELAN - contains annotations for emblems only
IVO语料库(IVO corpus)是一个收录约17万字转录文本的数据集,其内容为2021年7月至2022年7月期间录制的虚拟会议语音转写内容,详细条目记录于《IVO_corpus_file details》文件中。各段录音在时长、参会人数及会议类型上均存在差异。 “IVO核心会议语料库”涵盖编号1至15的会议,包含来自四类不同机构场景的15段录音,分别为市政议会会议(DCC)、艺术推广非营利组织(NCoL)、学术会议组委会(TaLC)以及顶尖软件开发公司(GitLab)。其中部分会议为混合式会议(即部分参会者处于同一线下场景)。此类会议均遵循既定议程,可归类为职场互动场景。 另有4段会议(编号16至19)更偏向访谈、培训课程或演示汇报,而非常规工作会议,因此未纳入IVO核心会议语料库。 IVO项目由利默里克圣玛丽学院(Mary Immaculate College, MIC)的Anne O’Keeffe(邮箱:anne.keeffe@mic.ul.ie)与卡迪夫大学语言与传播研究中心的Dawn Knight(邮箱:KnightD5@cardiff.ac.uk)联合领衔。项目全团队成员包括:2名首席研究员(PI——Anne O’Keeffe、Dawn Knight)、5名联合研究员(CI——Svenja Adolphs、Benjamin Cowan、Tania Fahey-Palma、Fiona Farr、Sandrine Peraldi)、1名博士后研究员与2名研究助理,全程参与项目推进。此外,项目还拥有9名学术顾问,相关信息可访问https://ivohub.com/gallery/。本项目由艺术与人文研究委员会(AHRC)与爱尔兰研究委员会(IRC)联合资助。 本语料库中的数据已通过手动与自动化结合的方式完成匿名化处理。除语音转写文本外,IVO核心会议语料库还针对选定的非语言特征进行了标注:包括针对每场视频首尾5分钟内的每位可见参会者的反馈信号(点头与口头反馈)、标志性手势与有意义手势的标注,相关标注文件保存为.eaf格式(可通过ELAN软件打开,详见https://archive.mpi.nl/tla/elan)。各段录音针对上述特征的标注范围详见《IVO_corpus_file_details》文件(不同文件的标注范围存在个体差异)。若某段录音同时标注了多种特征,则会将其整合为单个合并的.eaf文件。IVO核心会议语料库的所有.eaf文件均可在ELAN软件中打开或复用。 本数据集包含以下文件: 1. IVO_corpus_file_details:收录语料库录音文件的全部相关信息(即其中包含的会议场次) 2. Transcription conventions:语料库转写文本所遵循的标注规范指南 3. Youtube Links to GitLab Videos:指向语料库中转写所用源文件的YouTube链接,用户可通过该链接获取源文件,重建完整的多模态语料库 4. ivo_meetings:以单个.txt文件形式收录整个语料库的全部转写文本(不含时间戳),可上传至数字化语料检索工具(如Sketch Engine)进行进一步探索分析 5. ivo_meetingscore:以单个.txt文件形式收录核心语料库样本的全部转写文本(不含时间戳),可上传至数字化语料检索工具(如Sketch Engine)进行进一步探索分析 6. DCC1_emblems.eaf:可通过ELAN软件打开的DCC1标注文件,仅包含标志性手势标注 7. DCC2_combined.eaf:可通过ELAN软件打开的DCC2标注文件,包含视频首尾各5分钟的反馈信号标注与标志性手势标注 8. DCC3_combined.eaf:可通过ELAN软件打开的DCC3标注文件,包含有意义手势与标志性手势标注 9. DCC4_combined.eaf:可通过ELAN软件打开的DCC4标注文件,仅包含标志性手势标注 10. Git1_combined.eaf:可通过ELAN软件打开的Git1标注文件,包含有意义手势与标志性手势标注 11. Git2_combined.eaf:可通过ELAN软件打开的Git2标注文件,包含有意义手势与标志性手势标注 12. Git3_emblems.eaf:可通过ELAN软件打开的Git3标注文件,仅包含标志性手势标注 13. Git4_emblems.eaf:可通过ELAN软件打开的Git4标注文件,仅包含标志性手势标注 14. NCoL1_combined.eaf:可通过ELAN软件打开的NCoL1标注文件,包含有意义手势、标志性手势以及视频首尾5分钟的反馈信号标注 15. NCoL2_emblems.eaf:可通过ELAN软件打开的NCoL2标注文件,仅包含标志性手势标注 16. NCoL2_emblems.eaf:可通过ELAN软件打开的NCoL2标注文件,仅包含标志性手势标注 17. NCoL4_combined.eaf:可通过ELAN软件打开的NCoL4标注文件,包含有意义手势、标志性手势以及视频首尾5分钟的反馈信号标注 18. TaLC1_combined.eaf:可通过ELAN软件打开的TaLC1标注文件,包含有意义手势、标志性手势以及视频首尾5分钟的反馈信号标注 19. TaLC2_emblems.eaf:可通过ELAN软件打开的TaLC2标注文件,仅包含标志性手势标注 20. TaLC3_emblems.eaf:可通过ELAN软件打开的TaLC3标注文件,仅包含标志性手势标注




