遇见数据集

Partially Automatically Annotated Corpus To Predict Gestural Cues In Embodied Conversational Agents

收藏
Zenodo2020-09-20 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

#Structure of the corpus This corpus has been built using speeches of Spanish politicians freely available here along with their transcriptions. Each transcription has been analyzed in terms of: Surface Syntactic Structure* Deep Syntactic Structure* Morphology (Part of Speech)* Communicative Structure Gestures (beat vs. no gesture tags) have been annotated using the videos. *All those features have been automatically retrieved using the parser freely available in https://github.com/TalnUPF/miis. The other features have been annotated manually. #Concerns about the corpus This corpus has been mostly annotated manually. Annotation agreement has not been computed. Moreover, it is small. In order to extract reliable correlations from it, it should be extended.

# 语料库结构 本语料库基于可在此处免费获取的西班牙政客演讲及其配套转写文本构建而成。 每条转写文本均从以下维度展开分析: 表层句法结构(Surface Syntactic Structure) 深层句法结构(Deep Syntactic Structure) 词法(词性标注,Morphology (Part of Speech)) 交际结构(Communicative Structure) 研究人员通过对应视频对手势(分为节拍手势与无手势两类标签)进行了标注。 * 上述所有特征均通过https://github.com/TalnUPF/miis中公开可用的句法分析工具自动提取,其余特征则采用人工标注方式完成。 # 语料库现存问题 本语料库的标注工作以人工为主,尚未计算标注一致性指标。 此外,该语料库规模较小。若要从中提取可靠的关联结论,需对其进行扩充。

提供机构:
Zenodo
创建时间:
2018-07-19
二维码
社区交流群
二维码
科研交流群
商业服务