遇见数据集

NyU-BU contextually controlled stories Corpus: NUBUC

收藏
Mendeley Data2024-03-27 更新2024-06-27 收录
数据链接:
官方服务:

资源简介:

The success of a language experiment heavily relies on selecting appropriate stimulus materials. This selection process entails a critical trade-off between similarity to ‘real’ language (i.e. external validity) and experimental and analytic control (i.e. internal validity). In order to bridge these conflicting demands, we developed the NyU-BU contextually controlled stories Corpus (NUBUC) of spoken language. The corpus is both naturalistic and experimentally controlled, comprising 16 high-quality recordings of 8 unique stories, spoken both by a female and a male actor. Each story consists of 128 sentences (~2000 words per story) organized around critical keywords, which have been matched along multiple linguistic dimensions. The context surrounding each keyword is also parametrically manipulated, varying prior context (weak/strong), local context (weak/strong) and sentence position (early/late). Here we describe the corpus in detail, including how it compares to and builds on existent corpora. These materials showcase the ability to overcome the apparent dichotomy between control and generalizability, by presenting subjects with carefully curated linguistic materials in a naturalistic listening scenario.

语言实验的成功高度依赖于合适的刺激材料(stimulus materials)遴选。该遴选过程需在与“真实”语言的相似度(即外部效度(external validity))与实验及分析控制水平(即内部效度(internal validity))之间做出关键权衡。为调和这两种相互冲突的需求,我们开发了纽约大学-波士顿大学语境可控口语故事语料库(NyU-BU contextually controlled stories Corpus,简称NUBUC)。该语料库兼具自然性与实验可控性,包含8个独特故事的16段高质量录音,分别由一名女性演员与一名男性演员完成朗读。每个故事均围绕核心关键词构建,包含128个句子(单篇故事约2000词),这些关键词已在多语言维度上完成匹配。每个关键词所在的语境亦经过参数化操控,涵盖前置语境(弱/强)、局部语境(弱/强)以及句子位置(句首/句末)的变化。本文详细介绍了该语料库,包括其与现有语料库的对比分析及继承拓展路径。本套材料通过在自然化聆听场景中为被试呈现精心甄选的语言材料,展现了克服控制与可推广性之间看似不可调和的二分对立的可行性。

创建时间:
2023-06-28
二维码
社区交流群
二维码
科研交流群
商业服务