遇见数据集

Digital Archive of Southern Speech

收藏
Mendeley Data2024-01-31 更新2024-06-28 收录
官方服务:

资源简介:

Introduction Digital Archive of Southern Speech (DASS) was developed by the University of Georgia. It is a subset of the Linguistic Atlas of the Gulf States (LAGS), which is in turn part of the Linguistic Atlas Project (LAP). DASS contains approximately 370 hours of English speech data from 30 female speakers and 34 male speakers in .wav format and in .mp3 format, along with associated metadata about the speakers and the recordings and maps in .jpeg format relating to the recording locations. LAP consists of a set of survey research projects about the words and pronunciation of everyday American English, the largest project of its kind in the United States. Interviews with thousands of native speakers across the country have been carried out since 1929. LAGS surveyed the everyday speech of Georgia, Tennessee, Florida, Alabama, Mississippi, Arkansas, Louisiana, and Texas in a series of 914 audio-taped interviews conducted from 1968-1983. Interviews average approximately six hours in length the systematic LAGS tape archive amounts to 5500 hours of sound recordings. DASS is a collection of 64 interviews from LAGS selected to cover a range of speech across the region and to represent multiple education levels and ethnic backgrounds. This release is distributed on an external hard drive and contains instructions for using the media and navigating to the LICHEN program. Digital Archive of Southern Speech - NLP Version (LDC2016S05), an alternate version suitable for natural language processing and human language technology applications is also available. Data The DASS speakers average age is 61 years there are 30 women and 34 men from the Gulf States region represented in this release. The interviews cover common topics such as family, the weather, household articles and activities, agriculture and social connections. The interviews were originally recorded in the field on reel-to-reel audio tape. A digital version of every reel of tape was then made, one .wav file per reel, usually about one hour of sound. Each interview thus consists of a set of 3 to 13 reels, or roughly 3 to 13 interview hours. Personally identifying or sensitive information in the files was replaced with a tone to protect the privacy and to assure ethical treatment of speakers. Each .wav file is split into multiple .mp3 files based on the topic of conversation and labeled thusly. Included spreadsheets provide information about the speakers, the labels used for topics and the sound files. Also included in this release is a version of the LICHEN software developed at the University of Oulu, Finland. LICHEN allows users to browse and search through the audio data in a more advanced fashion using a graphical interface. Further information and instructions for LICHEN can be found within the docs folder of this release. Updates None at this time. Samples For an example of the data contained in this corpus, review this audio sample. Authorship The following people were involved with the DASS project: William A. Kretzschmar, Jr., Paulina Bounds, Jacqueline Hettel and Steven Coats University of Georgia Lee Pederson Emory University Lisa Lena Opas-Hänninen, Ilkka Juuso and Tapio Seppänen University of Oulu (Finland) Sponsorship The Atlas Data contained herein comprises information collected in the period spanning from the 1930s to 2010 and has been compiled from diverse sources, by, and under the direction of, Dr. William A. Kretzschmar, Harry and Jane Wilson Professor in Humanities at the Department of English of The University of Georgia. Compilation and digitalization of this work was funded, in part, by the US National Science Foundation and by the US National Endowment for the Humanities. Additional information about the Atlas Project can be obtained at http://www.lap.uga.edu/Home.html. Portions © 1982-2010 American Dialect Society, © 1986-2010 University of Georgia Research Foundation, © 2012 Trustees of the University of Pennsylvania

### 数据集简介 南方言语数字档案馆(Digital Archive of Southern Speech, DASS)由佐治亚大学开发。该数据集是海湾州语言图谱(Linguistic Atlas of the Gulf States, LAGS)的子集,而海湾州语言图谱本身又是语言图谱项目(Linguistic Atlas Project, LAP)的组成部分。DASS 包含约370小时的英语语音数据,采集自海湾州地区的30名女性发言者与34名男性发言者,数据格式为.wav与.mp3,同时附带与发言者、录音相关的元数据,以及用于标注录音地点的.jpeg格式地图。 语言图谱项目(LAP)是一系列针对日常美式英语词汇与发音的调研项目,是美国同类项目中规模最大的一项。自1929年起,研究团队已完成对全美数千名母语使用者的访谈。 海湾州语言图谱(LAGS)于1968年至1983年间开展,通过914份录音访谈,调研了佐治亚州、田纳西州、佛罗里达州、阿拉巴马州、密西西比州、阿肯色州、路易斯安那州与得克萨斯州的日常言语。该系列访谈单份平均时长约6小时,系统性的LAGS磁带档案总计包含5500小时的录音素材。 DASS 从LAGS中精选出64份访谈,旨在覆盖该区域内多样的言语特征,并代表不同教育水平与族裔背景的人群。本次发布的数据存储于外置硬盘中,附带媒体使用说明以及LICHEN程序的使用指引。此外,还提供适用于自然语言处理与人类语言技术应用的另一版本——南方言语数字档案馆自然语言处理版(Digital Archive of Southern Speech - NLP Version, LDC2016S05)。 ## 数据概况 本次发布的DASS发言者平均年龄为61岁,涵盖来自海湾州地区的30名女性与34名男性。访谈内容涵盖家庭、天气、家居用品与日常活动、农业以及社会关系等常见话题。访谈最初以开盘录音磁带在实地录制,随后每盘磁带均被转换为数字版本,每盘磁带对应一个.wav文件,时长通常约1小时。因此每份访谈由3至13盘磁带组成,对应约3至13小时的访谈内容。为保护发言者隐私并确保伦理研究规范,文件中所有可识别个人身份的敏感信息均已替换为提示音。每个.wav文件会根据对话主题拆分为多个.mp3文件,并据此命名。附带的电子表格提供了关于发言者、话题标签以及音频文件的相关信息。本次发布还包含由芬兰奥卢大学开发的LICHEN软件版本。LICHEN支持用户通过图形界面以更高级的方式浏览与检索音频数据。关于LICHEN的更多信息与使用说明可在本次发布包的docs文件夹中获取。 ## 更新说明 暂无更新计划。 ## 样本 如需查看该语料库中的数据示例,请参阅此音频样本。 ## 作者贡献 参与DASS项目的人员如下: 小威廉·A·克赖茨施马尔(William A. Kretzschmar, Jr.)、宝琳娜·邦兹(Paulina Bounds)、杰奎琳·赫特尔(Jacqueline Hettel)与史蒂文·科茨(Steven Coats)——佐治亚大学 李·佩德森(Lee Pederson)——埃默里大学 丽莎·莉娜·奥帕斯-汉宁宁(Lisa Lena Opas-Hänninen)、伊尔卡·尤苏(Ilkka Juuso)与塔皮奥·塞帕宁(Tapio Seppänen)——芬兰奥卢大学 ## 资助与编纂说明 本数据集包含的图谱数据采集于1930年至2010年期间,由佐治亚大学英语系Harry and Jane Wilson人文教授威廉·A·克赖茨施马尔博士牵头,从多种来源整合汇编而成。本作品的编译与数字化工作部分得到了美国国家科学基金会与美国国家人文基金会的资助。如需了解语言图谱项目的更多信息,可访问http://www.lap.uga.edu/Home.html。 本作品部分内容©1982-2010 美国方言学会,©1986-2010 佐治亚大学研究基金会,©2012 宾夕法尼亚大学托管会

创建时间:
2024-01-31
二维码
社区交流群
二维码
科研交流群
商业服务