遇见数据集

LDC Spoken Language Sampler - Sixth Release

收藏
DataCite Commons2023-08-22 更新2024-07-13 收录
官方服务:

资源简介:

LDC (Linguistic Data Consortium) Spoken Language Sampler - Sixth Release (LDC2023S07) contains samples from 20 different corpora published by LDC between 2020 and 2023. LDC distributes a wide and growing assortment of resources for researchers, engineers and educators whose work is concerned with human languages. Historically, most linguistic resources were not generally available to interested researchers but were restricted to single laboratories or to a limited number of users. Inspired by the success of selected readily-available and well-known data sets, such as the Brown University text corpus, LDC was founded in 1992 to provide a new mechanism for large-scale corpus development and resource sharing. With the support of its members, LDC provides critical services to the language research community that include: maintaining the LDC data archives, producing and distributing data via media or web download, negotiating intellectual property agreements with potential information providers and maintaining relations with other like-minded groups around the world. Resources available from LDC include speech, text, video and lexicons in multiple languages, as well as software tools to facilitate the use of corpus materials. For a complete view of LDC's publications, browse the Catalog. The LDC Spoken Language Sampler - Sixth Release provides speech and transcript samples and is designed to illustrate the variety and breadth of the speech-related resources available from the LDC Catalog. The sound files included in this release are excerpts. Most excerpts are truncated to be much shorter than the original files, typically about 2 minutes. Samples shorter than this typically represent the entirety of a single file.

创建时间:
2023-08-22
二维码
社区交流群
二维码
科研交流群
商业服务