Harmonic Frontier Audio - Plosives and Non-Lexical Consonant Bursts (Preview Pack v0.9)
收藏资源简介:
Harmonic Frontier Audio -- Plosives and Non-Lexical Consonant Bursts (Preview, v0.9) A high-fidelity human vocal dataset designed for AI training, speech research, and articulation-aware voice modeling. Plosives and Non-Lexical Consonant Bursts (Preview), created by Harmonic Frontier Audio, provides a compact reference set demonstrating the quality, formatting, and metadata conventions used in the Harmonic Frontier Audio Human Vocality Primitives series. 🔎 Summary This dataset provides high-quality, rights-cleared recordings of plosive articulations and short-duration non-lexical consonant burst gestures --- discrete vocal events produced through controlled vocal tract closure and rapid release. The recordings emphasize: - articulatory closure and release - transient airflow dynamics - burst intensity and envelope shape - non-linguistic consonant gestures These characteristics make the dataset valuable for AI speech and voice modeling, phonetics research, articulation-aware synthesis, onset modeling, and human-aligned vocal control systems. Developed by Harmonic Frontier Audio, this preview follows The Proteus Standard™ for dataset provenance, transparency, and ethical AI use.Learn more about the Proteus Standard → https://harmonicfrontieraudio.com/proteus-standard Full dataset details and licensing information are available at:https://harmonicfrontieraudio.com/datasets/plosives-non-lexical-consonant-bursts If you find this dataset useful, please consider giving it a 🤍 on Hugging Face to help others discover it. 🫁 About Plosives and Non-Lexical Consonant Bursts Plosives are produced by complete or near-complete closure of the vocal tract followed by a controlled release of air pressure, resulting in a short, high-energy acoustic burst.Non-lexical consonant bursts refer to similar transient gestures produced without linguistic intent or semantic content. These vocal behaviors are foundational to: - speech articulation and onset modeling - expressive and controllable voice synthesis - articulation-aware AI systems - phonetic and physiological research This dataset presents a neutral, non-linguistic, non-performative representation of plosive and consonant burst gestures.It is not designed to encode semantic speech content, but rather to isolate gesture-level acoustic primitives underlying consonant articulation. 📂 Contents Audio Files (.wav) Recorded at 96 kHz / 24-bit WAV format\ Exported as mono\ Fade-ins and fade-outs of 3--5 ms applied for consistency\ No compression, normalization, or creative processing applied\ High-pass filtered at ~60 Hz to reduce proximity effect and subsonic rumble This preview includes 3 representative audio files, selected to demonstrate: - clean pulmonic egressive plosive articulation - contrasting non-lexical consonant burst gestures - variation in burst intensity and release character Metadata (.csv) Includes structured fields for: - file name - sound source type - airflow type - phonation type - gesture and articulation descriptors - microphone and recording chain - sample rate, bit depth, and dataset version Metadata follows the Harmonic Frontier Audio -- Foundations schema and is a strict subset of the full production metadata. 🎤 Recording Notes Recorded in a treated studio environment using a single-mic setup: Microphone: RØDE NT1-A condenser microphone Recording chain: RØDE NT1-A → Zoom F8n Pro Captured at 96 kHz / 32-bit float, rendered as 96 kHz / 24-bit mono WAV for release. Natural transient dynamics were preserved to maintain articulatory realism ⚡ Usage This preview pack is designed for: Evaluation of Harmonic Frontier Audio dataset quality and structure\ Testing AI systems that model consonant articulation and onset behavior\ Research in phonetics, speech production, and expressive voice modeling\ Creative sound design involving transient vocal gestures 👉 Note: This is not a full dataset.The complete Plosives and Non-Lexical Consonant Bursts dataset includes a broader and more balanced articulatory inventory and is available for licensing. 💡 Full Dataset Availability This is a preview pack of the Plosives and Non-Lexical Consonant Bursts Dataset.The complete dataset is available for commercial licensing. For licensing inquiries:📩 info@harmonicfrontieraudio.com 📜 License Released under CC BY-NC 4.0. Free for non-commercial use, testing, and research\ Commercial licensing available via Harmonic Frontier Audio\ A formal rights declaration is included in this dataset bundle 📧 Contact Harmonic Frontier Audio📩 info@harmonicfrontieraudio.com🌐 https://harmonicfrontieraudio.com/ 🗒️ Release Notes Version 0.9 (Jan. 2026) -- Initial Preview Pack release for Plosives and Non-Lexical Consonant Bursts.See CHANGELOG.md for detailed version history. Citation If you use this dataset in your research, please cite: Pullen, B. (2026). Plosives and Non-Lexical Consonant Bursts Dataset (Preview) [Data set]. Harmonic Frontier Audio. Zenodo. https://doi.org/10.5281/zenodo.18499679 ORCID: https://orcid.org/0009-0003-4527-0178
谐波前沿音频(Harmonic Frontier Audio)——爆破音(Plosives)与非词汇辅音爆发音(Non-Lexical Consonant Bursts)(预览版v0.9) 本数据集为高保真人类语音数据集,专为AI训练、语音研究及可感知发音的语音建模而设计。 由谐波前沿音频制作的本预览版数据集,旨在展示谐波前沿音频人类语音原语(Human Vocality Primitives)系列数据集所采用的质量标准、格式规范与元数据约定。 🔎 数据集概述 本数据集提供高质量、已获得版权许可的爆破音发音与短时非词汇辅音爆发音录音——这些是通过可控声道闭合与快速释放产生的离散语音事件。 录音重点关注以下特性: - 发音的闭合与释放过程 - 瞬态气流动力学特征 - 爆发强度与包络形状 - 非语言辅音动作 这些特性使得本数据集可广泛应用于AI语音与语音建模、语音学研究、可感知发音的语音合成、起始建模以及与人类对齐的语音控制系统。 本预览版数据集遵循普洛透斯标准(The Proteus Standard™),该标准涵盖数据集溯源、透明度与伦理AI使用规范。了解更多普洛透斯标准信息:https://harmonicfrontieraudio.com/proteus-standard 完整数据集详情与许可信息请访问:https://harmonicfrontieraudio.com/datasets/plosives-non-lexical-consonant-bursts 若您认为本数据集对您有所帮助,欢迎在Hugging Face上为其点赞🤍,以帮助更多用户发现该资源。 🫁 爆破音与非词汇辅音爆发音说明 爆破音是指声道完全或近乎完全闭合后,通过可控释放气压产生的短时间高能声学爆发。非词汇辅音爆发音则指无语言意图或语义内容的同类瞬态语音动作。 这些语音行为是以下领域的基础: - 语音发音与起始建模 - 富有表现力且可控的语音合成 - 可感知发音的AI系统 - 语音学与生理学研究 本数据集呈现中立、非语言、非表演性的爆破音与辅音爆发音动作表现,其设计初衷并非编码语义化语音内容,而是分离辅音发音背后的动作级声学原语。 📂 数据集内容 #### 音频文件(.wav) 录制格式为96 kHz / 24位WAV文件,导出为单声道格式,为保证一致性添加了3~5 ms的淡入与淡出效果,未进行压缩、归一化或创意性处理,通过约60 Hz的高通滤波以降低近讲效应与亚声隆隆声。 本预览版包含3个代表性音频文件,用于展示以下内容: - 纯净的肺部呼气爆破音发音 - 对比性非词汇辅音爆发音动作 - 爆发强度与释放特性的变化 #### 元数据(.csv) 元数据包含以下结构化字段: - 文件名 - 声源类型 - 气流类型 - 发声类型 - 动作与发音描述符 - 麦克风与录制链路信息 - 采样率、位深度与数据集版本 元数据遵循谐波前沿音频——基础(Harmonic Frontier Audio -- Foundations)规范,为完整制作元数据的严格子集。 🎤 录制说明 本数据集录制于经过声学处理的录音棚,采用单麦克风录制方案: 麦克风:RØDE NT1-A电容麦克风 录制链路:RØDE NT1-A → Zoom F8n Pro 原始录制参数为96 kHz / 32位浮点,最终发布版本渲染为96 kHz / 24位单声道WAV文件。录制过程保留了自然的瞬态动力学特征,以维持发音的真实感。 ⚡ 使用场景 本预览包适用于以下场景: - 评估谐波前沿音频数据集的质量与结构 - 测试针对辅音发音与起始行为建模的AI系统 - 语音学、语音生成与富有表现力的语音建模研究 - 涉及瞬态语音动作的创意音效设计 👉 注意:本数据集并非完整版。完整的爆破音与非词汇辅音爆发音数据集包含更广泛且均衡的发音清单,可通过许可获取。 💡 完整数据集获取 本预览包为爆破音与非词汇辅音爆发音数据集的预览版本,完整数据集可通过商业许可获取。 如需咨询许可事宜,请发送邮件至:📩 info@harmonicfrontieraudio.com 📜 许可证 本数据集采用CC BY-NC 4.0协议发布: - 可免费用于非商业用途、测试与研究 - 商业许可需通过谐波前沿音频获取 - 数据集包内包含正式的权利声明文件 📧 联系方式 谐波前沿音频 📩 邮箱:info@harmonicfrontieraudio.com 🌐 官网:https://harmonicfrontieraudio.com/ 🗒️ 版本说明 版本v0.9(2026年1月)——爆破音与非词汇辅音爆发音首个预览包发布。详细版本历史请参阅CHANGELOG.md。 ### 引用格式 若您在研究中使用本数据集,请按以下格式引用: Pullen, B. (2026). 爆破音与非词汇辅音爆发音数据集(预览版)[数据集]. 谐波前沿音频. Zenodo. https://doi.org/10.5281/zenodo.18499679 ORCID:https://orcid.org/0009-0003-4527-0178



