遇见数据集

Harmonic Frontier Audio - Breathy and Semi-Modal Phonation (Preview Pack, v0.9)

收藏
Zenodo2026-03-04 更新2026-05-26 收录
官方服务:

资源简介:

Harmonic Frontier Audio – Breathy & Semi-Modal Phonation (Preview, v0.9) A high-fidelity human vocal dataset designed for AI training, speech research, and expressive voice modeling. Breathy & Semi-Modal Phonation (Preview), created by Harmonic Frontier Audio, provides a compact reference set demonstrating the quality, formatting, and metadata conventions used in the Harmonic Frontier Audio Human Vocality Primitives series. 🔎 Summary This dataset provides high-quality, rights-cleared recordings of breathy and semi-modal phonation gestures — soft, partially turbulent voicing behaviors that sit between fully modal phonation and non-voiced airflow noise. The recordings emphasize: turbulent / aspirate onsets airy harmonic structure (mixed periodic + aperiodic energy) controlled transitions between breathy and more stable voicing low-intensity phonation suitable for fine-grained onset and timbre modeling These characteristics make the dataset valuable for AI speech modeling, phonetics research, voice synthesis, breath-aware phonation modeling, and human-aligned vocal control systems. Developed by Harmonic Frontier Audio, this preview follows The Proteus Standard™ for dataset provenance, transparency, and ethical AI use.Learn more about the Proteus Standard → https://harmonicfrontieraudio.com/proteus-standard Full dataset details and licensing information are available at:https://harmonicfrontieraudio.com/datasets/breathy-semi-modal-phonation If you find this dataset useful, please consider giving it a 🤍 on Hugging Face to help others discover it. 🌬️ About Breathy & Semi-Modal Phonation Breathy phonation occurs when the vocal folds vibrate with incomplete closure, allowing continuous airflow leakage and adding a strong noise component to the sound.Semi-modal phonation refers to a stable but softer voicing mode that retains harmonic structure while remaining closer to breathy or airflow-adjacent behavior than fully modal speech voice. These phenomena are foundational to: expressive voice production and timbral control onset and aspiration modeling in speech synthesis phonation-type classification and physiological modeling controllable voice systems that need interpretable “softness,” “airiness,” or “effort” dimensions This dataset presents a neutral, non-linguistic, non-performative representation of breathy and semi-modal phonation.It is not designed to encode semantic speech content, but rather to isolate acoustic primitives that underlie airy voicing and turbulent onset behavior. 📂 Contents Audio Files (.wav) Recorded at 96 kHz / 24-bit WAV format Exported as mono Fade-ins and fade-outs of 3–5 ms applied for consistency No compression, normalization, or creative processing applied High-pass filtered at ~40 Hz to remove subsonic rumble This preview includes 5 representative audio files, selected to demonstrate: soft phonation with audible airflow leakage turbulent onsets and breath-dominant attacks airy harmonic profiles with mixed periodic/aperiodic structure subtle transitions toward more stable voicing Metadata (.csv) Includes structured fields for: file name sound source type airflow type phonation type gesture and articulation descriptors microphone and recording chain sample rate, bit depth, and dataset version Metadata follows the Harmonic Frontier Audio – Foundations schema. 🎤 Recording Notes Recorded in a treated studio environment using a single-mic setup: Microphone: Rode NT1-A condenser microphone Recording chain: Rode NT1-A → Zoom F8n Pro Captured at 96 kHz / 32-bit float, rendered as 96 kHz / 24-bit mono WAV for release. Natural room tone and low-level breath noise were preserved to retain acoustic realism. ⚡ Usage This preview pack is designed for: Evaluation of Harmonic Frontier Audio dataset quality and structure Testing AI systems that model phonation type, onset behavior, and airy vocal timbre Research in phonetics, speech synthesis, and expressive vocal modeling Creative sound design involving breathy voice textures and soft vocal gestures 👉 Note: This is not a full dataset.The complete Breathy & Semi-Modal Phonation dataset includes a substantially larger and more varied set of phonation primitives and is available for licensing. 💡 Full Dataset Availability This is a preview pack of the Breathy & Semi-Modal Phonation Dataset.The complete dataset is available for commercial licensing. For licensing inquiries:📩 info@harmonicfrontieraudio.com 🔗 Explore More from Harmonic Frontier Audio Breathy & Semi-Modal Phonation (Preview) Whisper and Aspiration (Preview) Plosives and Non-Lexical Consonant Bursts (Preview) Scottish Smallpipes (Preview) Highland Bagpipes (Preview) Irish Tin Whistle in D (Preview) Subharmonic Phonation / Vocal Fry (Preview) Kalimba (Preview) Kazoo (Preview) Overtone Singing (Preview) (All datasets follow The Proteus Standard™ for ethical dataset provenance and licensing.) 📜 License Released under CC BY-NC 4.0. Free for non-commercial use, testing, and research Commercial licensing available via Harmonic Frontier Audio A formal rights declaration is included in this dataset bundle 📧 Contact Harmonic Frontier Audio📩 info@harmonicfrontieraudio.com🌐 https://harmonicfrontieraudio.com/ 🗒️ Release Notes Version 0.9 (March 2026) – Initial Preview Pack release for Breathy & Semi-Modal Phonation.See CHANGELOG.md for detailed version history.

Harmonic Frontier Audio – 气声与半模态发声(预览版v0.9) 本数据集为高保真人类语音数据集,专为人工智能训练、语音研究与表现力语音建模打造。由Harmonic Frontier Audio制作的《气声与半模态发声(预览版)》是一套精简的参考集,用于展示Harmonic Frontier Audio人类语音原语(Human Vocality Primitives)系列所采用的音质标准、格式规范与元数据(metadata)约定。 🔎 摘要 本数据集提供高质量、已完成版权清理的气声与半模态发声语音样本——这类发声属于柔和且带有部分湍流的语音行为,介于完全模态发声与无发声气流噪声之间。 本次录制重点涵盖: - 湍流/送气起始(turbulent / aspirate onsets) - 带空气感的谐波结构(混合周期性与非周期性能量) - 气声与更稳定发声间的可控过渡 - 适用于细粒度起始与音色建模的低强度发声 这些特性让本数据集在人工智能语音建模、语音学研究、语音合成、气感发声建模以及适配人类听觉感知的语音控制系统领域极具应用价值。 本预览版由Harmonic Frontier Audio开发,遵循《普罗透斯标准™》(The Proteus Standard™)以保障数据集溯源、透明度与伦理AI使用规范。了解更多《普罗透斯标准》信息,请访问:https://harmonicfrontieraudio.com/proteus-standard 完整数据集详情与授权信息请查阅:https://harmonicfrontieraudio.com/datasets/breathy-semi-modal-phonation 若您认为本数据集对您有所帮助,欢迎在Hugging Face上为其点赞🤍,以帮助更多用户发现该资源。 🌬️ 关于气声与半模态发声 气声发声指声带闭合不完全时的振动状态,此时会持续有气流泄漏,为声音引入显著的噪声成分。半模态发声则是一种稳定但更柔和的发声模式,保留谐波结构的同时,相比完全模态语音更接近气声或气流主导的发声状态。 这类发声现象是以下领域的基础: - 表现力语音生成与音色控制 - 语音合成中的起始与送气建模 - 发声类型分类与生理学建模 - 需要可解释“柔和度”“空气感”或“发声力度”维度的可控语音系统 本数据集以中立、非语言、非表演性的方式呈现气声与半模态发声,并非用于编码语义语音内容,而是用于隔离构成气感发声与湍流起始行为的声学原语(acoustic primitives)。 📂 内容 音频文件(.wav) 录制格式为96 kHz / 24位WAV,导出为单声道,添加3-5ms的淡入淡出(fade-in/fade-out)效果以保证一致性,未进行压缩、归一化或创意性处理,仅通过约40Hz的高通滤波(high-pass filter)移除亚低频隆隆声(subsonic rumble)。 本预览版包含5个代表性音频样本,用于展示以下特征: - 可闻气流泄漏的柔和发声 - 湍流起始与以气声为主的起音 - 混合周期性/非周期性结构的空气感谐波特征 - 向更稳定发声过渡的细微变化 元数据(metadata) 包含以下结构化字段: - 文件名 - 声源类型 - 气流类型 - 发声类型 - 发声动作与发音描述符 - 麦克风与录制链路信息 - 采样率、位深度与数据集版本 元数据遵循Harmonic Frontier Audio – Foundations(基础)架构规范。 🎤 录制说明 本数据集在经过声学处理的录音棚环境中录制,采用单麦克风配置: - 麦克风:Rode NT1-A 电容麦克风 - 录制链路:Rode NT1-A → Zoom F8n Pro 原始录制参数为96 kHz / 32位浮点格式,最终发布格式为96 kHz /24位单声道WAV。 保留了自然的房间环境音与低电平呼吸噪声,以保留声学真实感。 ⚡ 使用场景 本预览包适用于: - 评估Harmonic Frontier Audio数据集的音质与结构 - 测试针对发声类型、起始行为与气感语音音色建模的人工智能系统 - 语音学、语音合成与表现力语音建模领域的研究 - 涉及气声语音纹理与柔和语音动作的创意音效设计 👉 注意:本数据集并非完整版。完整的《气声与半模态发声》数据集包含数量更多、类型更丰富的发声原语(vocal primitives),可通过授权获取。 💡 完整数据集获取方式 本文件为《气声与半模态发声数据集》的预览包。完整数据集可通过商业授权获取。 授权咨询请联系:📩 info@harmonicfrontieraudio.com 🔗 探索Harmonic Frontier Audio的更多作品 - 气声与半模态发声(预览版) - 耳语与送气(预览版) - 爆破音与非词汇辅音爆发(预览版) - 苏格兰小风笛(预览版) - 高地风笛(预览版) - D调爱尔兰锡哨(预览版) - 次谐波发声/人声沙音(预览版) - 卡林巴琴(预览版) - 卡祖笛(预览版) - 泛音演唱(预览版) (所有数据集均遵循《普罗透斯标准™》,保障伦理数据集溯源与授权规范) 📜 授权协议 本数据集采用CC BY-NC 4.0协议发布。 - 可免费用于非商业用途、测试与研究 - 商业授权请通过Harmonic Frontier Audio申请 - 本数据集包中包含正式的权利声明文件 📧 联系方式 Harmonic Frontier Audio📩 info@harmonicfrontieraudio.com🌐 https://harmonicfrontieraudio.com/ 🗒️ 版本说明 版本0.9(2026年3月)——《气声与半模态发声》的首个预览包发布。详细版本历史请查阅CHANGELOG.md。

提供机构:
Zenodo
创建时间:
2026-03-04
二维码
社区交流群
二维码
科研交流群
商业服务