noxwano/ASMR-Archive-Processed-mini
收藏资源简介:
--- license: agpl-3.0 task_categories: - automatic-speech-recognition - text-to-speech language: - ja tags: - speech - audio - japanese - asmr - anime - not-for-all-audiences pretty_name: ASMR-Archive-Processed-mini size_categories: - 1M<n<10M --- # ASMR-Archive-Processed-mini ## Overview This dataset is a small subset of the original [OmniAICreator/ASMR-Archive-Processed](https://huggingface.co/datasets/OmniAICreator/ASMR-Archive-Processed) dataset, created so that when you want to use only a portion of the original dataset form every subdirectory, you can simply pass this dataset name. We randomly sampled about 10% of the data from each subdirectory of the original dataset. ## Dataset Contents & Preprocessing For detailed information regarding the specific contents of the data and the original preprocessing pipelines, please refer to the original [OmniAICreator/ASMR-Archive-Processed](https://huggingface.co/datasets/OmniAICreator/ASMR-Archive-Processed) dataset. ## Biases and Limitations Users should be aware of the following limitations inherited from the original dataset: * **NSFW Content**: This dataset contains a significant amount of data derived from content originally marked as NSFW. * **Gender Bias**: Due to the nature of the source material, the dataset is heavily skewed towards female voices. * **Overlapping Speakers**: Some audio segments may contain instances where multiple speakers are talking simultaneously. * **Inclusion of Sound Effects**: While the preprocessing pipeline is designed to isolate vocals, some segments may still contain residual sound effects commonly found in ASMR content. * **Potential Transcription Errors**: Transcriptions are generated automatically by AI models and have not been manually verified. They are likely to contain errors and inaccuracies. ## License & Usage This dataset inherits the **AGPL-3.0 license** from the source datasets. **Intended Use**: This dataset is intended strictly for educational and academic research purposes. **Disclaimer**: Use is at your own risk. You must ensure compliance with applicable laws. The dataset is provided "as is" with absolutely no express or implied warranty.
许可证:AGPL-3.0 任务类别: - 自动语音识别(automatic-speech-recognition) - 文本转语音(text-to-speech) 语言: - 日语(ja) 标签: - 语音 - 音频 - 日语 - 自发性知觉经络反应(ASMR, Autonomous Sensory Meridian Response) - 动画 - 非全受众适用 展示名称:ASMR-Archive-Processed-mini 规模类别:1M<n<10M # ASMR-Archive-Processed-mini ## 概述 本数据集为原始[OmniAICreator/ASMR-Archive-Processed](https://huggingface.co/datasets/OmniAICreator/ASMR-Archive-Processed)数据集的小型子集,旨在为仅需取用原始数据集各子目录部分数据的使用者提供便捷的数据集调用方式。我们从原始数据集的每个子目录中随机抽取约10%的数据构建本数据集。 ## 数据集内容与预处理 关于数据集具体内容与原始预处理流程的详细信息,请参阅原始[OmniAICreator/ASMR-Archive-Processed](https://huggingface.co/datasets/OmniAICreator/ASMR-Archive-Processed)数据集。 ## 偏差与局限性 使用者需注意本数据集继承自原始数据集的以下局限性: * **不适宜公开内容(NSFW, Not Safe For Work)**:本数据集包含大量源自原始标记为不适宜公开内容的素材数据。 * **性别偏差**:受源素材特性影响,本数据集的语音样本高度偏向女性发声。 * **多说话人重叠**:部分音频片段可能存在多位说话人同时发声的情况。 * **残留音效**:尽管预处理流程旨在分离人声,但部分片段仍可能残留ASMR内容中常见的辅助音效。 * **转录误差风险**:转录文本由人工智能模型自动生成,未经过人工校验,大概率存在错误与偏差。 ## 许可证与使用说明 本数据集继承源数据集的**AGPL-3.0许可证**。 **预期用途**:本数据集仅可用于教育与学术研究场景。 **免责声明**:使用者需自行承担使用风险,必须确保自身行为符合当地法律法规。本数据集按“现状”提供,不附带任何明示或暗示的担保。



