遇见数据集

malaysia-ai/crawl-youtube

收藏
Hugging Face2024-04-14 更新2024-03-04 收录
官方服务:

资源简介:

该数据集包含从马来西亚和新加坡的YouTube频道爬取的音频文件,总计约60k个音频文件,总时长约185k小时。数据集的特征包括文件名和URL,并提供了数据集的下载大小和总大小。数据集的加载可以通过提供的Python代码示例高效实现。数据集的使用遵循《版权法》第107条的合理使用条款,所有视频、歌曲、图像和图形均属于其各自的所有者。

This dataset comprises audio files crawled from YouTube channels across Malaysia and Singapore, with a total of approximately 60,000 audio files and an aggregate duration of roughly 185,000 hours. Its metadata includes file names and URLs, alongside the download size and total size of the dataset. Efficient loading of the dataset can be implemented using the provided Python code snippets. Usage of this dataset adheres to the fair use guidelines stipulated in Section 107 of the Copyright Act, and all videos, songs, images, and graphics remain the property of their respective owners.

提供机构:
malaysia-ai
原始信息汇总

数据集概述

  • 数据来源:马来西亚和新加坡的YouTube频道
  • 数据类型:音频文件
  • 数据量:总计60,000个音频文件
  • 总时长:总计185,000小时
二维码
社区交流群
二维码
科研交流群
商业服务