遇见数据集

FlorianD/metavoice

收藏
Hugging Face2023-12-06 更新2024-03-04 收录
官方服务:

资源简介:

# Data Engineer: Take home project ## Introduction The goal of this project is to evaluate your knowledge and skills in the design and implementation of a scalable data pre-processing pipeline. ## Problem statement - Reads audio data being populated by Metavoice product, `Studio`, into a CloudFlare R2 bucket - Runs two data transformation steps on the audio files: - Transcription - use [Whisper](https://github.com/openai/whisper) - Tokenisation - use mock code [here](https://gist.github.com/sidroopdaska/364e9f493d8dd9584eb9e1e9cae5715c) - Stores the results using the example schema below. ```<id - relative path of audio file>, <transcription>, <token array>``` ## Requirements - Install `ffmpeg` by following instructions [here](https://www.hostinger.com/tutorials/how-to-install-ffmpeg) - Use pipenv to install the required packages: ```pipenv install``` - Go to where the `main.py` file is located and run: ```python main.py ``` ## Notes For scalability, I decided to read the audio file with a given chunk_size, and so preprocess the audio file in chunks. This is to avoid memory issues when dealing with large audio files. The script is broken after a while (probably an audio file it does not like) as it shows: ```pydub.exceptions.CouldntDecodeError: Decoding failed. ffmpeg returned error code: 1``` I think there is a better solution, but by lack of time and not 100% sure if that feasible, that would be to: - create a HuggingFace [loading-script](https://huggingface.co/docs/datasets/audio_dataset#loading-script) - And so we could use the HF Dataset API to load the audio files and preprocess it. - For the Whisper model, HF provide useful functions to [preprocess](https://huggingface.co/learn/audio-course/chapter1/preprocessing) it: ``` from transformers import WhisperFeatureExtractor feature_extractor = WhisperFeatureExtractor.from_pretrained("openai/whisper-small") ... ```

# 数据工程师:居家实操考核项目 ## 项目简介 本项目旨在考核您在可扩展数据预处理流水线(scalable data pre-processing pipeline)的设计与实现方面的知识与技能。 ## 问题描述 - 读取由Metavoice旗下产品`Studio`生成并上传至CloudFlare R2 存储桶(CloudFlare R2 bucket)的音频数据 - 对音频文件执行两项数据转换步骤: - 转录(Transcription):使用[Whisper](https://github.com/openai/whisper)模型 - 分词(Tokenisation):使用此处[示例模拟代码](https://gist.github.com/sidroopdaska/364e9f493d8dd9584eb9e1e9cae5715c) - 按照下述示例格式存储处理结果: <唯一标识:音频文件相对路径>, <转录结果>, <Token数组(token array)> ## 项目要求 - 按照[此处指引](https://www.hostinger.com/tutorials/how-to-install-ffmpeg)安装`ffmpeg` - 使用pipenv安装所需依赖包: pipenv install - 进入`main.py`所在目录并执行以下命令运行脚本: python main.py ## 注意事项 出于可扩展性考量,本方案采用按指定块大小(chunk_size)读取音频文件的方式,对音频文件进行分块预处理,以此规避处理大型音频文件时可能出现的内存溢出问题。 但该脚本运行一段时间后会中断(大概率是遇到了无法正常解码的音频文件),报错信息如下: pydub.exceptions.CouldntDecodeError: 解码失败,ffmpeg返回错误码:1 本人认为存在更优解决方案,但受限于时间且无法完全确认其可行性,该方案思路如下: - 创建HuggingFace [加载脚本(loading-script)](https://huggingface.co/docs/datasets/audio_dataset#loading-script) - 即可通过HF(HuggingFace)数据集API加载音频文件并完成预处理 - 针对Whisper模型,HuggingFace提供了实用的[预处理工具](https://huggingface.co/learn/audio-course/chapter1/preprocessing),示例代码如下: from transformers import WhisperFeatureExtractor feature_extractor = WhisperFeatureExtractor.from_pretrained("openai/whisper-small") ...

提供机构:
FlorianD
原始信息汇总

数据集概述

问题描述

  • 读取由Metavoice产品Studio生成的音频数据,存储到CloudFlare R2桶中。
  • 对音频文件执行两个数据转换步骤:
    • 转录(Transcription):使用Whisper
    • 标记化(Tokenisation):使用模拟代码此处
  • 使用以下示例模式存储结果: <id - 音频文件的相对路径>, <转录文本>, <标记数组>

要求

  • 安装ffmpeg,按照此处的说明进行。
  • 使用pipenv安装所需的包: pipenv install
  • 进入main.py文件所在位置并运行: python main.py

注意事项

  • 为了可扩展性,决定以给定的chunk_size读取音频文件,并以块为单位预处理音频文件,以避免处理大音频文件时的内存问题。
  • 脚本在运行一段时间后会中断(可能是因为遇到了不喜欢的音频文件),显示错误: pydub.exceptions.CouldntDecodeError: Decoding failed. ffmpeg returned error code: 1
  • 建议解决方案:
    • 创建一个HuggingFace的加载脚本
    • 使用HF Dataset API加载和预处理音频文件。
    • 对于Whisper模型,HF提供了有用的预处理函数: python from transformers import WhisperFeatureExtractor feature_extractor = WhisperFeatureExtractor.from_pretrained("openai/whisper-small") ...
二维码
社区交流群
二维码
科研交流群
商业服务