yaya-sy/bimodal-ift
收藏官方服务:
资源简介:
这是一个用于语音到文本和文本到语音转换的指令数据集。数据集中的语音数据使用了SpeechTokenize方法进行标记化。该数据集适用于任何大型语言模型(LLM)的标准微调。
An instruction dataset for speech->text and text->speech. The speech data in this dataset is tokenized using the SpeechTokenize approach. This dataset is suitable for standard finetuning with any large language model (LLM).
提供机构:
yaya-sy


