相关数据集
typhoon-ai/typhoon-s-instruct-post-training
该数据集是用于Typhoon-S配方的后训练语料库,旨在构建高性能、区域和领域特定的大型语言模型(LLMs),这些模型保持本地化、可控和资源高效。它采用两部分混合策略:目标语言(泰语)对齐数据和通用英语指令+工具使用数据,以保持和增强泰语本地的指令跟随、文化/语言基础以及泰英代码转换的鲁棒性,同时传递广泛有用的助手行为。数据集包括SFT分割用于监督指令调优,以及distill分割用于策略上蒸馏(O
Hugging Face2026-01-28 更新120
typhoon-ai/thaimos-tts-annotation
--- dataset_info: features: - name: system dtype: string - name: file_id dtype: int64 - name: audio dtype: audio - name: text dtype: string - name: sound_quality dtype:
Hugging Face2025-06-13 更新80
typhoon-ai/ThaiOCRBench
--- dataset_info: features: - name: Id dtype: string - name: image dtype: image - name: Task dtype: string - name: question dtype: string - name: answer dtype: string
Hugging Face2025-12-02 更新70
typhoon-ai/thai-dialect-isan-dataset
--- language: - th license: apache-2.0 task_categories: - automatic-speech-recognition tags: - audio - speech-processing - isan-dialect pretty_name: Thai Dialect Isan Speech Corpus size_categories: -
Hugging Face2025-11-26 更新80
typhoon-ai/TVSpeech
TVSpeech是一个专门设计用于评估自动语音识别(ASR)模型在真实世界泰语音频中鲁棒性的基准数据集。该数据集包含570个从YouTube公共媒体频道精选的语音样本(总计3.75小时),涵盖了金融、技术和多样化的vlog等内容类别。数据集特别注重语音的声学和语义复杂性,包括领域特定术语、专有名词和技术行话等低频率词汇。所有样本均经过人工转录,采样率为16 kHz,并采用Creative Comm
Hugging Face2026-01-21 更新80



