--- pretty_name: Evaluation run of fradinho/llama-mistral dataset_summary: "Dataset automatically created during the evaluation run of model\ \ [fradinho/llama-mistral](https://huggingface.co/fradin
The Corpus of German Speech (CoGS) is a 51-million-word corpus of geolocated automatic speech recognition (ASR) YouTube transcripts from local government channels in Germany, created for the study of
该数据集包含输入文本(input_text)和目标文本(target_text)两个字段,适用于研究文本生成任务。数据集分为训练集和验证集,其中训练集有10000个样本,验证集有1024个样本。数据集用于支持论文《Roll the dice & look before you leap: Going beyond the creative limits of next-token predicti