S2L
收藏资源简介:
S2L数据集是一个大规模的开源数据集,包含超过66,000个人工标注的英文和俄语文本数学方程式和句子音频样本。数据集由两部分组成:S2L-sentences和S2L-equations,分别包含大约12,000个唯一的数学句子和10.7k个独立的方程。数据集收集了来自多个来源的数学方程式和句子,以及相应的参考发音,并进行了人工和人工生成的音频标注。数据集旨在解决将口头数学表达式和句子转换为LaTeX格式的问题,适用于教育和研究领域,如讲座转录或笔记创建。
The S2L dataset is a large-scale open-source dataset containing over 66,000 manually annotated English and Russian textual mathematical equation and sentence audio samples. The dataset consists of two subsets: S2L-sentences and S2L-equations, which contain approximately 12,000 unique mathematical sentences and 10.7k standalone equations respectively. The dataset collects mathematical equations and sentences from multiple sources along with their corresponding reference pronunciations, and has been annotated with both manual and machine-generated audio annotations. It aims to address the problem of converting spoken mathematical expressions and sentences into LaTeX format, and is applicable to educational and research fields such as lecture transcription or note creation.




