AlFrauch/im2latex
收藏资源简介:
该数据集是一组图像与其对应的LaTeX表达式代码对的集合。这些对是通过分析超过100,000篇自然科学和数学文章生成的,并且已经去除了重复项。数据集中包含约1,500,000张图像。
This dataset consists of pairs of images and their corresponding LaTeX expression codes. These pairs are generated by analyzing over 100,000 natural science and mathematics articles, with duplicate entries removed. The dataset contains approximately 1,500,000 images.
数据集卡片 for Dataset Name
数据集描述
数据集摘要
该数据集包含图像及其对应的LaTeX代码表达式的配对。这些配对是通过分析超过100,000篇自然科学和数学文章并生成相应的LaTeX表达式集合而产生的。该集合已清除重复项,包含约1,500,000张图像。
支持的任务和排行榜
[更多信息需要]
语言
LaTeX
数据集结构
数据实例
[更多信息需要]
数据字段
python Dataset({ features: [image, text], num_rows: 1586584 })
数据分割
[更多信息需要]
数据集创建
策划理由
[更多信息需要]
源数据
初始数据收集和规范化
[更多信息需要]
源语言生产者
[更多信息需要]
注释
注释过程
[更多信息需要]
注释者
[更多信息需要]
个人和敏感信息
[更多信息需要]
使用数据时的考虑
数据集的社会影响
[更多信息需要]
偏见的讨论
[更多信息需要]
其他已知限制
[更多信息需要]
附加信息
数据集策展人
[更多信息需要]
许可信息
[更多信息需要]
引用信息
plaintext @misc{alexfrauch_VSU_2023, title = {Recognition of mathematical formulas in the Latex: Image-Text Pair Dataset}, author = {Aleksandr Frauch (Proshunin)}, year = {2023}, howpublished = {url{https://huggingface.co/datasets/AlFrauch/im2latex}}, }
贡献
[更多信息需要]




