I2L-140K
收藏官方服务:
资源简介:
由 Singh、Sumeet S. 介绍。“教学机器编码:具有视觉注意的神经标记生成”。 ArXiv abs/1802.05415 (2018): n。页。用于 OpenAI 的 image-2-latex 系统任务的预构建数据集。包括总共约 140k 的公式和图像,分为训练集、验证集和测试集。 im2latex-100K 数据集的超集。
Introduced by Sumeet S. Singh, the paper *Teaching Machines to Code: Neural Token Generation with Visual Attention* was published as an arXiv preprint arXiv:1802.05415 (2018), n. pag. This is a pre-built dataset for the OpenAI image-2-latex system task, containing approximately 140,000 total formula-image pairs split into training, validation, and test sets. It is a superset of the im2latex-100K dataset.
提供机构:
OpenDataLab创建时间:
2022-05-23
搜集汇总
数据集介绍

背景与挑战
背景概述
I2L-140K是一个用于图像到Latex公式生成任务的公开数据集,包含约14万条公式和图像数据,分为训练、验证和测试集。该数据集基于Singh等人的研究,是im2latex-100K数据集的扩展版本,发布于2018年。
以上内容由遇见数据集搜集并总结生成



