遇见数据集

I2L-140K

收藏
OpenDataLab2026-07-12 更新2024-05-09 收录
官方服务:

资源简介:

由 Singh、Sumeet S. 介绍。“教学机器编码:具有视觉注意的神经标记生成”。 ArXiv abs/1802.05415 (2018): n。页。用于 OpenAI 的 image-2-latex 系统任务的预构建数据集。包括总共约 140k 的公式和图像,分为训练集、验证集和测试集。 im2latex-100K 数据集的超集。

Introduced by Sumeet S. Singh, the paper *Teaching Machines to Code: Neural Token Generation with Visual Attention* was published as an arXiv preprint arXiv:1802.05415 (2018), n. pag. This is a pre-built dataset for the OpenAI image-2-latex system task, containing approximately 140,000 total formula-image pairs split into training, validation, and test sets. It is a superset of the im2latex-100K dataset.

提供机构:
OpenDataLab
创建时间:
2022-05-23
搜集汇总
数据集介绍
I2L-140K 数据集图片
背景与挑战
背景概述
I2L-140K是一个用于图像到Latex公式生成任务的公开数据集,包含约14万条公式和图像数据,分为训练、验证和测试集。该数据集基于Singh等人的研究,是im2latex-100K数据集的扩展版本,发布于2018年。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务