遇见数据集

MMTAD: A Multilingual, Multi-domain Textual Attribute Dataset

收藏
Zenodo2025-07-11 更新2026-05-26 收录
官方服务:

资源简介:

MMTAD DATASET MMTAD is a dataset specifically curated for word-level textual attribute recognition in real-world documents. It comprises 1623 real-world document images—from legal notices and circulars to land records, legislative documents, textbooks and notary papers—captured under varied illumination, layouts, background textures, noise patterns, scanning artifacts and resolutions. It delivers 1,117,716 word-level annotations across eight languages (English, Hindi, Telugu, Marathi, Punjabi, Bengali, Spanish and others), labeling bold, italic, bold & italic, underline, strikeout and underline &strikeout alongside “normal” text. Overall MMTAD facilitates robust, multilingual, multi-domain detection of word-level text attributes—such as bold, italic, underline and strikeout—even under the noise and distortions typical of real-world document images. Project Page Link : TexTAR

提供机构:
Zenodo
创建时间:
2025-07-11
二维码
社区交流群
二维码
科研交流群
商业服务