CommonForms
收藏资源简介:
CommonForms是一个大规模、多样化的表单字段检测数据集,包含来自Common Crawl的超过59,000个文档和超过480,000个页面。该数据集通过筛选Common Crawl中的PDF文档,找出具有可填写元素的文档,并进行清洗和过滤,最终得到一个包含丰富语言和领域混合的数据集。CommonForms数据集旨在为表单字段检测提供高质量的训练数据,并支持开源模型FFDNet的训练和发布。FFDNet模型在CommonForms测试集上取得了很高的平均精度,并且能够预测文本、签名和选择按钮等表单字段的位置和类型。
CommonForms is a large-scale and diverse form field detection dataset, which contains over 59,000 documents and more than 480,000 pages sourced from Common Crawl. This dataset is constructed by screening PDF documents in Common Crawl to identify those with fillable elements, followed by cleaning and filtering processes, ultimately yielding a dataset with a rich mix of languages and domains. The CommonForms dataset aims to provide high-quality training data for form field detection tasks, and supports the training and release of the open-source model FFDNet. The FFDNet model achieves high average precision on the CommonForms test set, and is capable of predicting the positions and types of form fields such as text inputs, signature fields, and selection buttons.



