dh-unibe/towerbooks-line-test-test-with-inference
收藏资源简介:
该数据集源自dh-unibe/towerbooks-line-test-test-with-inference,并已通过推理结果进行丰富。它包含729个样本,分布在1个分割(train)中,涉及项目B_IX_490_duplicated。数据集包括图像到文本的转换功能,用于手写文本识别(HTR)和TrOCR模型推理,特征包括项目名称、文件名、区域ID、行ID、图像数据、文本内容、行和区域的坐标、阅读顺序以及推理输出(如inference_20260611_163016_468653_model_dh-unibe_trocr-towerbooks)。评估结果显示字符错误率(CER)为0.1696,基于237行数据。数据以parquet分片格式组织,便于通过HuggingFace Hub加载。
This dataset is derived from dh-unibe/towerbooks-line-test-test-with-inference and has been enriched with inference results. It contains 729 samples across 1 split (train), with the included project B_IX_490_duplicated. The dataset supports image-to-text tasks for handwritten text recognition (HTR) and TrOCR model inference, featuring attributes such as project name, filename, region ID, line ID, image data, text content, line and region coordinates, reading order, and inference outputs (e.g., inference_20260611_163016_468653_model_dh-unibe_trocr-towerbooks). Evaluation results show a character error rate (CER) of 0.1696 based on 237 rows. Data is organized in parquet shards for easy loading via the HuggingFace Hub.




