MeasureBenchClock-DiverseSynthetic
收藏资源简介:
MeasureBenchClock-Diverse-Synthetic是一个用于模拟时钟读数的合成数据集,基于修改和多样化的MeasureBench时钟生成代码创建。该数据集包含5800个样本,分为训练集(5000个样本)、开发集(300个样本)和测试集(500个样本)。每个样本包含生成的模拟时钟图像、基于图像的问题、标准答案以及多种文本形式的问题和答案变体。数据集还提供了详细的时间信息(小时、分钟、秒)、视觉属性(表盘形状、主题、样式)、几何属性(时钟位置、半径、裁剪情况)和生成元数据。答案格式根据秒针可见性自动调整:秒针可见时使用H:MM:SS格式,不可见时使用H:MM格式。该数据集旨在支持模拟时钟读数任务的监督微调和强化学习实验,特别适用于合成数据训练、多模态推理和模型鲁棒性研究。由于图像是合成生成的,可能不完全匹配真实世界时钟的视觉特征,但通过多样化的生成配置(包括分辨率、布局、颜色方案等)提高了数据的多样性。
MeasureBenchClock-Diverse-Synthetic is a synthetic dataset for simulating clock reading, created based on modified and diversified MeasureBench clock generation code. It contains 5800 samples, divided into a training set (5000 samples), a development set (300 samples), and a test set (500 samples). Each sample includes generated simulated clock images, image-based questions, standard answers, and multiple textual variants of questions and answers. The dataset also provides detailed time information (hours, minutes, seconds), visual attributes (dial shape, theme, style), geometric attributes (clock position, radius, cropping status), and generation metadata. The answer format is automatically adjusted based on second-hand visibility: using H:MM:SS format when the second hand is visible, and H:MM format when it is not. This dataset aims to support supervised fine-tuning and reinforcement learning experiments for clock reading tasks, particularly suitable for synthetic data training, multimodal reasoning, and model robustness research. Since the images are synthetically generated, they may not fully match the visual characteristics of real-world clocks, but data diversity is enhanced through diverse generation configurations, including resolution, layout, color schemes, etc.





