masscount-cf
收藏资源简介:
MassCount-CF是一个用于识别视觉语言模型中加性神经质量基数坐标的合成反事实计数语料库,作为论文《Counting Requires Mass: An Algebraic and Causal Account of Numerosity in Vision-Language Models》的配套数据集。该数据集完全通过合成生成,不包含真实图像或受版权保护的材料。数据规模包含99,989个主场景、326,623张图像和28,617,018个对象实例。数据结构以场景图为核心持久化构件,包含场景描述符(如凸包面积、覆盖率、最近邻统计、Ripleys K)和对象标准化坐标的parquet文件,图像可通过生成器按需重新生成任何分辨率。数据集包含12个场景家族,包括chain(44,400场景)、nuisance(19,915场景)、union_part(7,500场景)等。数据划分遵循FSC-147的类别不相交协议,按parent_chain_id分组分为测试集(62,499场景)、训练集(31,154场景)、验证集(6,332场景)和参考集(4场景)。数据集通过了完整性验证,包括场景ID唯一性、内容哈希唯一性、跨划分重复检查等。适用于计数、视觉语言模型、对象计数、数量感知等研究任务。
MassCount-CF is a synthetic counterfactual counting corpus for identifying additive neural mass cardinality coordinates in vision-language models, and it is the supporting dataset for the paper *Counting Requires Mass: An Algebraic and Causal Account of Numerosity in Vision-Language Models*. This dataset is entirely synthetic, containing no real images or copyright-protected materials. The dataset comprises 99,989 main scenes, 326,623 images, and 28,617,018 object instances. The core persistent building block of the data structure is the scene graph, including parquet files containing scene descriptors (such as convex hull area, coverage rate, nearest neighbor statistics, Ripley's K) and standardized object coordinates. Images can be regenerated at any resolution on demand via the generator. The dataset includes 12 scene families, such as chain (44,400 scenes), nuisance (19,915 scenes), union_part (7,500 scenes), etc. The data partitioning follows the class-disjoint protocol of FSC-147, and is grouped by parent_chain_id into the test set (62,499 scenes), training set (31,154 scenes), validation set (6,332 scenes), and reference set (4 scenes). The dataset has passed integrity verification, including checks for scene ID uniqueness, content hash uniqueness, cross-partition duplicate detection, and other items. It is applicable to research tasks such as counting, vision-language models, object counting, and numerosity awareness.





