Tejaskumar/Emergent-NCA-Sequences-5M
收藏资源简介:
Emergent NCA Sequences 5M 是一个大规模合成符号动力学数据集,包含超过500万个由随机冻结神经细胞自动机(NCA)生成的独特序列。每个序列模拟一个时空动力系统,包含500帧,网格尺寸可变(高度范围11-45,宽度范围9-46)。NCA的连续隐藏状态通过MiniBatch KMeans聚类被量化为一个全局共享的32符号词汇表,确保序列间具有结构可比性但动态独特。数据集设计用于促进序列推理、世界模型学习和长期预测任务,特别是通过稀疏长时程预测目标(如从间隔帧预测未来帧)来避免模型简单记忆。其高度多样的动态行为(包括周期性、混沌和波动模式)源于每个序列使用随机初始化的NCA权重,从而提供不受控的多样性,适用于测试模型泛化能力和规则内化能力。
Emergent NCA Sequences 5M is a large-scale synthetic symbolic dynamics dataset comprising over 5 million unique sequences generated by frozen random Neural Cellular Automata (NCA). Each sequence simulates a spatiotemporal dynamical system with 500 frames and variable grid sizes (height range 11-45, width range 9-46). The continuous hidden states of the NCA are quantized into a globally shared 32-symbol vocabulary via MiniBatch KMeans clustering, ensuring structurally comparable but dynamically unique sequences across rollouts. The dataset is designed to advance sequence reasoning, world model learning, and long-horizon prediction tasks, particularly through a sparse long-horizon prediction objective (e.g., predicting future frames from sparse context frames) to prevent simple memorization. Its highly diverse emergent behaviors (including periodic, chaotic, and wave-like patterns) stem from randomly initialized NCA weights per rollout, offering uncontrolled diversity suitable for testing model generalization and rule internalization capabilities.




