Single-Shot Optical Neural Network
收藏资源简介:
Data from classification of MNIST [1], Fashion-MNIST [2] and QuickDraw [3] images reported in "Single-Shot Optical Neural Network" by L. Bernstein et al. The networks of 784 -> N (-> N) -> 10 activations performed inference on the test sets through consecutive matrix products implemented on the optical hardware, with ReLU applied electronically between each layer (see main text for more details). Folders for each tested network contain the following text files: Inputs: 2D matrices of size B x 784 containing training (B = 50,000 or 100,000), validation (B = 10,000) and test (B = 10,000) sets. Each row is an input vector that can be reshaped into an input image of size 28 x 28. True labels: B-length vectors containing the true label of each input in the training, validation and test sets. Neural network weights: 2D matrices of size K x N used in inference experiments to classify the test sets. Weights were pre-trained on the training set using a digital electronic computer as described in Materials and Methods. Weight values were normalized such that all values fall between -255 and 255. In the optical neural network, the weighting SLM displays the absolute values of the weights (rounded to the nearest integer), and the negative weight signs are applied in post-processing. The 32-bit weight values were used for inference performed on the digital electronic computer for the ground truth comparison. Outputs (normalized): 2D output matrices of size 10,000 x 10 from the networks processed on a digital electronic computer (ground truth) and the optical neural network. Each row is an output vector where the position of the maximum value indicates the predicted label of the input in the same row of the test set. Predicted labels: 10,000-element vectors that represent the labels predicted by the networks processed on a digital electronic computer (ground truth) and the optical neural network. These predicted labels were used to generate the confusion matrices and calculate the classification accuracies (versus the true labels). The classes for the Fashion-MNIST dataset are the following: 0: T-shirt 1: Trouser 2: Pullover 3: Dress 4: Coat 5: Sandal 6: Shirt 7: Sneaker 8: Bag 9: Ankle boot And the randomly selected classes for QuickDraw are: 0: Hourglass 1: Saw 2: Golf club 3: See saw 4: Spoon 5: Horse 6: Onion 7: Light bulb 8: Harp 9: Flip flops [1] Y. LeCun, L. Bottou, Y. Bengio, P. Haffner, Gradient-based learning applied to document recognition. Proceedings of the IEEE 86, 2278–2324 (1998). [2] H. Xiao, K. Rasul, R. Vollgraf, Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms. Preprint at https://arxiv.org/abs/1708.07747 (2017). [3] J. Jongejan, H. Rowley, T. Kawashima, J. Kim, N. Fox-Gieg, The Quick, Draw! AI experiment, https://quickdraw.withgoogle.com/ (2016).
本数据集所用图像分类数据取自MNIST[1]、Fashion-MNIST[2]与QuickDraw[3]三类图像,相关研究成果刊载于L. Bernstein等人发表的《Single-Shot Optical Neural Network》一文。所采用的神经网络结构为784 → N(→ N)→ 10个激活单元,该网络通过光学硬件实现的连续矩阵乘法在测试集上完成推理,每一层之间通过电子方式施加修正线性单元(Rectified Linear Unit,ReLU)激活函数(详见正文部分)。 每个待测网络对应的文件夹中包含以下文本文件: Inputs:尺寸为B×784的二维矩阵,涵盖训练集(B=50,000或100,000)、验证集(B=10,000)与测试集(B=10,000)。每一行均为一个输入向量,可被重塑为尺寸28×28的输入图像。 True labels:长度为B的向量,包含训练集、验证集与测试集中每个输入的真实标签。 Neural network weights:推理实验中用于对测试集进行分类的、尺寸为K×N的二维矩阵。权重预先在训练集上完成预训练,预训练过程采用数字电子计算机完成,详见《材料与方法》部分。权重值已进行归一化处理,使所有取值介于-255至255之间。在光学神经网络中,加权空间光调制器(Spatial Light Modulator,SLM)会显示权重的绝对值(四舍五入至最接近的整数),负权重的符号则在后处理阶段施加。用于真值对比的数字电子计算机推理过程采用32位权重值。 Outputs (normalized):分别来自数字电子计算机(真值基准)与光学神经网络处理得到的、尺寸为10,000×10的二维输出矩阵。每一行均为一个输出向量,其中最大值所在的位置对应测试集同一行输入的预测标签。 Predicted labels:长度为10,000的向量,分别代表数字电子计算机(真值基准)与光学神经网络处理得到的预测标签。这些预测标签被用于生成混淆矩阵并计算分类准确率(与真实标签对比)。 Fashion-MNIST数据集的类别如下: 0: T恤(T-shirt) 1: 长裤(Trouser) 2: 套头衫(Pullover) 3: 连衣裙(Dress) 4: 外套(Coat) 5: 凉鞋(Sandal) 6: 衬衫(Shirt) 7: 运动鞋(Sneaker) 8: 手提包(Bag) 9: 踝靴(Ankle boot) 本次实验为QuickDraw数据集选取的随机类别如下: 0: 沙漏(Hourglass) 1: 锯子(Saw) 2: 高尔夫球杆(Golf club) 3: 跷跷板(See saw) 4: 勺子(Spoon) 5: 马(Horse) 6: 洋葱(Onion) 7: 灯泡(Light bulb) 8: 竖琴(Harp) 9: 人字拖(Flip flops) [1] Y. LeCun, L. Bottou, Y. Bengio, P. Haffner. 基于梯度的文档识别学习方法. 《IEEE汇刊》86卷,2278–2324页(1998年)。 [2] H. Xiao, K. Rasul, R. Vollgraf. Fashion-MNIST:一种用于机器学习算法基准测试的新型图像数据集. 预印本:https://arxiv.org/abs/1708.07747(2017年)。 [3] J. Jongejan, H. Rowley, T. Kawashima, J. Kim, N. Fox-Gieg. 《The Quick, Draw!》人工智能实验. https://quickdraw.withgoogle.com/(2016年)。



