Explorations of Neural Network Architectures and Training Paradigms
收藏资源简介:
This thesis explores three methods for improving neural networks, respectively making them more interpretable, efficient, and effective. ‘Regression Networks’ generalize existing methods to offer explainable machine learning, but require careful use due to subtle complexities while interpreting their predictions. ‘Iterative Permanent Dropout’ shrinks neural networks while preserving accuracy. Finally, an analysis of the neural network architecture behind systems like ChatGPT motivates a new design more suitable to advanced intelligent systems, named 'Architecture Designed for Artificial General Intelligence' (or 'ADAGI'), while a discussion of broader ethical and practical safety concerns leads to some proposed regulations designed for safer AGI development.



