This book presents deep learning as both a mathematical discipline and an applied field, connecting theoretical principles with the modeling and design choices that shape modern neural systems. It explores how neural networks are formulated and how they learn, from multilayer perceptrons, activation and loss functions, optimization, and backpropagation to questions of expressivity and generalization. These ideas are applied to neural architectures for different forms of data, including convolutional, recurrent, and graph neural networks, with attention to the mathematical principles, training challenges, and design choices that distinguish them. The book then turns to generative models and sequential decision-making, covering variational and adversarial models, diffusion, and reinforcement learning. The later parts follow the evolution of neural networks toward increasingly general systems, exploring attention and transformer architectures, multimodal models, large language models, retrieval-augmented systems, and autonomous agents. The book also considers directions beyond conventional deep learning, including continuous-time and physics informed neural networks, biologically inspired and neuromorphic models, Bayesian approaches to uncertainty, and quantum neural networks. The term “theoretical” in the title, Neural Networks: A Concise Theoretical Foundation, refers to the mathematical formulation of neural networks, not to a comprehensive, proof-oriented treatment of neural network theory. Organized in four parts, the book connects mathematical descriptions with the architectures, learning methods, and emerging directions that continue to shape neural computation.
...
Read Full Text