Artificial Neural Networks
Layered computational models that learn distributed representations by adjusting weighted connections from data.
- Revision
- 2
- Created by
- SCIENDIA Knowledge Desk
- Updated by
- SCIENDIA Knowledge Desk
- Last updated
- 17.08.2026 18:43
Built by the community
Members can improve this article. Every saved change remains visible in the revision ledger.
Overview
An artificial neural network maps inputs to outputs through connected processing units. Each unit combines values using adjustable weights, applies a nonlinear transformation and passes a result forward. Deep networks stack many such transformations.
Technical foundations
A neural network is a parameterised composition of affine transformations and nonlinear activation functions. Convolution imposes translation-related weight sharing, recurrent connections represent sequential state, and attention computes data-dependent weighted interactions between tokens or spatial locations. Training minimises an empirical risk, often augmented by regularisation, using gradients obtained by reverse-mode automatic differentiation. Optimisers such as stochastic gradient descent or Adam update millions to billions of parameters, while normalisation, residual connections and careful initialisation improve signal propagation through deep computational graphs.
How it works
During training, a loss function measures prediction error. Backpropagation efficiently computes how the loss changes with each parameter, and an optimiser updates the parameters. Repeated examples allow internal layers to form representations useful for classification, generation, control or estimation.
Measurement and research methods
Rigorous evaluation separates training, validation and held-out test data and prevents leakage across related samples. Classification studies report calibrated probabilities, precision-recall behaviour and subgroup performance rather than accuracy alone. Generative models require task-specific factuality, diversity and safety assessments. Ablation studies test which components are causally useful, while uncertainty can be explored with ensembles, Bayesian approximations or conformal prediction. Reproducibility depends on documenting preprocessing, random seeds, hyperparameter search, compute budget and stopping criteria; benchmark contamination can otherwise produce deceptively strong results.
Key ideas
- Architecture determines how information can flow and which patterns are easy to represent.
- Training performance and real-world generalisation are different objectives.
- Data quality, evaluation design and deployment context matter as much as model size.
Current research frontier
Research is moving toward more efficient architectures, multimodal systems and models that integrate symbolic constraints or scientific priors. Scaling laws describe average performance trends but do not guarantee reliability on a particular distribution. Mechanistic interpretability examines internal circuits, representation geometry and causal effects of activations, whereas robustness research studies adversarial perturbations and distribution shift. Deployment adds monitoring, access control, privacy and human-oversight requirements. Open problems include faithful reasoning verification, continual learning without catastrophic forgetting, energy-efficient training, and distinguishing genuine generalisation from sophisticated interpolation over memorised data.
Why it matters
Neural networks power advances in vision, speech, language, scientific modelling and decision support. Their ability to learn features from complex data reduces the need to hand-design every intermediate rule.
Limits and open questions
Networks can inherit bias, fail outside their training distribution and produce confident errors. Interpretability, energy use, robustness, privacy and reliable uncertainty estimates remain important technical and social challenges.
Explore through connected concepts
This article is indexed with 20 technical tags. Select a tag to explore the Wiki by concept.