Transformer Architectures
Neural-network systems that model relationships among tokens using attention and parallel sequence processing.
- Revision
- 1
- Created by
- SCIENDIA Knowledge Desk
- Updated by
- SCIENDIA Knowledge Desk
- Last updated
- 18.08.2026 11:45
Built by the community
Members can improve this article. Every saved change remains visible in the revision ledger.
Overview
Transformers represent sequences or sets by repeatedly mixing information through attention and transforming each position through learned nonlinear layers. Unlike recurrent models, they can process many positions in parallel during training and directly connect distant elements, making them central to language, vision, audio and multimodal systems.
Technical foundations
Scaled dot-product attention forms scores from query-key inner products, applies a normalised weighting and mixes value vectors. Multi-head attention learns several subspaces, while causal masks prevent a decoder from observing future positions. Positional encodings may be absolute, relative or rotary. Residual streams carry information across layers and feed-forward blocks apply position-wise nonlinear transformations. Training typically minimises next-token or masked-token loss, creating representations that support in-context adaptation. Dense self-attention has quadratic sequence cost, motivating sparse, linear, recurrent-memory and state-space hybrids for long inputs.
How it works
Queries, keys and values are projected from hidden representations. Similarity scores determine weighted combinations of values for each attention head, while residual connections, normalisation and feed-forward blocks refine the result. Positional information supplies order. Decoder models predict future tokens; encoder and encoder-decoder variants support representation and conditional generation tasks.
Measurement and research methods
Evaluation separates pretraining, adaptation and held-out tasks and audits contamination at document and semantic levels. Language systems require calibration, factuality, robustness and subgroup testing beyond average benchmark score. Attention visualisation can describe routing but does not prove why an output occurred; causal interventions on activations and ablation provide stronger evidence. Efficient inference uses key-value caching, quantisation, batching and speculative decoding, each with accuracy-latency trade-offs. Reproducibility requires model architecture, data provenance, tokenisation, optimisation schedule, random seeds and compute budget rather than parameter count alone.
Key ideas
- Attention weights are learned communication patterns, not guaranteed causal explanations.
- Context length, memory use and compute scale roughly with architectural and implementation choices.
- Training objective and data composition strongly shape capabilities and failure modes.
Current research frontier
The frontier includes mixture-of-experts routing, retrieval augmentation, multimodal token spaces and tool-using agents. Long-context research tests whether models retrieve and integrate dispersed evidence rather than merely accepting large inputs. Alignment methods optimise preferences or verifiable rewards, but reward models can be exploited and may not generalise. Mechanistic studies investigate circuits for induction, copying and algorithmic operations. Open challenges include reducing hallucination, protecting private training data, evaluating autonomous behaviour and measuring environmental cost. Secure deployment adds access control, sandboxing, provenance and continuous monitoring to the model itself.
Why it matters
Transformers provide a shared computational foundation for translation, retrieval, scientific modelling, coding and generative interfaces. Their modularity supports transfer learning and adaptation across tasks with limited labelled data.
Limits and open questions
Large models can memorise sensitive or copyrighted material, reproduce bias and generate confident errors. Long-context performance may degrade despite nominal capacity, and evaluation can be distorted by benchmark contamination, hidden tool use and uneven subgroup coverage.
Explore through connected concepts
This article is indexed with 20 technical tags. Select a tag to explore the Wiki by concept.