Optimal Transport
A mathematical framework for transforming one distribution into another while minimising a defined movement cost.
- Revision
- 1
- Created by
- SCIENDIA Knowledge Desk
- Updated by
- SCIENDIA Knowledge Desk
- Last updated
- 18.08.2026 12:20
Built by the community
Members can improve this article. Every saved change remains visible in the revision ledger.
Overview
Optimal transport formalises how mass, probability or resources should be rearranged between distributions. The chosen cost encodes geometry and application priorities, producing distances and maps that account for where probability lies rather than comparing bins independently.
Technical foundations
The Monge problem seeks a map pushing a source measure to a target while minimising integrated cost, but maps may not exist for discrete or splitting allocations. Kantorovich relaxes the problem to a joint measure whose marginals equal source and target. For metric costs, the minimum defines Wasserstein distances with meaningful geometry. Duality replaces couplings with potential functions constrained by the cost, yielding theoretical certificates and algorithms. Dynamic formulations describe transport as a density flow satisfying continuity while minimising kinetic action.
How it works
A transport plan assigns how much mass moves from each source location to each target. The Kantorovich formulation optimises over couplings with fixed marginals; under suitable conditions a deterministic Monge map exists. Dual potentials and regularisation make large problems computationally tractable.
Measurement and research methods
Linear programming solves small discrete instances, while Sinkhorn iterations add entropy and alternate scalable matrix normalisations. Regularisation strength trades computational speed and smoothness against bias. Sliced Wasserstein methods project to one dimension, and multiscale solvers exploit geometry. Evaluation tests marginal error, objective gap, runtime and sensitivity to sample size. In applications, preprocessing and ground metric require justification: arbitrary feature scaling changes transport paths. Unbalanced formulations permit creation or destruction of mass when totals or detection rates differ.
Key ideas
- The solution depends on the ground cost, not only the two distributions.
- A coupling represents joint allocation and need not be a one-to-one map.
- Regularisation improves computation while changing the exact optimisation problem.
Current research frontier
Research connects transport with generative flows, domain adaptation, inverse problems and distributionally robust optimisation. Barycentres summarise several distributions while retaining geometry, and Gromov-Wasserstein methods compare relational structures without shared coordinates. Statistical work addresses high-dimensional sample complexity through structural assumptions and regularisation. Open challenges include causal constraints, fairness and interpretable costs in social applications. Efficient differentiable solvers make transport a component of neural systems, but numerical gradients and regularisation can obscure whether a learned model still solves the intended mathematical problem.
Why it matters
Optimal transport supports imaging, economics, climate analysis, generative modelling and comparison of structured data. Wasserstein geometry also provides a language for flows of probability distributions.
Limits and open questions
High-dimensional sample complexity can be severe, and learned costs may encode bias. Entropic approximations blur fine structure, while causal or capacity constraints require extensions beyond unconstrained mass movement.
Explore through connected concepts
This article is indexed with 20 technical tags. Select a tag to explore the Wiki by concept.