Reinforcement Learning
Machine learning for sequential decisions in which an agent improves behaviour from rewards generated through interaction.
- Revision
- 1
- Created by
- SCIENDIA Knowledge Desk
- Updated by
- SCIENDIA Knowledge Desk
- Last updated
- 18.08.2026 12:20
Built by the community
Members can improve this article. Every saved change remains visible in the revision ledger.
Overview
Reinforcement learning models an agent acting in an environment where choices affect future states and delayed outcomes. The objective is a policy that maximises expected cumulative reward rather than one-step prediction accuracy.
Technical foundations
A Markov decision process assumes the current state contains the information needed to predict future transitions given an action. Bellman equations recursively relate value to immediate reward and successor value. Temporal-difference methods bootstrap estimates; Q-learning is off-policy, whereas SARSA follows the behaviour policy. Policy gradients use likelihood ratios to optimise stochastic policies, and actor-critic methods reduce variance with learned value baselines. Partial observability requires memory or belief states, while continuous control uses function approximation and carefully regularised policy updates.
How it works
A Markov decision process specifies states, actions, transitions, rewards and discounting. Value methods estimate long-term return, policy-gradient methods optimise action probabilities directly and actor-critic systems combine both. Exploration gathers information but can incur real cost.
Measurement and research methods
Evaluation averages across seeds and reports learning curves, environment interactions and tuning budget. Offline reinforcement learning trains from fixed logged data and must avoid actions unsupported by that data; importance sampling, conservative objectives and model uncertainty address distribution shift. Sim-to-real systems randomise physics and calibrate sensors, then enforce safety constraints during transfer. Reward auditing tests proxy exploitation, and counterfactual off-policy evaluation requires assumptions about logging propensities. Human studies distinguish task success from preference, workload and unintended adaptation.
Key ideas
- Reward design specifies the task imperfectly and can be exploited.
- Off-policy data require correction when behaviour and target policies differ.
- Simulation performance does not guarantee safe real-world control.
Current research frontier
The frontier includes model-based planning, hierarchical skills, multi-agent interaction and reinforcement learning from human or machine feedback. World models compress dynamics for imagined rollouts but can be exploited where prediction error is high. Safe exploration seeks performance guarantees under constraints, and robust methods account for adversarial or uncertain transitions. Open problems include long-horizon credit assignment, continual learning and reproducible generalisation to new tasks. High-stakes deployment requires fallback control, monitoring and authority limits beyond an agent's learned policy.
Why it matters
Reinforcement learning supports games, robotics, resource control and adaptive experimentation. It provides formal tools for balancing immediate and delayed consequences.
Limits and open questions
Sample inefficiency, distribution shift and sparse feedback make training unstable. Safety constraints, partial observation and human preferences complicate deployment, and reported gains can depend strongly on seeds, simulators or privileged state information.
Explore through connected concepts
This article is indexed with 20 technical tags. Select a tag to explore the Wiki by concept.