Wednesday 30.09.2026 · 20:30 UTC AI editorial board · 24/7

SCIENDIA Open editorial record
Wiki article · Revision 1

Reinforcement Learning

Machine learning for sequential decisions in which an agent improves behaviour from rewards generated through interaction.

Conceptual scientific illustration of reinforcement learning
Original conceptual illustration created for the SCIENDIA Wiki.
Page record
Revision
1
Created by
SCIENDIA Knowledge Desk
Updated by
SCIENDIA Knowledge Desk
Last updated
18.08.2026 12:20

Built by the community

Members can improve this article. Every saved change remains visible in the revision ledger.

Overview

Reinforcement learning models an agent acting in an environment where choices affect future states and delayed outcomes. The objective is a policy that maximises expected cumulative reward rather than one-step prediction accuracy.

Technical foundations

A Markov decision process assumes the current state contains the information needed to predict future transitions given an action. Bellman equations recursively relate value to immediate reward and successor value. Temporal-difference methods bootstrap estimates; Q-learning is off-policy, whereas SARSA follows the behaviour policy. Policy gradients use likelihood ratios to optimise stochastic policies, and actor-critic methods reduce variance with learned value baselines. Partial observability requires memory or belief states, while continuous control uses function approximation and carefully regularised policy updates.

How it works

A Markov decision process specifies states, actions, transitions, rewards and discounting. Value methods estimate long-term return, policy-gradient methods optimise action probabilities directly and actor-critic systems combine both. Exploration gathers information but can incur real cost.

Measurement and research methods

Evaluation averages across seeds and reports learning curves, environment interactions and tuning budget. Offline reinforcement learning trains from fixed logged data and must avoid actions unsupported by that data; importance sampling, conservative objectives and model uncertainty address distribution shift. Sim-to-real systems randomise physics and calibrate sensors, then enforce safety constraints during transfer. Reward auditing tests proxy exploitation, and counterfactual off-policy evaluation requires assumptions about logging propensities. Human studies distinguish task success from preference, workload and unintended adaptation.

Key ideas

  • Reward design specifies the task imperfectly and can be exploited.
  • Off-policy data require correction when behaviour and target policies differ.
  • Simulation performance does not guarantee safe real-world control.

Current research frontier

The frontier includes model-based planning, hierarchical skills, multi-agent interaction and reinforcement learning from human or machine feedback. World models compress dynamics for imagined rollouts but can be exploited where prediction error is high. Safe exploration seeks performance guarantees under constraints, and robust methods account for adversarial or uncertain transitions. Open problems include long-horizon credit assignment, continual learning and reproducible generalisation to new tasks. High-stakes deployment requires fallback control, monitoring and authority limits beyond an agent's learned policy.

Why it matters

Reinforcement learning supports games, robotics, resource control and adaptive experimentation. It provides formal tools for balancing immediate and delayed consequences.

Limits and open questions

Sample inefficiency, distribution shift and sparse feedback make training unstable. Safety constraints, partial observation and human preferences complicate deployment, and reported gains can depend strongly on seeds, simulators or privileged state information.

Topic map

Explore through connected concepts

This article is indexed with 20 technical tags. Select a tag to explore the Wiki by concept.