How agents learn to maximize cumulative reward through trial and error — model-free vs. model-based RL, and the value-based, policy-based, and actor-critic algorithm families that power modern agents.
Reinforcement learning (RL) algorithms are the decision-making frameworks that train an agent to maximize cumulative reward through trial-and-error interaction with an environment. A standard taxonomy splits them by what the agent builds: a map of its surroundings, a direct target on behavior, or a hybrid of both.
The agent repeats this loop millions of times, updating its behavior to increase long-run cumulative reward.
Every RL algorithm answers one question first: does the agent try to learn the underlying mechanics of the environment, or learn purely from direct experience?
The agent treats the environment as a black box — it acts, observes the outcome, and adapts without modeling the physics or logic behind the scenes. Highly effective for chaotic, high-dimensional settings like autonomous driving.
The agent actively learns a predictive model of the environment's transitions and rewards, then uses that "internal simulation" to plan actions before executing them — highly data-efficient but computationally intensive.
Model-free RL splits into three architectures based on what the agent prioritizes updating.
These algorithms learn a value function (or "Q-function") that maps how valuable a state or action is for long-term reward. The policy is implicit — the agent simply picks the highest-valued action available.
Instead of scoring states, these algorithms directly optimize the policy — the agent's strategy — to maximize expected return. They excel in smooth, continuous action spaces (like robotic joint control) where value functions bottleneck.
These algorithms merge value-based and policy-based methods using two networks: the Actor (controls behavior and updates the strategy) and the Critic (estimates the value function to tell the actor how good its moves actually were).
A quick reference across all ten algorithms covered above.
| Algorithm | Family | Policy type | Action space | Known for |
|---|---|---|---|---|
| Q-Learning | Value-based | Off-policy | Discrete | Tabular, foundational method |
| SARSA | Value-based | On-policy | Discrete | Conservative, safer updates |
| DQN | Value-based | Off-policy | Discrete (high-dim input) | Atari / raw-pixel control |
| REINFORCE | Policy-based | On-policy | Discrete or continuous | Monte Carlo policy gradients |
| TRPO | Policy-based | On-policy | Continuous | Strict, bounded updates |
| PPO | Policy-based | On-policy | Discrete or continuous | Modern general-purpose default |
| A2C / A3C | Actor-critic | On-policy | Discrete or continuous | Parallelized training efficiency |
| DDPG | Actor-critic | Off-policy | Continuous | Industrial robotics control |
| TD3 | Actor-critic | Off-policy | Continuous | Stabilized successor to DDPG |
| SAC | Actor-critic | Off-policy | Continuous | Max-entropy exploration |
RL's advantages have been recognized across many industries. A few concrete examples:
Robotic arms trained with RL assist with warehouse management, packaging, quality testing, and defect inspection. Self-driving cars use RL to learn how to behave across countless driving situations.
From AlphaGo — the first program to beat a human professional at Go — to modern game AI, RL has repeatedly pushed the boundary of what game-playing agents can achieve.
RL can optimize trading strategies, support portfolio management, and help balance risk against profit in ways that adapt as market conditions change.
RL helps personalize treatment plans and plays a growing role in drug discovery and testing, helping the field move faster from research to patient care.