AI Infra Notes logo — agent-environment feedback loop AI Infra Notes

Reinforcement Learning Algorithms Explained

How agents learn to maximize cumulative reward through trial and error — model-free vs. model-based RL, and the value-based, policy-based, and actor-critic algorithm families that power modern agents.

2Core Paradigms
10+Algorithms Covered
4Real-World Domains

01 · FoundationsWhat Reinforcement Learning Algorithms Do

Reinforcement learning (RL) algorithms are the decision-making frameworks that train an agent to maximize cumulative reward through trial-and-error interaction with an environment. A standard taxonomy splits them by what the agent builds: a map of its surroundings, a direct target on behavior, or a hybrid of both.

Agent
chooses an action
→
Environment
transitions to a new state
→
Reward + State
fed back to the agent

The agent repeats this loop millions of times, updating its behavior to increase long-run cumulative reward.

02 · The Fundamental SplitModel-Free vs. Model-Based RL

Every RL algorithm answers one question first: does the agent try to learn the underlying mechanics of the environment, or learn purely from direct experience?

Model-Free RL

The agent treats the environment as a black box — it acts, observes the outcome, and adapts without modeling the physics or logic behind the scenes. Highly effective for chaotic, high-dimensional settings like autonomous driving.

This page's focus. Splits into value-based, policy-based, and actor-critic families below.

Model-Based RL

The agent actively learns a predictive model of the environment's transitions and rewards, then uses that "internal simulation" to plan actions before executing them — highly data-efficient but computationally intensive.

Examples: MuZero (plans by learning a model of the environment's dynamics rather than being given the rules, used to master Go, chess, and Atari) and DreamerV3 (trains actor-critic policies inside a learned "world model" via imagined rollouts, generalizing across many task domains without per-task tuning).

03 · Model-Free FamiliesThree Ways Model-Free Agents Learn

Model-free RL splits into three architectures based on what the agent prioritizes updating.

01

Value-Based Algorithms

These algorithms learn a value function (or "Q-function") that maps how valuable a state or action is for long-term reward. The policy is implicit — the agent simply picks the highest-valued action available.

Q-Learning
An off-policy, tabular method that uses the Bellman equation to iteratively find the best future actions regardless of the agent's current strategy. Best suited to simple, discrete environments with a manageable state-action table.
SARSA (State-Action-Reward-State-Action)
An on-policy alternative that updates its values using the action it is actually about to take next, rather than assuming a perfectly greedy choice — generally more conservative and stable during training.
DQN (Deep Q-Network)
Combines Q-learning with deep neural networks to handle high-dimensional state spaces, such as raw pixels in Atari games, using techniques like experience replay and target networks for training stability.
02

Policy-Based Algorithms

Instead of scoring states, these algorithms directly optimize the policy — the agent's strategy — to maximize expected return. They excel in smooth, continuous action spaces (like robotic joint control) where value functions bottleneck.

REINFORCE
A foundational Monte Carlo policy-gradient method that raises or lowers action probabilities based on whether a full finished episode yielded a positive return.
TRPO (Trust Region Policy Optimization)
An earlier, mathematically rigorous predecessor to PPO that keeps each policy update strictly bounded within a designated "trust region" to guarantee stable improvement.
PPO (Proximal Policy Optimization)
The de facto modern default (notably at OpenAI). It uses a clipped objective function to prevent destructively large policy updates while remaining far simpler to implement and tune than TRPO.
03

Actor-Critic Algorithms

These algorithms merge value-based and policy-based methods using two networks: the Actor (controls behavior and updates the strategy) and the Critic (estimates the value function to tell the actor how good its moves actually were).

A2C / A3C (Advantage Actor-Critic)
Multiple agents interact with multiple environment copies synchronously (A2C) or asynchronously (A3C) to decorrelate experience and accelerate training.
DDPG (Deep Deterministic Policy Gradient)
Extends actor-critic methods to continuous action spaces with a deterministic policy, making it a long-standing choice for high-precision industrial robotics.
TD3 (Twin-Delayed DDPG)
Stabilizes DDPG with three fixes: clipped double-Q learning (two critics, take the smaller estimate) to curb value overestimation, delayed policy updates, and target policy smoothing — now generally preferred over plain DDPG.
SAC (Soft Actor-Critic)
An off-policy method that adds an entropy bonus to the reward, encouraging exploration so the agent doesn't prematurely settle into a suboptimal strategy.

04 · At a GlanceAlgorithm Comparison

A quick reference across all ten algorithms covered above.

AlgorithmFamilyPolicy typeAction spaceKnown for
Q-LearningValue-basedOff-policyDiscreteTabular, foundational method
SARSAValue-basedOn-policyDiscreteConservative, safer updates
DQNValue-basedOff-policyDiscrete (high-dim input)Atari / raw-pixel control
REINFORCEPolicy-basedOn-policyDiscrete or continuousMonte Carlo policy gradients
TRPOPolicy-basedOn-policyContinuousStrict, bounded updates
PPOPolicy-basedOn-policyDiscrete or continuousModern general-purpose default
A2C / A3CActor-criticOn-policyDiscrete or continuousParallelized training efficiency
DDPGActor-criticOff-policyContinuousIndustrial robotics control
TD3Actor-criticOff-policyContinuousStabilized successor to DDPG
SACActor-criticOff-policyContinuousMax-entropy exploration

05 · In PracticeReal-World Applications of Reinforcement Learning

RL's advantages have been recognized across many industries. A few concrete examples:

Robotics & Automation

Robotic arms trained with RL assist with warehouse management, packaging, quality testing, and defect inspection. Self-driving cars use RL to learn how to behave across countless driving situations.

Gaming & Entertainment

From AlphaGo — the first program to beat a human professional at Go — to modern game AI, RL has repeatedly pushed the boundary of what game-playing agents can achieve.

Finance & Trading

RL can optimize trading strategies, support portfolio management, and help balance risk against profit in ways that adapt as market conditions change.

Healthcare & Medicine

RL helps personalize treatment plans and plays a growing role in drug discovery and testing, helping the field move faster from research to patient care.

06 · Quick GuideChoosing an Algorithm Family

Simple, discrete environment?
Start with Q-Learning or SARSA — tabular value-based methods are easy to reason about and debug.
High-dimensional input (pixels)?
DQN and its successors bring deep networks to value-based learning for complex perception.
Continuous control (robotics)?
Reach for DDPG, TD3, or SAC — off-policy actor-critic methods built for continuous action spaces.
Need a stable general-purpose default?
PPO is the most common starting point in modern RL, balancing simplicity, stability, and performance.
Sample efficiency is critical?
Model-based methods (MuZero, DreamerV3) learn more from fewer environment interactions, at higher compute cost.
Want faster wall-clock training?
A2C/A3C parallelize environment interaction across workers to speed up on-policy learning.