DDPG on Reacher: Twenty Arms Chasing a Balloon

DDPG completes the Unity Reacher task. Twenty robot arms use one shared replay buffer. This post shows the actor-critic setup and why the training was easier than Tennis.

On this page

Reacher was the easiest of the three environments. It also gave the most important lesson, because its result was the opposite of the result on Tennis one year before.

The environment has twenty robot arms. Each arm has two joints, and each arm has a target balloon that moves around it. An arm gets a reward for each step that its hand is in the target. The task is to keep the hand on the balloon while the balloon moves.

The code is in the repo.

The environment

Each arm gets 33 numbers: position, rotation, velocity and angular velocity. Each arm controls 4 continuous values, which are the torques on its two joints. All values are continuous, thus DDPG is the correct method, not DQN.

The task is complete when the average score of all twenty arms is +30 or more for 100 episodes in sequence.

The twenty arms are the important part. They do not compete and they do not cooperate. They are twenty copies of the same problem that run at the same time. All of them send their experience to one shared replay buffer. This gives two advantages:

  • The agent gets twenty times more experience in each second.
  • The sampled batches have more variety, because they come from twenty arms in twenty different conditions.

This shared buffer is a large part of the result.

The model

The model has the same shape as the Tennis agents. The actor gets a state and gives an action directly. The critic gives a value for the state and the action. Each network has six fully-connected layers:

  • Actor: 400 → 300 → 200 → 100 → 50 units with ReLU, and a tanh output. The tanh function keeps the torques in their limits.
  • Critic: The state goes through 400 units. The action joins at the second layer. Then the layers become smaller and give one value.

The settings are:

  • Replay buffer: 1e6
  • Batch size: 1024
  • Gamma: 0.99
  • Tau: 1e-3
  • Actor LR: 1e-4, critic LR: 1e-3
  • Weight decay: 0, optimizer AdamW
  • Exploration noise: Ornstein-Uhlenbeck

The exploration noise

The first version had smaller networks with three layers, and it used Gaussian exploration noise. After 2,000 episodes, it did not get to the target. Deeper networks made the result a little better. But after 1,000 more episodes, the task was still not complete.

An Ornstein-Uhlenbeck (OU) process for the noise completed the task in 101 episodes.

On Tennis, this change had the opposite result. OU noise prevented the convergence, and Gaussian noise gave a solution. The reason is this: OU noise has a correlation between timesteps. It pushes the agent to continue in the same direction. A robot arm that moves to a target needs this behavior, because a smooth and continuous movement is the correct behavior. An agent that must react to a tennis ball does not need it.

There is no general rule here. Try the two types of noise. Also, the default values in a paper are correct for the environment in that paper. They are not always correct for a different environment.

Results

cd scripts
python train_agent.py --unity-app Reacher.app --target-score 30
Episode 95  Average Score: 29.26
Episode 96  Average Score: 29.35
Episode 97  Average Score: 29.43
Episode 98  Average Score: 29.52
Episode 99  Average Score: 29.61
Episode 100 Average Score: 29.69
Episode 101 Average Score: 30.08

Environment solved in 101 episodes! Average Score: 30.08

The average score for each episode. The score increases quickly in the first 30 episodes, then stays a little above 30.

The agent completed the task in 101 episodes. This curve is very different from the other two curves. It increases quickly in the first thirty episodes, then it stays near the maximum. After that point, the agent is already good. The 100-episode average needs more time to show it.

Reacher gives a reward at almost each step. Thus, there is no long period of zero rewards, as in Tennis.

This video compares the untrained arms and the trained arms. The untrained arms move randomly. The trained arms find their targets and follow them:

Twenty untrained arms, then twenty trained arms that follow their targets.

The next method to try is D4PG. The ddpg and mlagents scripts come from the Udacity deep-reinforcement-learning repo (MIT).

Written by Sina Fathi-Kazerooni

Lead AI Engineer. I write about reinforcement learning, transformers, multi-agent systems, and the tools I build along the way.