DDPG on Reacher: Twenty Arms Chasing a Balloon
DDPG completes the Unity Reacher task. Twenty robot arms use one shared replay buffer. This post shows the actor-critic setup and why the training was easier than Tennis.
On this page
Reacher was the easiest of the three environments. It also gave the most important lesson, because its result was the opposite of the result on Tennis one year before.
The environment has twenty robot arms. Each arm has two joints, and each arm has a target balloon that moves around it. An arm gets a reward for each step that its hand is in the target. The task is to keep the hand on the balloon while the balloon moves.
The code is in the repo.
The environment¶
Each arm gets 33 numbers: position, rotation, velocity and angular velocity. Each arm controls 4 continuous values, which are the torques on its two joints. All values are continuous, thus DDPG is the correct method, not DQN.
The task is complete when the average score of all twenty arms is +30 or more for 100 episodes in sequence.
The twenty arms are the important part. They do not compete and they do not cooperate. They are twenty copies of the same problem that run at the same time. All of them send their experience to one shared replay buffer. This gives two advantages:
- The agent gets twenty times more experience in each second.
- The sampled batches have more variety, because they come from twenty arms in twenty different conditions.
This shared buffer is a large part of the result.
The model¶
The model has the same shape as the Tennis agents. The actor gets a state and gives an action directly. The critic gives a value for the state and the action. Each network has six fully-connected layers:
- Actor: 400 → 300 → 200 → 100 → 50 units with ReLU, and a
tanhoutput. Thetanhfunction keeps the torques in their limits. - Critic: The state goes through 400 units. The action joins at the second layer. Then the layers become smaller and give one value.
The settings are:
- Replay buffer: 1e6
- Batch size: 1024
- Gamma: 0.99
- Tau: 1e-3
- Actor LR: 1e-4, critic LR: 1e-3
- Weight decay: 0, optimizer AdamW
- Exploration noise: Ornstein-Uhlenbeck
The exploration noise¶
The first version had smaller networks with three layers, and it used Gaussian exploration noise. After 2,000 episodes, it did not get to the target. Deeper networks made the result a little better. But after 1,000 more episodes, the task was still not complete.
An Ornstein-Uhlenbeck (OU) process for the noise completed the task in 101 episodes.
On Tennis, this change had the opposite result. OU noise prevented the convergence, and Gaussian noise gave a solution. The reason is this: OU noise has a correlation between timesteps. It pushes the agent to continue in the same direction. A robot arm that moves to a target needs this behavior, because a smooth and continuous movement is the correct behavior. An agent that must react to a tennis ball does not need it.
There is no general rule here. Try the two types of noise. Also, the default values in a paper are correct for the environment in that paper. They are not always correct for a different environment.
Results¶
cd scripts
python train_agent.py --unity-app Reacher.app --target-score 30
Episode 95 Average Score: 29.26
Episode 96 Average Score: 29.35
Episode 97 Average Score: 29.43
Episode 98 Average Score: 29.52
Episode 99 Average Score: 29.61
Episode 100 Average Score: 29.69
Episode 101 Average Score: 30.08
Environment solved in 101 episodes! Average Score: 30.08

The agent completed the task in 101 episodes. This curve is very different from the other two curves. It increases quickly in the first thirty episodes, then it stays near the maximum. After that point, the agent is already good. The 100-episode average needs more time to show it.
Reacher gives a reward at almost each step. Thus, there is no long period of zero rewards, as in Tennis.
This video compares the untrained arms and the trained arms. The untrained arms move randomly. The trained arms find their targets and follow them:
Links¶
- Repo: github.com/sina5/reacher
- Report — full architecture and hyperparameters
- Continuous Control with Deep Reinforcement Learning
The next method to try is D4PG. The ddpg and mlagents scripts come from the Udacity deep-reinforcement-learning repo (MIT).