Sina's Blog

CS and AI

AI

6 posts

3 min readAIProgramming

Two DDPG Agents Learning to Play Tennis

After the Banana Collector I moved on to Unity’s Tennis environment, which is harder in two ways. The actions are continuous, so DQN’s trick of taking the max over four actions doesn’t work anymore. And there are two agents, each one sitting inside the other’s environment.

2 min readAIProgramming

Training a DQN Agent to Collect Bananas in Unity

This was the first RL agent I trained end to end. The environment is Unity’s Banana Collector: a square arena full of yellow and blue bananas. Yellow is worth +1, blue is worth −1, and the agent has to work out which is which from a stream of numbers.