Pushman · an RL devlog

How not to train a bot to play your game

I wanted smart opponents for my little ring-out fighting game. I got them, and found an impressive number of wrong ways to train a reinforcement-learning agent on the way there. Here are the ones worth learning from.

scroll ↓

01 · the game in 60 seconds

Three moves on a triangle, six states underneath.

Pushman is a two-player ring-out fight. Push beats dodge, dodge beats block, block beats push, rock-paper-scissors with a stamina meter on top. Six states under the hood, and you commit when you act: no moving while you charge a push, no pushing while you block. The whole game is reading which one your opponent is about to commit to. That's what I wanted the bots to learn.

Combat triangle
The combat triangle. Each move beats one and loses to another.
Player state machine
Six states. You commit on entry, which is what makes the reads matter.

02 · the first pass

I gave it the basics and pressed go.

The agent sees where it is, where the opponent is, both stamina meters, who's in which state. It can move, turn, and pick one of three actions. For rewards I did the obvious thing: a little for landing hits, a little for surviving, a little for facing the opponent, and a real prize for winning. I used PPO, nothing exotic. It trained, and it played badly, which for a first pass is exactly what you'd expect. The interesting part is how it was bad.

Observation table
What the agent sees, grouped. The last row is the one I forgot (more on that later).
Reward table
What it's paid for. This is the final version, and getting here took a few wrong turns.

03 · mistake — the reward economy

The agent learned to dodge, and forgot how to do anything else.

The rewards were sane and yet the bots only ever dodged. Two-thirds of my rock-paper-scissors went extinct, and the training curve looked fine the entire time. (scroll the panel)

~ SAME REWARD PER HIT DODGE PUSH 1 input 4-step commit extinct THE FIX — WIN DWARFS HITS WIN ×10 edge hit center hit

01

Two moves paid the same.

A dodge and a push paid about the same reward. On paper, that's a fair fight.

02

One move is far more work than the other.

A dodge is one input. A push is a four-step commitment: charge, hold, release, and still be standing there when it lands.

03

So it stopped pushing.

Earning the same payout the hard way made no sense, so the push dropped out of the policy. Block only counters a push, so block disappeared with it.

04

Two-thirds of the triangle disappeared.

Rock-paper-scissors collapsed to just rock. The reward curve looked healthy the whole time.

05

The fix was in the ratios.

It was making the win dwarf any run of hits, and paying the most for hits that shove someone toward the edge, where ring-outs happen.

Reward values, before and after
Every signal retuned, and the dodge-on-empty-stamina exploit removed.

04 · mistake — the missing observation

Even with sane rewards, the bots couldn't aim.

They dodged fine, but their pushes and blocks pointed at nothing. The agent knew where the opponent was, and it knew its own heading, but I never gave it the one number that matters for aiming: whether it was pointed at them. A push only lands in a narrow cone in front of you, and I was asking a handful of neurons to derive that alignment from raw angles, to learn trigonometry from scratch, which it couldn't. I fixed it by computing the dot product myself and handing it over.

I'd been paying the bot to face the enemy long before I gave it any way to know whether it was.

Drag the opponent. Is the push aimed?LANDS
~70° push window SELF
facing · directionToOpponent0.92 · aimed ✓
The whole aim problem in one number: Vector2.Dot(transform.up, dirToOpp). Precompute it instead of making the network learn the trig.

05 · the catch

Fixing an observation meant starting over.

The moment you add an input, the network's first layer changes shape, and every weight you'd trained is now the wrong size. You can't carry it forward. So adding that one aiming number meant throwing the model out and starting fresh, with rebalanced rewards and retuned stamina all at once. That's fine when you plan for it, and miserable when you find out by surprise.

Changed a reward?♻ warm-start, keep the weights you trained
Changed an observation?✗ start over, the first layer is now the wrong shape

06 · what worked — failing forward

You don't always have to start from scratch.

As long as you don't change what the network sees, and keep the rewards on the same scale, you can keep refining: re-weight the rewards, warm-start from the weights you already have, bump the learning rate so the policy explores again, and let it go. You're starting from a local optimum instead of a blank page. The boring discipline that gets you this: pick a representative set of observations and rewards early, so most of your iteration becomes tuning instead of rebuilding.

Warm-start from a local optimum
Keep the weights, raise the learning rate, and climb out of the local optimum.

07 · what worked — variety

One network, five opponents.

I wanted several bots worth fighting, instead of one optimal one. The textbook move is self-play (fight past copies of yourself), but it collapses. The agent finds one opening that beats the ghost pool, and the matches stop being competitive. So instead I trained one network against five hand-defined personalities, each a small reward bias with its identity fed in as a one-hot input. Facing five pressures every episode, no single strategy beats all of them, so the network stays flexible. It's slower at first, but it holds up. Because personality is an input, I can pick which bot you face, and how hard it reacts, after training.

Self-play vs round-robin
Self-play converges to one best response; round-robin forces a multi-modal policy.

↑ personality fed as a one-hot network input. Swap the active row to swap the bot, with no retraining.

reward emphasis

Five personalities, one network. Each is a reward bias and a one-hot identity input. Pick which bot the player faces at runtime, with no retraining.
Difficulty · MEDIUM Lag 6 fr  ·  Noise ±9 px
BOT true position ghost (what bot sees)
EASY EXPERT
Difficulty is a single scalar. Easier bots observe a stale, noisy ghost of your position rather than where you are.

08 · proof

And then they got smart.

The bots learned to bait a block, wait out the opponent's stamina, and punish the drop. That's the policy reading its opponent, which is exactly what I set out to build.

09 · the last mistake

A good bot, but not a fun game.

I won, with a varied, sharp set of bots that read you and punished you at whatever difficulty I picked. Then I sat down to play my own game, as a human, and it wasn't fun. The stamina economy the whole thing was built on was the problem. The bots were a perfect solution to a game I hadn't finished designing.

I optimized hard for a reward I never checked was the right one. I misspecified my own objective. It took the bots three million steps to stop ringing themselves out. Hopefully it doesn't take me that many to remember to play the thing before I train against it.

if you're starting your own

Build something fun to play against first.

The cheapest, highest-payoff thing is one honest opponent worth losing to, rough numbers for stamina and force and distance, and an hour of playing it before you train anything. The network will perfectly optimize whatever you hand it, including your mistakes.