Coldopen · skill prediction without a match history

How do you rank a player who has never played? I trained bots to answer that.

A new account has no results, so ranked games guess, and the first few matches are miserable for everyone while they do. You could instead read skill from how somebody plays. But to learn that reading you need labeled players, and a game that hasn't launched has none. So I train a ladder of bots from hopeless to expert, learn the mapping from their behavior, and point it at people. The idea works. The hard part was that four separate measurements each told me it didn't, or that it did for a reason that turned out to be false. Every one of them is playable below.

scroll
The setup

Bots as stand-ins for players you don't have yet.

I built two games as simulators I could train agents inside: a Tetris sprint (clear 40 lines as fast as you can, a ranked mode on TETR.IO, a big online Tetris platform) and Minesweeper. Both are real games with real public leaderboards. That matters, because the whole point is to check the bots against actual humans at the end.

The recipe is one line. Make agents of varying skill, record how they play (how fast, how efficiently, how often they set up the big clears), then fit a model from behavior to skill. Apply it to a human and you have a rank estimate from their very first game, with no match history at all.

Then you check it against 396 real Tetris players whose ranks you already know.

Lie #1

The baseline that cheated.

I got this scoreboard. Each method predicts an ordering of 396 ranked players from their play alone, and the score is a rank correlation: 1.0 means the predicted order matches their real ranks exactly, 0 means it might as well be random.

predicting human rank · 396 TETR.IO playerssimulator loses
methodscore
just sort players by pieces-per-second0.932
sort by keystrokes-per-piece0.773
a model fitted on the humans themselves, scored on players it never saw0.923
my agent-trained model0.660
Sorting players by a single number beat the whole simulator by 0.27. I wrote it up as a clean negative result.

That conclusion was wrong.

To sort players by pieces-per-second, you first have to know that pieces-per-second is the right thing to sort by. That knowledge comes from looking at players whose rank you already know.

But that is exactly the thing you don't have at a cold start. I had let the baseline read the answer key and then scored the honest method against it.

The agent-trained model, given zero human labels, put all the weight on pieces-per-second and ignored the other four features. It found the right answer on its own. That is the baseline itself, derived without the labels the baseline needs.

So the honest question is how much labeled data you would need to match this without the simulator. Drag the slider.

label efficiency · same held-out players, 200 random splits each—
fitted on humans — fitted on agents, zero labels — gap —
The amber line never learns anything, because it has no labels to learn from. It's flat because it's the same model at every point on the axis. The steel line is what you could do yourself, if you had that many labeled players. It never catches up within the range the data can measure.

At five labeled games the simulator is worth about 0.10 of correlation. At three hundred it's still ahead. That's the actual result, and it appears only once you stop asking the wrong question.

Lie #2

The check that couldn't see a bot doing nothing.

Before you compare bots to humans, you have to confirm the bots are even in the same neighborhood, that they produce the range of behavior people produce. For each bot I measure five behaviors: speed, keystroke efficiency, quad rate, use of the hold slot, and back-to-back streaks (quads chained without a lesser clear between them). My check asked whether each bot's numbers fell inside the human range. It came back 100% inside on three of the five.

All three were false. The clearest is quad rate, the fraction of your lines cleared four at once. It's the signature move of a good Tetris player.

quad rate · 396 humans vs my bot—
—

Human quad rates span nearly the whole 0-to-1 range, so a bot sitting at 0.01 is technically "inside the range", and simultaneously below every human alive. A range check cannot see the difference between covering a distribution and clustering at the bottom of it.

Asking "which human rank does this look like?" instead exposed it in one line. The cause was simple. My bot's scoring function paid it for any line clear, so it cashed in singles the moment it could and never stacked four rows. Teaching it to hold out for quads moved its quad rate from 0.01 to 0.61.

Lie #3

The failing grade with no passing grade.

The next metric asks whether, for one bot, all five of its measured features agree about how good it is. If its speed says "expert" and its stacking says "beginner", it isn't imitating any real person.

I scored my ladder and got eight ranks of disagreement on an 18-rank scale. I started rewriting the bot.

Then I ran the identical test on the humans.

internal disagreement · lower is more consistent18-rank scale
A single human sprint disagrees with itself by the same eight ranks. Eight is what this measurement reads when nothing is wrong. One game is a small sample, and any real player looks inconsistent across five features. My bots were no worse than a real game.

The number was an uncalibrated measurement. I nearly rebuilt a working system because of it.

Something narrower survived the comparison, and it was more useful: four of the five features sat inside the human spread, and exactly one didn't. My bot used the hold slot ("save this piece for later") like a beginner while playing like an expert everywhere else. Fixing that made it better at the actual game.

Any measurement of "how human is this?" is uninterpretable until you run it on humans and see what score they get.
Lie #4

The dial that was secretly a switch.

To build a ladder of bots you need a knob that makes them worse by degrees. An obvious one is to train a bot to copy an expert and stop training early. A half-trained copy should play half as well.

It doesn't. Drag the accuracy slider. This is a perfect player whose decisions I corrupt at a known rate, so nothing else varies.

corrupting a perfect player · 40-line sprint—
lines cleared of 40 — games actually finished —
Every human record on the leaderboard is a finished game. So the entire human-relevant band is squeezed into the top 5% of this axis, and everything below 90% is indistinguishable failure. A 90%-accurate copy is not a player at all.

A ladder needs a knob whose middle settings produce middle play. Imitation accuracy isn't one, because a sprint is a hundred decisions in a row and one bad one ends it. The knob that worked changes what the bot is trying to do (expert bots hold out for quads, beginner bots grab any clear going), which degrades smoothly across the whole range.

A footnote: a trained-but-imperfect bot is worse than a randomly corrupted one at the same error rate. Random mistakes cancel out. A model's mistakes are systematic. It's wrong the same way on similar boards, so its errors compound.

The design lesson

One knob makes bots that all look the same.

The agent-trained model got confused for a structural reason. I built the Tetris ladder with a single skill knob, so everything moved together: faster bots were also better stackers, better planners, tidier with their keystrokes. Real players don't work like that: some are fast and sloppy, others slow and careful.

When every feature moves in lockstep, a model can't tell which one matters. Mine gave pieces-per-second a negative weight, in a population where speed is the strongest predictor of skill.

I built the second game with three independent knobs instead. The plot below shows both games on the same axes:

speed vs judgment · each dot is one bot configuration—
—

Left: a diagonal smear. Speed and judgment are effectively the same number, so a model fitted on it learns a relationship that falls apart on humans. Right: a cloud. You can be fast and careless, or slow and careful, independently, the way people are.

What a better test looks like

Minesweeper can tell you if a click was a mistake.

The problem with the Tetris sprint is that it's a time trial. Rank correlates 0.93 with raw speed, so there was never room for a cleverer method to add anything. I picked Minesweeper next because it has a skill measure that logic can settle.

After any position, a solver can determine which hidden squares are provably safe. So for every click a player makes, you can ask whether a guaranteed-safe square was available and whether they took it. That's a judgment score. It has nothing to do with speed, and it's computable for bots and humans identically.

Play a few clicks. Hold the hint button to see what the solver knows.

minesweeper · 9×9, 10 minesclick any square to start
The first click is always safe.
provably safe 0 avoidable guess 0 forced guess 0 hit a known mine 0
Green outline: provably safe. Red outline: provably a mine. No outline on a hidden square means the solver can't tell either. If you click one of those and there was a safe square available, that's an avoidable guess. This is the exact measurement applied to the bots, and it would be applied to human replays unchanged.

My best Minesweeper bot reveals a provably-safe square 83% of the time and clicks a known mine 1% of the time. Turn its "lapse rate" up and blunders climb to 19% while its speed doesn't change. That's two independent axes, exactly what the Tetris ladder failed to produce.

An aside

What taught the bot to play.

I spent two million steps of reinforcement learning on Minesweeper and got a bot that never won a single game. It learned something, though: to never click anything. With 99 mines hidden among 480 squares, an uninformed click is bad enough that flagging squares forever is the better strategy. It found the optimum of the reward I gave it, and that optimum was to do nothing.

What worked instead was changing the amount of information per step. The simulator knows where its mines are. So instead of one thumbs-up-or-down per click, give the network the answer for all 480 squares at once and ask it to predict them.

same network, same game, two training signalslearning
Reinforcement learning: 0 wins after 2,000,000 attempts. Supervised on dense labels: winning two thirds of its games after 12,000. This is the same network fed a signal that says something.
The screen, run for real

The screen ruled out the next game before I built it.

The rule I came out of Tetris with was: before building a simulator for a game, check how much of rank one obvious statistic already explains. The next candidate was Tetris versus mode: two players, garbage attacks (your line clears dump junk rows on the opponent), and a rank that comes from a rating fed by match outcomes rather than any single number. That was exactly the property the sprint lacked. I already had the data.

So I ran the screen before writing any of it.

how much of rank does one number already explain?both fail
Anything above the dashed line has no room left for a cleverer method. Versus mode, judged on a single round (the actual cold-start situation), scores 0.922 from pieces-per-second alone, statistically the same as the sprint's 0.932. Tetris skill is speed-limited in every mode.

A rating built from wins and losses didn't help, because the thing that rating measures is still mostly speed. Building the versus simulator would have inherited the exact ceiling that made the first one uninformative.

A label that is itself a performance measure (a time, a score, a words-per-minute) will always be predictable from the behavior that produces it. Headroom exists only in labels that are latent and have several causes.

That one sentence disqualifies whole genres without downloading anything. Typing sites rank you by words-per-minute, so rank and WPM are the same quantity. Rhythm games compute your rating as a formula over accuracy and chart difficulty. The entire speed-and-accuracy family fails identically.

It applies to my own remaining hope too. Minesweeper's leaderboard sorts by time, and its speed stat, 3BV/s, is board value (the minimum clicks a board needs) divided by time. This is the same trap a third time. The label with real room in it is win rate: surviving an expert board is mostly about not guessing, while time is mostly about clicking fast. So the thing to ask a data holder for is win rates, and the games people lost.

Where this stands

The limits of a one-game result.

What holds up: a model trained purely on bots was worth more than 300 labeled human games on the one test I could run. That's a real result and it survives the corrected comparison.

What doesn't: that test ran on a single game, and that game turned out to be nearly the worst possible choice: a time trial whose ranking is 93% explained by one obvious number. No method had headroom to fill. I picked it because I could get the data, which is not the same as picking it because it could answer the question.

Minesweeper is the game where the question can be answered, and the bot side is finished. The human half is stuck: the site's data isn't reachable without hammering their servers in ways I'm not willing to, and my email asking nicely hasn't been answered yet.

Why the stall is structural

Stack the three requirements together. A game has to be simulatable to build bots for. It has to have headroom above its own obvious statistic. And its human data has to be reachable. Public sources fail at least one every time: games rich enough to have interesting telemetry (shooters, MOBAs, card games) are too complex to simulate honestly, while games simple enough to simulate are simple because their skill is thin.

A game studio removes all three constraints at once. They own the simulator. Their ranks come from matchmaking ratings rather than leaderboard times, so the labels are latent. They record losses. And access is a conversation instead of a scrape. They also have the motive, since cold start is their problem: new players quitting in the first week is a number they already track.

So this is the method, the toolkit, and most usefully the one-script screen that tells you in minutes whether your game has any room for it. It is not a solved method.

For anyone starting this: before you build a simulator for a game, check how well the single most obvious statistic already predicts rank. If it's 0.9, nothing is left for you to win.
I published and then retracted three claims in this project: a bug I asserted without checking, a validation gate I believed was green because I'd read it off a terminal rather than the file, and a metric that was measuring a stuck program rather than a trained one. They're all still in the repository's history. A control caught every one.
None of the human data is in the repository. It belongs to the people who published it, and it's personal data even after the names are hashed. Everything you can see on this page is an aggregate.

The failure mode is a confident wrong number.

It looks decisive and points the wrong way. Every trap on this page produced a confident, plausible, wrong conclusion, and each one was caught by asking what the measurement reads when nothing is wrong.

Orbitope · coldopen