Live demo: our self-play AI, still learning this game live on a small cloud CPU. The game counter starts at zero so you can watch it play in real time. Our full training runs on dedicated hardware and has played 2.3M+ games. tworobots.com
two robots
self-play training
inside the network ↓ game sense ↗
connecting
—

Games played

tonight's moon
—
of —
—games / min
eta—
finish—
—
—

Am I getting better?

last greedy benchmark · the number that counts
No greedy benchmark round has been measured yet (scripts/autopilot.sh writes one to benchmarks.log).
live seat vs frozen past selves · exploring, not a strength number
—
waiting for the first decided games
50%
Decisions
—/ s
—
Simultaneous games
—
—
Gradient steps
—/ s
—
Corpus written
—MB / s
—
Colourless warbots
—of built
—
Policy commits when it can
—of turns
—
Scoring commits / game
—
—
Wasted round trips / game
—
—

Learner

—
learning rate—
epsilon—
explore—
difficulty—
loss—
TD error—
mean Q—
updates—
batch—
samples / s—
replay store—
GPU—
replay buffer
—

What Adam is doing

—
—

How games are won

—
—

Action mix

—

When it can, does the policy take it?

—

Commit economy

—
—

Warbot colours

—

Time to milestone

round = ceil(turn / 2) · the seat's own turns

Per mechommander

per turn played · learner seat only
ult won = win rate of the games in which that mechommander's ultimate fired

Hero matchups

row = my mechommander · column = opponent's

What I learned

—
first note after a few minutes of watching

Card balance

—
strongest when played
weakest when played

Hero powers

—
Win rate of the games in which the power was used, against that mechommander's own win rate. "Used" is per decision it was ready.

Waste budget

—
—backend round trips per game that produced nothing

At risk

Ongoing

newest first
nothing to report yet

Inside the network

asleep until you scroll here

Nothing below is decorative. The constellations are the network's own card embedding table, the 32 numbers it has learned for every card, read straight out of its weights. The trees are the actual moves it just weighed and the value it gave each one. Calibration checks its predictions against what really happened in games it has not trained on.

Card constellations

Each star is one card. Its position is its learned 32-d embedding squashed to 3-D with PCA, so cards the network treats alike sit close together. Colour is the card's colour, and brightness is how often the learner played it. Faint lines join each card to its nearest neighbour of the same colour, like a star chart: while the network ignores colour they sprawl across the sky, and they will draw tight as it learns that colour matters. Every frame is a saved snapshot of the network, rotated to line up with the one before; press play to watch it reorganise. Drag to orbit, scroll to zoom, hover over a star.

the constellations rise when you scroll here

Live decisions

The network deciding, laid out like the network itself, in six layers: the mechommander, the phase it is in, what it reads, the moves it weighed, which expert decided, and the value. On the left is the mechommander it is playing. Next is what it reads in the position: the foresight layer works out, from the rules, whether a commit point is on this turn, whether a hero hit is open or lethal, whether the opponent can score or reach its hero next turn, whether a bot of its own is one hit from dying, and more. Mint situations are openings, rose ones are threats. Then come the kinds of move it weighed, and on the right the value it gave them: Q, its estimate of the final result, from −1 (loss) to +1 (win). Each real decision sends a soft light along the path it took. The lines from a situation to a move thicken as that response becomes a habit, so you can see, for example, whether reading score now makes it commit. Across the top is what its next-turn heads expect: that it scores before the turn ends, that the opponent scores, that it loses hero HP or a bot on their turn, and how it thinks its opening will turn out (a bot by its round 3, two engines and full HP by round 5, a point by round 6). Those are learned from what really happened, not worked out. Under them is the phase layer: which of its own rounds this is, whether it moved first, whether the encoder calls this the opening, the development, the midgame or the late game, and how its three experts (one head each for early, mid and late play, behind a learned gate) share the decision. When the gate stops following the calendar, that is the network having learned a phase boundary of its own.

waiting for decisions…

Where the brain hands over

Waiting for decisions with a phase reading…

Calibration

Three small heads on the network predict how the game will end, four more predict the next turn, and its Q value predicts the result. Here those predictions face the real endings. Held-out takes a copy of the network, waits a minute, then scores only games that finished after the copy, which it has never trained on. In-sample scores recent rows it has trained on, so it measures fit, not foresight.

calibration wakes when you scroll here