Inspiration
During COVID the UFC was one of the only sports still running, and I became a fan for a while. I'm not much of one anymore, but something I noticed then stuck with me: fighters who had been around longer often beat newcomers who looked more talented. Experience seemed to matter more than it should.
I wanted to test that with an Elo rating, the chess number, applied to fighters. The catch is that a UFC newcomer has no UFC history, so a rating built only on UFC fights knows nothing about them. But they've all fought somewhere. So I built the rating from every professional bout I could find, in any promotion, amateur fights excluded.
What I built
A scraper that collected fight records from public fighter pages: about 143,000 professional bouts across 50,000 fighters, from 1984 to mid 2025, with method of victory, round, and time. That was the hardest and least glamorous part, and eventually the reason I stopped: keeping it current was a chore.
On top of that, an Elo rating computed fight by fight in date order, everyone starting at 1500. The sanity check passed: the top of the table was Jon Jones, Islam Makhachev, and Khabib Nurmagomedov.
Then the part I'd half forgotten. The fighters and their bouts form a graph, so I trained a graph neural network on it, with each fighter's Elo, height, reach, weight, age, and base style as features. It has two outputs for any matchup: who wins, and how, by knockout, submission, or decision.
Tap a fighter to pick, then the result reveals. Names and records only; no photos.
How it did
I ran it on one card before the fights happened: UFC Baku, June 2025, predictions made three days out. It called 6 of 12 winners and 6 of 12 methods. The plain Elo rating, no neural network, would have called 7. So on the one honest test, the simple thing beat the clever thing.
It did have one moment. On the prelims, Ko Seok-hyeon was making his UFC debut against an undefeated Oban Elliott, and almost nobody expected him to win. The model picked Ko, at 55 percent. Ko won all three rounds. I'd love to say it saw something. More likely it saw a fighter with a long record in smaller promotions, which is exactly the thing I built the rating to notice, and got lucky that it mattered that night.
What didn't work
The honest summary is that it was a coin flip. It was fine at calling favorites, which anyone can do, and no better than a guess in close fights, which is the only place a model earns its keep. A coin-flip model is as good as a coin.
Going back through the code for this page, I can see why. Most of the features I was proud of were empty: the scraper only captured date of birth for about 7 percent of fighters, fighting style for 14 percent, reach for under half. So the network was mostly learning from the shape of the graph, and a graph of who beat whom is roughly what Elo already summarizes. Two bugs made it worse. The network was trained on each fighter's rating after the fight instead of before it, which leaks the answer into the features, and the prediction script fed every matchup the dataset's average features instead of the two fighters' own. The 6 of 12 owes more to the graph than to anything I added on top of it.
The bigger lesson is about method. I had one hypothesis, that experience wins, and I never tested it on its own before wrapping a neural network around it. More features, a chronological split, a baseline to beat, and a dozen cards instead of one would have told me whether there was anything there. Predicting fights is a hard problem. I made it harder by not being scientific about it.
A rating that agrees with your eyes is the easy part. A model that beats it is the whole job, and mine didn't.