AWScompeting

AWS Agentic Football Cup

Five AI agents, one pitch, ten matches a week. I coach them with prompts.

multi-agent systems, prompt engineering, experiment design
football, games · AWS Bedrock agents, browser instrumentation, match telemetry
2ndof 465 teams · League A · as of 27 Sept 2026live leaderboard ↗
1 Bugle Summits27–0–31020
2 Bugle Lances24–0–6924
3 Bugle Mariners25–1–4903
4 Chalk Sharks24–0–6884
5 Chalk Gannets23–1–6864

Why I entered

I stumbled onto this one. I play and watch a lot of football, I'd been taking on AI projects all year, and here was a competition that was both: build a five-a-side team of AI agents and coach them, in plain English, against everyone else's. The prize didn't hurt either. Win your league and AWS flies you to Las Vegas for the finals at re:Invent in December, with a share of a fifty-thousand-dollar pool on the line.

My team is Bugle Lances. I didn't pick it; every team gets a name assigned.

What it is

Six leagues of up to two thousand teams, seven weeks, ten matches a week from Friday to Sunday. Before each round you write a role prompt for each of your five players and pick the model they run on. Then the match engine plays your agents against another team's, automatically, and the results land on a public leaderboard. Top your league after week seven and you're on the plane. There are also wildcard weeks, with a Vegas spot for things like the fastest first goal or the longest winning streak.

You never touch the players during a match. Everything you can do, you do before kickoff, in the prompts.

How I'm approaching it

As an experiment, not a vibe. Early on I realized I was changing prompts on feel and had no idea which changes were working, so I built a way to record what my players actually do during a match: where they are, what they choose, when they pass, shoot, mark, or stand still. That gave me something to look at after every ten-match weekend besides the score.

From there it's the boring loop. Change one thing. Play the round. Look at the data. Keep it or revert it. Two things that turned out to matter more than I expected: which model each player runs on, and keeping prompts short enough that the agents decide fast. A player that thinks for too long is a player standing still.

I'm holding the specifics until the season ends. It's a competition, and the field can read.

What I'm learning

The competitive part is what makes it honest. Every prompt idea gets tested against four hundred and sixty other teams' prompt ideas within seventy-two hours, and the leaderboard doesn't care how clever the idea sounded. It's the most direct feedback on prompting I've had.

what I learned

Prompting is an experiment, not a vibe. Change one thing, play ten matches, look at the data.