For the 2018 and 2022 World Cups, I ran a small private pool for friends and family. Nothing fancy, but it always added some fun to the tournament. I managed the 2018 pool offline and used RunYourPool in 2022. As the 2026 kickoff drew closer, I looked around again but did not find anything that fit how I wanted invitations, scoring, deadlines, bracket flow, player comparisons, and payments to work. So I decided to vibe code a custom app with my dear friends Claude and Codex.

It took several days, many iterations, and a fair amount of production support. By the end, the app handled invites, passwordless login, group and knockout picks, standings, payment tracking, announcements, score updates, and the final results. For the most part, it worked as designed, although there were definitely glitches and bugs to fix. The custom flow gave the group more reasons to get involved with the tournament, compare picks, and follow matches that might otherwise have passed without much attention.

The project made custom, one-off software feel much more practical. It was also a good reminder that generated code is only one part of getting a useful system into people’s hands.

App vs. Platform

I started with the idea of building an app, then got a bit more ambitious midway through the build and tried to extend it into a reusable multi-pool and multi-tournament platform. That was a bad detour, and I quickly course-corrected. Scope creep can derail you even if coding agents are doing all the code generation. Constraining the pool to one tournament was critical to shipping on time. I could make decisions around the exact bracket, scoring rules, lock times, and organizer workflow instead of designing an architecture capable of supporting any knockout competition. That was possible, but not within the time and attention budget I was willing to allocate to this project.

Ease of use mattered most. Players joined with a pool code or invite link, signed in through a magic link, and filled out group and knockout picks. Drafts saved automatically. Completing the knockout bracket also produced podium picks, which kept those answers consistent. Once matches began, the same workspace showed standings, results, and other players’ picks.

The organizer side tracked participants, payment status, prize eligibility, announcements, exports, and audit checks. Money stayed outside the app; it recorded who had paid and how the pot worked. That small boundary avoided turning a social tool into a payment product.

The result was specific in ways an off-the-shelf product could not be. If a rule confused people, I could change the copy or the behavior that day. If the group wanted a different email update, I could add it. Control was the payoff.

The completed World Cup pool knockout bracket, showing results from the round of 32 through the final
The completed knockout bracket. Select the image to view it at full size.

Where the agents earned their keep

Coding agents were most useful when the goal was concrete and the feedback loop was short. A rough request could become a working flow, then tests and production feedback could push it through several revisions. That was especially effective for UI changes, scoring rules, operator scripts, and tracing a visible bug back through the system.

One example was passwordless login. Email scanners can open a magic link before the recipient does and consume its token. The fix was small but important: the link opened a confirmation page, and only a deliberate POST consumed the token and established the session. The agent helped trace the symptom, change the flow, and cover the behavior with tests.

The tournament itself produced more corrections. Users spotted inconsistencies between the app bracket and the official bracket, so I had to correct the bracket topology. A small number of players needed an emergency unlock without reopening matches they had already completed. I could make those changes because I was both the app author and the pool administrator; an off-the-shelf app would not have given me that control. The agents made it practical to diagnose and ship the fixes on the tournament’s schedule.

The parts that needed care

The app uses Next.js on Vercel, with Supabase for authentication and durable data. Row Level Security keeps pool membership, organizer privileges, picks, payments, and announcements inside the right boundaries. A few trusted server paths use the Supabase service role for operations such as accepting invites, exporting organizer data, and synchronizing results. Those paths stay narrow and server-only.

Picks live in relational tables rather than one large document. Final submission runs through a database function so the group, knockout, podium, and submission records update together. Scoring remains pure domain logic, which made it easier to test the parts most likely to cause arguments.

Standings are recomputed from the current picks and results instead of being stored as a separate leaderboard cache. That costs some computation, but it removes another stateful thing to repair when a result changes. For a small private pool, that was a useful trade.

Results arrived through a protected sync endpoint that normalized an ESPN scoreboard feed. Supabase cron called the endpoint on the tournament schedule, with a daily Vercel job as a fallback. The sync could update result and read-model tables, but it could not touch anyone’s picks. Dry-run modes and manual scripts let me inspect fixture and result changes before writing them.

This architecture was reliable enough for its job because it was explicit about the tournament. Twelve groups, the best third-place teams, fixed match numbers, and lock dates all appear in the domain. Reusing the app for a different tournament would require substantial work across the schema, bracket rules, scoring, copy, and operations.

Architecture diagram showing player and organizer browsers connecting to a Next.js app, Supabase authentication and database services, email, scheduled result synchronization, and live standings
How the app, database, email, and result-sync paths fit together. Select the diagram to view it at full size.

The work did not disappear

The agents made the build feasible for one person, but I still owned the design, production behavior, and support.

UI polish still took repeated subjective correction. Sports feeds needed human cross-checks. Email delivery, cron configuration, database migrations, and deployment settings all needed operator attention. Production bugs arrived on the tournament’s schedule, not mine. Supporting the app meant switching between product manager, designer, tester, data auditor, support desk, and on-call engineer.

The final standings deserved special care. Before sending the winner announcement, I checked the result independently from the raw picks, match results, payment status, and eligibility rules. Tests made the scoring logic safer, but a visible audit path made the outcome easier to trust.

The final pool leaderboard with player names redacted, showing the scoring breakdown by podium, group, and knockout picks
The final leaderboard. The scoring breakdown made the result easier to audit.

The most useful operational features were unglamorous: exports, dry runs, manual backups, audit records, and a clear way to correct data. They mattered because this was a social app. A scoring mistake would not take down a company, but it could spoil something friends and family had enjoyed throughout the tournament.

What I would do again

The specificity kept the project manageable and made rapid feedback useful. I would also add operational checks earlier. Auth, external data, email, lock transitions, and correction paths carried more risk than the main screens suggested.

I would continue to keep scoring and submission rules separate from the screens that display them. Each rule has its own tests, so the same calculation and deadline apply everywhere in the app. When prizes and bragging rights depend on the standings, “the page looked right” is not enough. I would keep the result sync away from user picks, retain dry runs for every production write, and plan the final audit while designing the first schema.

I would resist generalizing the app until a second real use appeared. The code now carries assumptions from this tournament, and hiding them behind generic names would not make them disappear. A reusable product would need deliberate work on configuration, bracket formats, provider failures, replayable imports, backup and restore, and monitoring.

For this pool, that second pass was unnecessary. The app made the tournament more engaging for the people in it, and building something shaped around our own rules was part of the fun. That was enough to make the trade worthwhile.