The Inklands: a storybook adventure for young writers. A glowing open book on a wooden stand in an enchanted forest at dusk, with a castle in the distance.

Six Lessons from Building a Game with AI

This summer I built a video game for my eight-year-old daughter, almost entirely with AI agents. This is what went wrong along the way, what I did about each problem, and the tool I ended up writing because of it.

Sunghwan Yoo · October 2026

This is the written version of a talk. If you'd rather watch, the 24-minute video is right below, with the slides just under it.

Download the slides (PDF, 4.8 MB)

This summer I built a video game for my eight-year-old daughter, almost entirely with AI agents. This is what went wrong along the way, what I did about each problem, and the tool I ended up writing because of it. It's one person's experience on my own projects, so read it as a field report, not research.

Why a game

My daughter had just started summer vacation. I didn't want her watching TV all day, and she needed practice with reading and writing, so we bought practice books. She didn't want to do them at all.

I'm a software engineer, so I started to wonder whether I could solve this with software. She loved the Professor Layton games, where the puzzles are wrapped inside a story. So I decided to make a game where the puzzles are reading and writing quests: you read and write to help the characters. The boss fights are writing duels. You get three words, you write a story with them, and your story becomes an illustrated picture book.

Her birthday was three weeks away, so it would be a surprise present. I wanted a real story, something that would keep her going for 40 to 50 hours, running on an ordinary computer, with the engine built from scratch. As the deadline got closer, it turned into a real software project with a tight deadline, like any project at work. Looking back, that was not a three-week project.

Three weeks became two months

Commits per week, with her birthday and the full game marked

What she got on her birthday, July 6, was about ten percent of the game: one region. The whole game, all eight regions playable end to end, arrived on August 18. By the end of August there were 463 commits. The peak, 152 commits in one week in mid-August, was the week I started running three AI agents at once. That's part two of this essay.

Here's what shipped: 8 regions, 171 scenes, 683 quests and 110 characters, with about 1,430 AI-generated images and 34 short films. Almost all of the art, music, video, quests and dialog were generated with AI, because I'm not an artist. The engine is React and Vite, written from scratch, and AI agents wrote most of that code too.

You explore a storybook world and talk to animals. Every quest is reading or writing in disguise: you answer questions about a passage, or you write sentences that have to clear a quality bar. If a sentence doesn't, an AI coach explains how to improve it, instead of fixing it for you. Between the stories there are jigsaws, crosswords, word searches and coloring pages.

Did it work?

My daughter playing the game

She loved what she calls "the daddy game", and asked for more every single day. Of course, she doesn't know it's educational software in disguise. She also became my junior test engineer. She'd tell me things like "I can't understand this question", "This button doesn't work" or "That name is wrong". I sat next to her and wrote everything down, and that list drove most of the fixes.

Two months in, the game was playable from start to finish. She finished it, and it took her 51 hours.

The Mouse's Sticky House, her picture book

The part she loved most is that her own writing becomes a picture book. At the end of each region she beats the boss by writing a better story from three given words. She wrote this one, The Mouse's Sticky House. The game split her sentences into pages, illustrated them with the same mouse on every page, and put the book on her shelf. Her words are untouched, spelling and all. Even the cliffhanger on the last page is hers: "Find out in our next book."

If I'm honest, I'd give the game 65 out of 100. The systems are solid, and the images are good for a kids' game. The dialog is readable, but often bland. The story world is fine, but the storytelling is weak, and the gameplay needs another round of polish. Video is the weakest part: characters drift, ghost and grow extra fingers. That's today's video models, and I think it will take a year or two. Still, I'm satisfied. She loved it, she played it to the end, and her writing and typing got visibly better over the summer.

But the game isn't the main point here. The more interesting part is the problems I ran into, and how I solved them.

Part one: six problems, six lessons

1. Models differ, so learn where each one is strong

I started in Google's Antigravity, with Gemini 3.5 Flash. When I told it exactly what to do, it did it. But it missed corner cases, and sometimes said it had fixed something when it hadn't. So I tried Claude as well, and split the work: Claude on the hard tasks, Flash on the straightforward ones. Four differences stood out:

  • Capability. Claude writes long dialog and scripts easily, where Flash was often too short. But Flash can listen to music, and Claude can only measure an audio file.
  • English tone. Claude's English has habits. It loves words like pinned, prose, rung and load-bearing, words real people rarely say, and my daughter kept asking what they meant. I sometimes had another model rephrase Claude's dialog to remove them. (A test finally stopped it; see lesson 4.)
  • Autonomy. Claude chases corner cases and failing tests without being asked.
  • Churn. Flash would say "done", and break something else. One script fix took Flash five rounds, and Claude fixed it in one.

The lesson: learn where each model is strong, and adapt how you prompt it. Both could have built this game, and the gap between them narrows with every release.

2. Too many assets: generate the prompts, and treat them as code

I needed 110 characters, over 300 backgrounds, hundreds of quests and thousands of lines of dialog. And volume was only half of it. Almost every asset needed several rounds: change the hair color, move a figure. Every round meant opening a prompt, editing it by hand and regenerating. Writing and rewriting one prompt per asset doesn't scale.

The answer is meta-programming, applied to prompts. Normal prompting goes from a prompt to an outcome. Meta-prompting goes from a prompt, to prompts, to outcomes: you use a prompt to generate the prompts. For example: "Generate ten prompts for animal character sprites for a children's game, all different animals, and store them in a JSON file so an AI can edit them later."

One stored prompt, compiled from three layers

Each character is built from three layers. The first is the same for every image in the game: the house style, plus the rules every character gets. The second is one character, say Hazel, a squirrel shopkeeper: her description, her height and a face lock. The third is one line per pose or expression. Her base pose is generated from text, and every other pose is an edit of that picture, with the same face lock each time, so she stays the same squirrel. That's one line of authoring per pose, for 110 characters and 587 poses.

By the end, the stored prompts ran to 17,000 lines, and a one-line change to the house style showed up as sixty changed prompts in the diff, before I spent any quota. Prompts became code.

The lesson: don't hand-write prompts one at a time. Generate them, store them as data, and review a change to them the way you'd review code.

3. Usage limits: Max isn't always the answer

I kept hitting the five-hour limit on Claude Pro. In practice it ran out after about an hour of heavy use, and then I waited four hours. The obvious answer is Claude Max, at $100 a month. Should I buy it?

Max: five times the five-hour window, three and a half times the week

Max costs as much as five Pro subscriptions. It gives you five Pros' worth of the five-hour window, but only about three and a half Pros' worth of the weekly limit. So Max is great for a busy day or a short sprint. But if you work long hours every day, the weekly limit stops you first, and you never use all the five-hour capacity you paid for. You pay for five and use three and a half. Five Pros, for the same money, give you five of both.

So I ended up with three Claude Pro subscriptions, plus Antigravity: two or three prompts running in parallel, with three five-hour windows and three weekly limits to draw from, for $60 a month. That solved the wall. It also created new problems, which are part two.

The lesson: check what a plan actually limits, the burst and the week, before you upgrade.

4. Agents repeat the same mistakes: make the rule a test

I'd tell an agent to stop doing something, and it stopped, for that session. The next session, it was back. A prompt is forgotten; a test is not. So the fix is to make the rule a test that fails, not a line in a prompt.

Two examples, and neither one is about code. They test the content:

  • The AI kept making the correct multiple-choice answer the longest one, so a child could win by picking the longest button without reading. Now a test checks the length of every answer, and its failure message says why. The agent reads its own failure and fixes it, with no prompt from me.
  • Claude's favorite odd words, like crate and load-bearing. One even came back as a scene name. Same fix: a banned-word list, and a test that fails on any of them.

The test became the policy, and the error message became the lesson. By the end, the game had almost 5,000 tests.

The lesson: a rule that has to survive the next session belongs in a failing test, with an error message written for the agent that will read it.

5. Context is never big enough: give agents a library, not a memory

Every agent has a context limit, around a million tokens today, and for a large, long-running project that isn't enough. The project doesn't fit, and every new session starts from zero. It's like a new hire with amnesia every morning: it forgets the decisions we already made, repeats mistakes we already fixed, and quietly brings old bugs back. So the real problem isn't getting a bigger context. It's structure: the agent needs a way to read only the part it needs.

My answer was markdown files as a knowledge base. The agent writes its plan into a doc first, then executes. The docs record why each decision was made, what's left, and the pitfalls, so the next agent doesn't repeat a mistake. And no agent reads all of it. AGENTS.md is an index: before you touch a quest, read this doc; before you generate art, read that one. Folder and file names do the rest, so each agent pulls in only the handful of files its task needs.

The rule that makes it work: the docs are always up to date, and keeping them that way is every agent's job. It adds up. For every two lines of application code, there's about one line of documentation written for agents to read.

The lesson: don't wait for a bigger context window. Write the project down in small, indexed docs, and make keeping them current part of every task.

6. Humans in the loop: automate yourself out

A loop is one development iteration: prompt, wait, verify. The next prompt is the re-prompt. A game has many kinds of loop (code, sprites, backgrounds, music, video), and one fix takes three or four of them. Every prompt and every verify was me, and that's where my time went.

Take making the art. Even with the prompt ready, every image and every video meant copy, paste, wait, download and rename, by hand. About two thousand times, for this game, and every one of them was me.

So I built a tool that types the prompt into the browser, downloads the result to the right place, and post-processes it. Asset generation got about ten times faster, and I could let it run overnight. That became the pattern: every tool I built removed one place where I had to sit in the loop.

Then I wrote agent skills that do my checks for me: play the game and report what breaks; read a region as a story; look at every frame of every video; and ask whether each question can actually be answered from what's on the screen. They try to copy my own thought process. Alongside them are helper tools, like contact sheets from video and a way for an agent to click through the game itself.

What the video audit sees: a frame every half second

This is what the video audit reads: a frame every half second. It shows two real defects in one clip, both live in the game at the time: a watermark in the corner, and the same bird drawn two different ways. The agent finds them, edits the prompt, re-renders and audits again, with no human until it passes. The watermark, by the way, had been drawn in by the watermark remover itself, running on footage that was already clean. Tools fail too, so the audit checks their output.

Before, every loop needed me twice, to prompt and to verify. After, I write the first prompt, the agents run the loops in between, and I verify once at the end. It didn't catch everything, and the quality varied. But it took 30 to 50 percent of the obvious issues off my plate, and sometimes it found hidden ones too.

The lesson: take the human out of the loop wherever the quality stays as good or better. I think this is a glimpse of where software engineering is going.

Part two: many agents

With those tools, the game got done. But she asked for more every day, and I was working until three in the morning. To keep up, I kept several Claude instances running at once, on the extra subscriptions from lesson 3. That became its own problem. Running three agents at once brought four new problems, and every one of them is really a cost problem.

Merging hurts

At first all three agents worked in the same checkout, and each one overwrote the others' work. The fix is git worktrees: each branch checked out in its own directory, so they can't collide. But two long conversations still produce big conflicts at the end, and each agent spends a long time merging. The longer branches live, the bigger the conflicts, and every conflict is an agent re-reading both sides, at full price, to fix something nobody was building.

Which model, which effort?

A less powerful model won't automatically save you money. It reads more tokens and churns more, so you can end up paying more. On my own tasks, Sonnet read 10.7 million tokens per task and Opus 7.3 million, and Opus scored higher. Effort is the same story: Gemini 3.8 Flash on medium effort read 29 million tokens a task, and on high, 19 million, and high scored higher. A weaker setting wanders: it re-reads, retries, and sometimes has to be re-run. Databricks found the same on their own codebase: Sonnet 5 is cheaper per token than Opus 4.8, but cost more per task, because it used almost twice the tokens.

How much quota is left?

Every account has two windows, the five-hour one and the weekly one, and I had three accounts. Subscription quota is subsidized, so whatever you don't use before it resets is simply gone, and running out means waiting.

Three accounts' five-hour and weekly windows

In this example, all three accounts have about the same weekly quota left. But account two's weekly limit resets tonight, and whatever it hasn't spent by then is gone. The five-hour windows all differ: account one is nearing 80 percent, and account three has barely started. So which account gets the next task? Account two: spend it before it resets. It's a scheduling decision, and the answer changes every hour.

Is the cache warm?

One day I took a lunch break, came back, and sent Claude one message to continue. My five-hour usage jumped from zero to eleven percent, on one message. Usually one message barely moves it. That's how I found out about prompt caching.

Every request re-sends the whole conversation. The part the model has already seen is cached, so it's read again at about a tenth of the price. Only the new part costs about double, to be cached for next time. That's why a long conversation stays cheap. But the cache only lives for a limited time, about an hour on a Claude subscription. Go to lunch and it's gone, and the whole conversation is cached again, at double. That's about twenty times the warm price, for the same message.

A cheat sheet for before you walk away:

  • Claude, on a subscription: about an hour
  • Codex: about 30 minutes
  • Muse, or the Claude API by default: 5 minutes
  • If a session has already gone cold, don't try to rescue it. Start a new one, if you can.
  • If it's still warm but you'll be gone longer than its cache lasts, on Claude run /compact before you go, so you come back to a short summary instead of a huge cold history. Everywhere else, just let it go.
  • On a five-minute cache, even a coffee is a cost decision.

It's a scheduling problem

Put the four together, and with several agents, every task becomes a decision. Is there a warm session? How much of the five-hour and weekly windows is left? Is there quota that expires if I don't use it? Which model, and which effort? Should I compact before lunch?

That's a scheduling problem, and I was doing it in my head. It doesn't scale, so I decided to build a tool. Besides, my daughter has already asked for another game for her ninth birthday. With math, this time.

Warmstart

There are something like two hundred AI orchestrators on GitHub, and none did exactly what I needed. So I built Warmstart. It's open source, under the Apache 2.0 license.

Warmstart: the fleet of agent accounts with their quota, and the task list

This is my real window: the fleet of agent accounts across the top, each with its quota, and the game's task list below.

The key idea is to turn long conversations into small, independent tasks. My real pattern was: playtest, write down the fixes, and prompt them in parallel. That's not a conversation, it's a bug list. Basically, an issue tracker for AI agents. Single-turn tasks bring three big benefits:

  1. Merging. Short tasks branch, finish quickly and merge cleanly, where long conversations drift for days.
  2. Dependencies. If one task can only run after another, say a backend change before the screen that calls it, the second waits as blocked, and unblocks itself the moment the first lands.
  3. Exact accounting. Every task has its own price in dollars, its own active time, and its own landed commit. In one long conversation, you never know which turn cost what.

One task shape doesn't fit every job, so Warmstart has five ways to file one, all on the same fleet and the same queue:

The five ways to file a task: Single Task, Conversation, Plan and Execute, Plan and Split, Debate
  • Single Task, the default and what I use most: one autonomous turn that codes, verifies and lands, and can stop to ask me a question.
  • Conversation, for when I don't know what I want yet: it pauses after every turn.
  • Plan and Execute: a strong model writes the plan, and a cheaper model implements it.
  • Plan and Split: a planner cuts large work into dependent pieces, several agents build them in parallel, and the planner integrates them.
  • Debate: two to five agents, often from different vendors, answer the same question blind, and an organizer weighs where they agree and where they dissent.

Measuring quality: agents grading agents

By now I could see the speed and the cost of every task, but not its quality. Not a benchmark score: the quality of my own work. So I borrowed peer review from academia. Agents grade each other's finished tasks against a rubric, from zero to ten. So far that's 836 reviews of 472 finished tasks. A caveat: this is an anecdote, not a real benchmark. But it came out surprisingly close to the public benchmarks.

Quality against cost per task

Here's quality against cost, both averaged per task, from Warmstart's own statistics page. Two things jump out. First, the best model isn't the most expensive: Opus 5.5 scores 9.2, at about eight cents a task, and GPT-6 Sol 9.0, at nine cents. Second, a cheaper-looking model isn't always cheaper: Gemini 3.8 Flash costs fourteen cents a task for 8.1, more than Sonnet 5 at seven cents for 8.4. Muse is the value pick, at 8.8 for four cents, but that's Meta's Contributor plan, where you agree to donate your conversations and output.

The score each model gives when grading others

The most interesting finding is that models grade each other very differently. GPT is the strictest, at 7.7 on average. Gemini is the most generous, at 9.2. My guess is that this is part of the quality gap. To improve, you need self-reflection in the loop: you have to be able to find the defect in your own work. That's one fleet, my own, so take it as a hint, not a verdict.

Where this goes

Models and tools keep getting stronger, and we'll be asked to tighten the loop, with fewer humans in it. Steps that used to need a person, like reviewing code, debugging or triaging a failing test, are already moving to agents, and as long as quality holds, that will only accelerate. So three things to take away:

  1. Turn conversations into small, single-turn tasks.
  2. Treat quota, cache and cost as a scheduling problem, and give it to a tool.
  3. Measure quality too, with agents grading agents.

I'm deeply rooted in coding, and I think in three to five years most of that goes away. We'll write specs and requirements, with agents. But the job doesn't disappear. We go from coding to solving problems at a higher level. I'm a problem solver, and I love that part.

Try it

Warmstart is open source. Try it, and tell me what breaks; I'm sure a lot does.

The game isn't open to the public yet. It needs more polish, and a children's game needs a few more safeguards before anyone else's child plays it. You can watch the trailer at theinkland.com.

No comments yet