Post two · a project write-up
AI World Gen: thousands of decisions, and a loop that graded them
A map generator that asks a model one question per tile. A new kind of model that answers in milliseconds for a fraction of a cent. And the evaluation loop that took it from a green blur to a village, one measured change at a time, most of them written by an agent.
The short version, if you only have a minute:
- Deciding one tile at a time (cell by cell) is the hardest way to get a good map. Version 1 drew maps where almost every cell was the same, and the judge (a model that grades each map) gave them 0.09 out of 1, the worst of the three methods in this post. Measured changes, one at a time, took it to villages (0.69).
- The biggest step took work away from the model. In version 8, code draws the rooms first, and the model only fills them in. This shows how hard it is for a model to build a shape when it makes one decision after another, instead of seeing the whole map at once.
- An LLM was good with no tuning at all. In the LLM method, a language model draws the whole map in one answer. It scored 0.68 with only what version 1 knew. At version 11, it trails cell by cell by 0.04 (0.65 against 0.69), a gap smaller than the noise between two runs (0.05). It is also up to 15 times faster, at the same price. Cell by cell keeps a clear lead where code controls each cell: the rules (0.99 against 0.76), routes to doors (0.82 against 0.61) and whether the map reads as a place (0.88 against 0.74). Batched, where Jev (the decision model) gets the questions for all the cells in one request, did worst: each cell is decided without knowing what its neighbours got, so the map contradicts itself (0.57).
Contents
A model that decides instead of writes
AI World Gen started as a way to test Jev, a new kind of model made by TypeSafe. I wanted answers to three questions. How much of the context can it take into account? How much does it really cost? And can you tune its prompt, the way you tune the prompt of a large language model (LLM)?
The idea is simple: describe a place, and watch an AI build its map one tile at a time. The hard question was never “can a model draw a map?”. It was “can it draw a map that is useful and logical?”.
Asking a typical chat model (an LLM) for every tile is slow and expensive, and you get a sentence back when you only want one word. Jev works differently. You send it a state (a JSON object that describes the situation) and a typed question: “pick one of these options”. It answers with a probability for each option, and it writes no text at all. It is built for speed: TypeSafe quotes tens of milliseconds for the model itself. It also costs about $0.04 per million input tokens (a token is a piece of a word, about four letters), roughly fifty times less than the chat models that I would otherwise use. The real bill has its own section below.
The generator splits the work in two. First, one call to a language model invents the vocabulary: the six to twenty-four kinds of tile that this place is made of, and the rules for each one. Then Jev decides each tile, with the vocabulary as the options and the map so far as the state.
After that, the generator decides what goes in each cell, one cell at a time, and so builds the entire map. It selects an empty cell, reads the full map state (the contents of all the cells filled so far), and asks Jev: “What is the most probable content for this empty cell?” The model decides, the cell is filled, and the generator moves on to the next empty cell. I call this way of drawing a map cell by cell.
What it costs, and how that grows
The bill for the whole project comes from the usage page of OpenRouter, the service that sells access to Jev and to the other models. It is $5.72 in total. Jev cost $3.85: $3.11 for the 39.2 thousand decisions of cell by cell, $0.51 for the five batched runs in this post, and $0.23 for a repeat run and tests. Claude Sonnet 5 cost $1.84: $1.24 for the five LLM runs, $0.33 for the 7 calls that wrote vocabularies, and $0.27 for tests. The last $0.03 was two calls to a third model, DeepSeek V4 Pro. So one cell by cell decision costs eight thousandths of a cent, and one vocabulary costs $0.047.
Two other questions matter more than the total: what does one map cost, and how does that cost change when the map gets bigger? This post draws a map in three ways, and their costs grow differently.
- Cell by cell: Jev, one request per cell. Each question sends the state again, and the state includes the whole map. So a bigger map means more questions, and each question is longer.
- Batched: Jev, all the cells in one request. The state that all cells share (the instruction, the world, the map and its balance sheet) is sent once per request, and each cell adds only its own part. But one request holds at most about 65,000 tokens. So a bigger map needs more requests, and each request sends the shared state again.
- LLM: a language model draws the whole map in one answer. The prompt is sent once, and the model answers with about one letter per cell. So a bigger map only means a longer answer.
Cell by cell and LLM cost about the same at 8 by 8. Above that size, LLM is cheaper, and the difference grows with the map: LLM is 8 times cheaper at 32 × 32, and 16 times cheaper at 64 × 64.
Batched costs three to five times less than cell by cell at every size, because it sends the shared state once per request instead of once per cell. Up to about 16 by 16, it is the cheapest of the three ways. Above that, it grows faster than LLM. Each cell’s question still carries its own list of options with their rules, about 600 to 750 tokens per cell, while LLM writes about one token per cell. So at 64 × 64, a batched map costs $0.13 (in 59 requests), against $0.042 for LLM.
Jev is fifty times cheaper per token, but cell by cell sends the map again with every question, so it spends most of that saving on repetition. Still, Jev's price is what makes cell by cell possible at all. The same cell by cell method with Claude Sonnet 5 instead of Jev would cost $1.22 for one 16 × 16 map.
A new description also pays $0.047 once, at the start, for its vocabulary, whichever way draws the map. For a small map, that is more than the drawing itself costs. A preset ships its vocabulary and pays nothing, and every test case in this post uses a preset.
| Map | Cells | Cell by cell | Batched | LLM |
|---|---|---|---|---|
| 8 × 8 | 64 | $0.0054 | $0.0016 | $0.0051 |
| 16 × 16 | 256 | $0.023 | $0.0074 | $0.0069 |
| 32 × 32 | 1,024 | $0.11 | $0.031 | $0.014 |
| 64 × 64 | 4,096 | $0.68 | $0.13 | $0.042 |
Version one was a green blur
The first version came from an afternoon of vibe coding (fast building with an AI assistant, with little planning). You describe a place, press the button, and watch the tiles appear one by one. It worked, in the sense that every call got an answer and a map came out at the end. Then I looked at the map.
The map was useless. In the first test set, thirteen of the fifteen maps used one tile type for more than 85% of the cells: a village that was all grass, a station that was all corridor.
The cause is simple. The model decides a cell mostly from its neighbours. If the neighbours are grass, it picks grass. Each decision is reasonable on its own, but together they make a map with nothing in it.
I could have looked at three maps, changed the prompt, and looked at three more. I have done that before, and it usually ends the same way: the change that fixes the village breaks the station. So before I changed the generator, I built a way to measure it. Then I built a loop around that measurement: change the code, measure every map, compare with the last version, and change again. Every version is scored on the same maps, so an improvement cannot quietly undo an older one.
Measuring instead of looking
The measurement lives in Galtea, an evaluation platform for AI products. I work there, so read this as a demonstration as much as a recommendation. In Galtea, AI World Gen is a product, and everything below belongs to it:
- A specification is one sentence about what a good map does, such as “Every walkable area is reachable, and the map is playable.” This project has seven.
- A metric measures one part of a specification as a number from 0 to 1. Each specification has several metrics, and in this post its score is the mean of them. There are two kinds. Twenty-eight metrics are plain code in this project: “doors sit in walls” is a loop over the door tiles, and “one walkable region” is a flood fill. The code computes them and sends the scores to Galtea. The last metric is the judge. It scores what code cannot check: does this look like a village? Galtea runs it with another model, GPT-5.2, which reads each finished map as text (one letter per cell, with a legend) and gives it a score from 0 to 1. The post always shows the judge on its own, outside the mean.
- A test case is one fixed input: a preset (the described place), a grid size, a fill order (the order in which the generator visits the cells) and a random seed. In this post, I also call a test case a seed. Galtea groups test cases into datasets: here, one dataset of five test cases for each specification, plus one dataset of three 16 by 16 maps. The same 38 test cases run against every version, so two versions are always compared on the same maps.
- A version is one change to the generator. When a version runs on all the test cases, each map is saved as a session: the test case goes in, and the finished map comes out. Each score in a session is an evaluation: one metric applied to that one map.
I added specifications as new failures appeared. At the start there were three: barriers form structures, paths form routes, and every walkable cell is reachable. After version 4, the maps were no longer random, but they were wrong in a new way, so I added a fourth: a place reads as a place. After version 7, I added three more: the vocabulary's own rules hold, routes lead to doors and off the map, and each landmark appears once, without filling the map.
How the generator evolved
Each chart below is one specification. Each point is the mean score of one version over the same 38 seeds. The exact numbers are in the folded table under the charts.
The solid orange line is cell by cell, and it is the main story. The two dashed lines are the two other ways to fill the map, which I tried after the last version. LLM (green squares): a language model draws the whole map in one answer. Batched (blue diamonds): Jev gets the questions for all the cells in one request. I ran both with the code of five versions (v1, v4, v7, v8 and v11). So each dashed point shows what that method did with what cell by cell knew at that version. Their own sections come later. Hover a version to read its values.
The numbers behind the charts
| Run | Structures | Paths | Reachable | Reads as a place | Rules hold | Routes to doors | Landmarks | Judge (GPT-5.2) |
|---|---|---|---|---|---|---|---|---|
| v1, cell by cell | 0.18 | 0.74 | 0.69 | 0.48 | – | 1.00 | 0.07 | 0.09 |
| v2, cell by cell | 0.39 | 0.53 | 0.75 | 0.43 | – | 0.77 | 0.61 | 0.24 |
| v3, cell by cell | 0.68 | 0.66 | 0.82 | 0.50 | – | 0.62 | 0.79 | 0.45 |
| v4, cell by cell | 0.82 | 0.80 | 0.77 | 0.50 | – | 0.66 | 0.66 | 0.48 |
| v5, cell by cell | 0.93 | 0.82 | 0.81 | 0.53 | – | 0.70 | 0.76 | 0.54 |
| v6, cell by cell | 0.95 | 0.81 | 0.81 | 0.51 | – | 0.60 | 0.68 | 0.48 |
| v7, cell by cell | 0.95 | 0.80 | 0.82 | 0.52 | 0.99 | 0.61 | 0.78 | 0.48 |
| v8, cell by cell | 0.99 | 0.75 | 0.84 | 0.89 | 0.99 | 0.78 | 0.77 | 0.65 |
| v9, cell by cell | 0.98 | 0.80 | 0.87 | 0.87 | 0.98 | 0.87 | 0.84 | 0.69 |
| v10, cell by cell | 0.98 | 0.78 | 0.87 | 0.88 | 0.99 | 0.86 | 0.83 | 0.71 |
| v11, cell by cell | 0.98 | 0.77 | 0.88 | 0.88 | 0.99 | 0.82 | 0.83 | 0.69 |
| v1-llm, LLM | 0.92 | 0.75 | 0.87 | 0.65 | – | 0.54 | 0.91 | 0.68 |
| v4-llm, LLM | 0.91 | 0.79 | 0.89 | 0.66 | – | 0.51 | 0.93 | 0.68 |
| v7-llm, LLM | 0.92 | 0.76 | 0.90 | 0.62 | 0.80 | 0.65 | 0.90 | 0.68 |
| v8-llm, LLM | 0.94 | 0.76 | 0.82 | 0.78 | 0.76 | 0.65 | 0.84 | 0.65 |
| v11-llm, LLM | 0.94 | 0.75 | 0.85 | 0.74 | 0.76 | 0.61 | 0.90 | 0.65 |
| v1-batch, batched | 0.58 | 0.61 | 0.79 | 0.45 | – | 1.00 | 0.21 | 0.20 |
| v4-batch, batched | 0.78 | 0.55 | 0.79 | 0.38 | – | 0.79 | 0.70 | 0.32 |
| v7-batch, batched | 0.64 | 0.57 | 0.80 | 0.39 | 0.94 | 0.64 | 0.47 | 0.25 |
| v8-batch, batched | 0.98 | 0.66 | 0.85 | 0.87 | 0.99 | 0.82 | 0.71 | 0.57 |
| v11-batch, batched | 0.97 | 0.47 | 0.87 | 0.91 | 0.83 | 0.78 | 0.66 | 0.57 |
The judge never sees the metrics, but it agrees with them. Its score goes from 0.09 at version 1 to 0.69 at version 11. It rises most at the same versions where the metrics rise most: versions 2 and 3 (0.09 to 0.45), and version 8, when code started to draw the rooms (0.48 to 0.65).
The table below shows one version per row: what changed, and the maps it drew. It shows one test case at a time; use the buttons to move to the next one. Under each version where LLM and batched ran, a folded note describes those runs. LLM and batched, at every stage has their numbers.
rules-village-clusteredmedieval village, clustered order, 8 by 8 cells · 1 of 38
paths-village-frontiermedieval village, frontier order, 8 by 8 cells · 2 of 38
story-village-frontiermedieval village, frontier order, 8 by 8 cells · 3 of 38
routes-village-spiralmedieval village, spiral order, 8 by 8 cells · 4 of 38
structures-village-spiralmedieval village, spiral order, 8 by 8 cells · 5 of 38
large-village-spiralmedieval village, spiral order, 16 by 16 cells · 6 of 38
paths-station-branchingspace station, branching order, 8 by 8 cells · 7 of 38
structures-station-frontierspace station, frontier order, 8 by 8 cells · 8 of 38
rules-station-randomspace station, random order, 8 by 8 cells · 9 of 38
reach-station-spiralspace station, spiral order, 8 by 8 cells · 10 of 38
story-station-spiralspace station, spiral order, 8 by 8 cells · 11 of 38
large-station-frontierspace station, frontier order, 16 by 16 cells · 12 of 38
routes-mansion-frontierhaunted mansion, frontier order, 8 by 8 cells · 13 of 38
reach-mansion-randomhaunted mansion, random order, 8 by 8 cells · 14 of 38
coherence-mansion-spiralhaunted mansion, spiral order, 8 by 8 cells · 15 of 38
story-mansion-spiralhaunted mansion, spiral order, 8 by 8 cells · 16 of 38
structures-mansion-spiralhaunted mansion, spiral order, 8 by 8 cells · 17 of 38
large-mansion-clusteredhaunted mansion, clustered order, 16 by 16 cells · 18 of 38
paths-market-clusteredmarket, clustered order, 8 by 8 cells · 19 of 38
routes-market-clusteredmarket, clustered order, 8 by 8 cells · 20 of 38
coherence-market-frontiermarket, frontier order, 8 by 8 cells · 21 of 38
rules-market-frontiermarket, frontier order, 8 by 8 cells · 22 of 38
reach-jungle-clusteredjungle temple, clustered order, 8 by 8 cells · 23 of 38
story-jungle-clusteredjungle temple, clustered order, 8 by 8 cells · 24 of 38
paths-jungle-frontierjungle temple, frontier order, 8 by 8 cells · 25 of 38
rules-jungle-spiraljungle temple, spiral order, 8 by 8 cells · 26 of 38
rules-desert-branchingdesert outpost, branching order, 8 by 8 cells · 27 of 38
reach-desert-frontierdesert outpost, frontier order, 8 by 8 cells · 28 of 38
story-desert-frontierdesert outpost, frontier order, 8 by 8 cells · 29 of 38
coherence-desert-spiraldesert outpost, spiral order, 8 by 8 cells · 30 of 38
structures-arctic-clusteredarctic base, clustered order, 8 by 8 cells · 31 of 38
coherence-arctic-frontierarctic base, frontier order, 8 by 8 cells · 32 of 38
reach-arctic-randomarctic base, random order, 8 by 8 cells · 33 of 38
routes-arctic-spiralarctic base, spiral order, 8 by 8 cells · 34 of 38
structures-cyberpunk-branchingcyberpunk block, branching order, 8 by 8 cells · 35 of 38
coherence-cyberpunk-clusteredcyberpunk block, clustered order, 8 by 8 cells · 36 of 38
routes-cyberpunk-frontiercyberpunk block, frontier order, 8 by 8 cells · 37 of 38
paths-cyberpunk-spiralcyberpunk block, spiral order, 8 by 8 cells · 38 of 38
| Version and change | Cell by cell | LLM | Batched |
|---|---|---|---|
v1 · baseline The starting point: Jev sees only the neighbouring tiles, so each map comes out almost all one tile. The judge gives 0.09. LLM and batched, with v1’s codev1-llm · LLM. Claude Sonnet 5 draws the whole map, with only what version 1 gave Jev: the vocabulary and four sentences of rules. There are no target shares and no plan. The judge gives 0.68, against 0.09 for cell by cell. Many more types are on the map (coverage 0.83 against 0.15), and 42% of the maps have a closed room, with no plan in the prompt. v1-batch · batched. Version 1's own questions, sent together. The 38 maps took 86 requests: every cell is offered every type, so an 8 by 8 map needs two requests. The judge gives 0.20, against 0.09. Version 1 failed because each cell copied its neighbours. In a batch, there are no neighbours to copy. | ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() | ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() | ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() |
v2 · confetti Jev now also gets a balance sheet: how much of each tile type the map has, against its target. More types appear (coverage, the share of types that reach the map, goes from 0.15 to 0.44), but as scattered single tiles. | ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() | not run | not run |
v3 · floods Jev now gets hints such as “a wall line reaches this cell, continue it”. It follows them too far: one 8 by 8 station is 58 of its 64 cells corridor. | ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() | not run | not run |
v4 · lines Code now checks each hint before it is sent. Walls form lines and closed shapes, but only about a quarter of the doors stand in a wall (0.26). LLM and batched, with v4’s codev4-llm · LLM. Now with the target shares, the lines and the patches, but still no plan. The judge gives 0.68, against 0.48 for cell by cell. Every specification is within 0.05 of v1-llm: the balance sheet and the lines taught the LLM almost nothing new. v4-batch · batched. Version 4's fix, a line that continues, needs the cells decided before, and a batch has none. Paths score 0.55 against 0.80, and only 5 of the 38 maps have a landmark. The judge gives 0.32, against 0.48. | ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() | ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() | ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() |
v5 · the model reads the sketch, a little Jev now sees the whole map, one letter per cell. It helps only a little, because one cell cannot see where a rectangle should end. | ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() | not run | not run |
v6 · less noise The random pick from Jev’s probabilities now drops options below a third of the best one. Ground forms patches, but rare tiles almost disappear (coverage 0.76 to 0.55). | ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() | not run | not run |
v7 · the local rules hold, the rooms do not Typed rules, such as “never next to X”, now remove forbidden tiles before Jev answers. The rules hold (0.99), but rooms still do not form. LLM and batched, with v7’s codev7-llm · LLM. The typed rules are now written in the prompt. They hold at 0.80, against 0.99 in cell by cell, where code offers only the types that the rules allow. The judge gives 0.68, against 0.48. v7-batch · batched. Code still applies the typed rules, so they hold at 0.94. But the neighbours they check are not there yet, and no door stands in a wall (0.00). The judge gives 0.25, against 0.48. | ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() | ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() | ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() |
v8 · rooms The blueprint: before Jev starts, code places each building as a rectangle with a door. Doors in walls go from 0.15 to 0.96. The price: fewer tile types (coverage 0.51). LLM and batched, with v8’s codev8-llm · LLM. The blueprint goes into the prompt, room by room. This is the version that moved the LLM most: “a place reads as a place” rises from 0.62 to 0.78, and 74% of the maps have a closed room, against 34% at v7. The judge gives 0.65, the same as cell by cell. v8-batch · batched. The blueprint is code, applied to each cell, so a batch gets all of it. Each cell is offered only the types of its part, so all 64 questions now fit in one request. Structures go from 0.64 to 0.98, and every door is in a wall. The judge gives 0.57, against 0.65. | ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() | ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() | ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() |
v9 · the world remembers what it lacks Jev now sees what the map still lacks, a landmark stops being offered once it is placed, and a route is suggested outside each door. Every landmark appears once (1.00). | ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() | not run | not run |
v10 · the cap that could not fire A ground tile that covers twice its target is no longer offered. This changed nothing: in the station, the flooding tile was the corridor, which was also the only floor. | ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() | not run | not run |
v11 · a floor that is not a road The station gets an open deck and the city a plaza, so corridors and streets become lines. The corridors are now short stubs, which is the next problem. LLM and batched, with v11’s codev11-llm · LLM. The same seeds, vocabularies and blueprint, drawn by Claude Sonnet 5 in one answer per map. A 16 by 16 map takes 6 seconds instead of 90, and the whole run costs the same. But only 47% of the maps have a closed room, against 92%, and the typed rules hold at 0.76 against 0.99. Its own section has the details. v11-batch · batched. Version 11’s questions, unchanged, sent together, in one request per map where it fits. That is 61 times fewer requests, at a third of the price. But each cell is decided without knowing what the other cells in the same request got: 181 cells sit next to a tile that their “never next to” rule forbids, and each of those broken pairs is inside one request, and one mansion has 39 staircases where the rules ask for one. Its own section has the details. | ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() | ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() | ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() |
rules-village-clusteredmedieval village, clustered order, 8 by 8 cells · 1 of 38
paths-village-frontiermedieval village, frontier order, 8 by 8 cells · 2 of 38
story-village-frontiermedieval village, frontier order, 8 by 8 cells · 3 of 38
routes-village-spiralmedieval village, spiral order, 8 by 8 cells · 4 of 38
structures-village-spiralmedieval village, spiral order, 8 by 8 cells · 5 of 38
large-village-spiralmedieval village, spiral order, 16 by 16 cells · 6 of 38
paths-station-branchingspace station, branching order, 8 by 8 cells · 7 of 38
structures-station-frontierspace station, frontier order, 8 by 8 cells · 8 of 38
rules-station-randomspace station, random order, 8 by 8 cells · 9 of 38
reach-station-spiralspace station, spiral order, 8 by 8 cells · 10 of 38
story-station-spiralspace station, spiral order, 8 by 8 cells · 11 of 38
large-station-frontierspace station, frontier order, 16 by 16 cells · 12 of 38
routes-mansion-frontierhaunted mansion, frontier order, 8 by 8 cells · 13 of 38
reach-mansion-randomhaunted mansion, random order, 8 by 8 cells · 14 of 38
coherence-mansion-spiralhaunted mansion, spiral order, 8 by 8 cells · 15 of 38
story-mansion-spiralhaunted mansion, spiral order, 8 by 8 cells · 16 of 38
structures-mansion-spiralhaunted mansion, spiral order, 8 by 8 cells · 17 of 38
large-mansion-clusteredhaunted mansion, clustered order, 16 by 16 cells · 18 of 38
paths-market-clusteredmarket, clustered order, 8 by 8 cells · 19 of 38
routes-market-clusteredmarket, clustered order, 8 by 8 cells · 20 of 38
coherence-market-frontiermarket, frontier order, 8 by 8 cells · 21 of 38
rules-market-frontiermarket, frontier order, 8 by 8 cells · 22 of 38
reach-jungle-clusteredjungle temple, clustered order, 8 by 8 cells · 23 of 38
story-jungle-clusteredjungle temple, clustered order, 8 by 8 cells · 24 of 38
paths-jungle-frontierjungle temple, frontier order, 8 by 8 cells · 25 of 38
rules-jungle-spiraljungle temple, spiral order, 8 by 8 cells · 26 of 38
rules-desert-branchingdesert outpost, branching order, 8 by 8 cells · 27 of 38
reach-desert-frontierdesert outpost, frontier order, 8 by 8 cells · 28 of 38
story-desert-frontierdesert outpost, frontier order, 8 by 8 cells · 29 of 38
coherence-desert-spiraldesert outpost, spiral order, 8 by 8 cells · 30 of 38
structures-arctic-clusteredarctic base, clustered order, 8 by 8 cells · 31 of 38
coherence-arctic-frontierarctic base, frontier order, 8 by 8 cells · 32 of 38
reach-arctic-randomarctic base, random order, 8 by 8 cells · 33 of 38
routes-arctic-spiralarctic base, spiral order, 8 by 8 cells · 34 of 38
structures-cyberpunk-branchingcyberpunk block, branching order, 8 by 8 cells · 35 of 38
coherence-cyberpunk-clusteredcyberpunk block, clustered order, 8 by 8 cells · 36 of 38
routes-cyberpunk-frontiercyberpunk block, frontier order, 8 by 8 cells · 37 of 38
paths-cyberpunk-spiralcyberpunk block, spiral order, 8 by 8 cells · 38 of 38
Version 8: the code draws the rooms
Version 8 deserves a closer look, because it is the version that took work away from the model.
Versions 1 to 7 all tried to get rooms from the model: more context, the whole map as text, less random sampling, typed placement rules. The doors still landed in open ground. The problem is that a model that decides one cell at a time cannot draw a rectangle. It never sees the whole shape that it is building. So it turns a wall line into a solid block, and it puts the door wherever the neighbours allow one.
So version 8 stops asking the model for shapes. Before the first cell is decided, code draws a plan of the map, called the blueprint. The blueprint is plain geometry, with no model involved:
- The vocabulary lists its buildings. For example, a cottage is a ring of cottage wall around cottage floor, with one cottage door.
- Code places each building as a rectangle, with a random size and number inside the limits that the vocabulary sets, and with empty cells between them. Together, the buildings never cover more than half the map.
- Each rectangle gets one door, unless the vocabulary marks it sealed. The door is in a wall, never in a corner, on the side that faces the middle of the map. So it opens onto the map, not off its edge.
- Every cell now has a part: a wall, a door, the inside of a named building, or outside all buildings.
The model still decides every cell, but the question is narrower. It is no longer “what goes here?” but “what goes in this wall?”, and the answer can be a window. An inside cell can be a floor or a chest. An outside cell can be anything that belongs outdoors. The blueprint comes from the test case’s seed, so a test case always gets the same blueprint.
No other change moved “a place reads as a place” this far. Doors in walls went from 0.15 to 0.96, 37 of the 38 maps got a room you can enter, and “a place reads as a place” went from 0.52 to 0.89. The price was variety. Coverage of the vocabulary fell to 0.51, because a cell inside a room is never offered the outdoor types.
The general lesson: give the model the decisions that need judgement, such as what stands in this room, who walks here, or what this place is for. Give code the decisions that are geometry.
Regressions, and what the loop caught
The loop also caught two problems that I would not have seen by looking at pictures.
A new metric can change old scores. Version 1 scored 0.99 on “paths form routes”. At first, that specification only checked whether path cells join up. In a map that is all corridor, every path cell touches another one, so the score was almost perfect. Later I added a metric that checks how much of the map the paths cover, and version 1 fell to 0.74.
A table of versions is only fair if every version is scored by the same rules. So every time I added a metric, I ran it on every old map. Every time I added a test case, I drew it again with the code of each old version. The charts therefore show today’s rules applied to the old maps.
A fix in one place can lower a score in another. Version 6 dropped the options below a third of the best one. So the generator picked Jev’s first choice more often, and an unlikely tile less often. Open ground stopped being full of odd single tiles, and “ground in patches” rose from 0.22 to 0.48. But the rare types were now almost never picked, so coverage fell from 0.76 to 0.55. One mansion test case shows both effects.
Version 11 made the same kind of trade somewhere else. It gave the space station an open deck as its ground, so the corridor shrank from half the map to a tenth. The corridors that remain are short stubs. In the same run, the path-share score (1 when paths cover 8% to 30% of the map) rose from 0.65 to 0.84. But path continuity (the share of path cells in the middle of a line, not at an end) fell from 0.73 to 0.58.
In the one map that I would have looked at, I would have missed both problems. The mean over 38 maps showed both.
The same maps, drawn by an LLM
That is where the loop stopped. The next two sections try two other ways to fill the map: the dashed lines in the charts. Both use the code of version 11 here, and a later section runs them with earlier code.
The first question: can one answer from a language model replace many small decisions from Jev? In the LLM method, Claude Sonnet 5 draws the whole map in one answer. Its prompt holds everything that cell by cell gives Jev: the tile types and their rules, the rooms of the blueprint, and the rules of a good map. The model answers with one letter per cell. If the answer has the wrong shape, the generator tries again, up to three attempts in all. I ran it on the same 38 test cases, as v11-llm.
One tip before the results: turn the model’s thinking off. With thinking on, the first call spent 68 seconds and $0.064 on an 8 by 8 village, and returned nothing.
The numbers behind the chart
| Specification | v11, cell by cell | v11-llm, LLM | Change |
|---|---|---|---|
| Structures | 0.98 | 0.94 | -0.04 |
| Paths | 0.77 | 0.75 | -0.02 |
| Reachable | 0.88 | 0.85 | -0.04 |
| Reads as a place | 0.88 | 0.74 | -0.14 |
| Rules hold | 0.99 | 0.76 | -0.23 |
| Routes to doors | 0.82 | 0.61 | -0.20 |
| Landmarks | 0.83 | 0.90 | +0.07 |
| Judge (GPT-5.2) | 0.69 | 0.65 | -0.04 |
The numbers behind the chart
| What | cell by cell | LLM |
|---|---|---|
| 8 × 8 map, cost | $0.005 | $0.006 |
| 16 × 16 map, cost | $0.023 | $0.010 |
| 8 × 8 map, time | 22 s | 4 s |
| 16 × 16 map, time | 90 s | 6 s |
| 38-map run, cost | about $0.26 | $0.25 |
What the results show:
- Maps: worse where structure matters. Only 47% of its maps have a closed room (a room with a wall all around it), against 92% for cell by cell. Only 60% of its doors are in a wall, against 87%. When the model writes 256 letters in one answer, it makes small slips: a wall one column off, a door in the grass. Cell by cell cannot slip like this, because code applies the blueprint to every cell.
- Rules: broken more often (0.76 against 0.99). Jev is almost never offered a tile that a rule forbids. A language model can write any letter.
- Landmarks: better (0.90 against 0.83). A model that sees the whole map places the one well and the one gate.
- The judge: a tie. It gives 0.65 against 0.69, a gap smaller than the noise between two runs (0.05).
- Time: much faster. A 16 by 16 map takes 6 seconds instead of 90, because 256 requests become one.
- Money: the same. The whole run cost $0.25 with the LLM and about $0.26 with cell by cell.
So which one should you use? Use the LLM for a quick map with every landmark in place; it is now an option in the generator’s AI Setup. Use cell by cell for rooms that close, doors in walls and rules that never break, and to watch the map fill. The LLM got most of this right on its first try, while cell by cell needed many rounds of tuning and a blueprint drawn by code. Price is not the difference, because both cost the same. The difference is control over each cell.
The same decisions, batched
The LLM is fast because it sends one request. It also gives up the typed decision, which is the part of the project that I care about most. So I tried a third way. Jev’s decisions endpoint (the web address that answers typed questions) accepts a list of questions, not only one. I kept every per-cell question exactly as it was, and sent them all in one request. I call this batched: an 8 by 8 map usually becomes one request with 64 questions, instead of 64 requests.
Every question is built by the same code as in cell by cell. The catch: all the answers come back together, so no question knows what the other cells in its request will get. Every question sees the same map, with empty cells where the rest of the request is. So the rules that look at other cells, such as “never next to a portrait” or “exactly one staircase”, have nothing to check.
One request holds at most about 65,000 tokens. To fit a whole 8 by 8 map, I moved the instruction that is the same for every cell into the shared part of the request, so it is sent only once. A 16 by 16 map still needs four or five requests, sent one after another.
I ran it on the same 38 test cases, as v11-batch. Jev answered all 3,008 questions, and none came back broken.
The numbers behind the chart
| Specification | v11, cell by cell | v11-batch, batched | Change |
|---|---|---|---|
| Structures | 0.98 | 0.97 | -0.01 |
| Paths | 0.77 | 0.47 | -0.30 |
| Reachable | 0.88 | 0.87 | -0.01 |
| Reads as a place | 0.88 | 0.91 | +0.03 |
| Rules hold | 0.99 | 0.83 | -0.16 |
| Routes to doors | 0.82 | 0.78 | -0.04 |
| Landmarks | 0.83 | 0.66 | -0.17 |
| Judge (GPT-5.2) | 0.69 | 0.57 | -0.12 |
The effect is clearest on one map: the same 16 by 16 mansion, from the same test case, drawn both ways.
The 39 staircases are not a bug. Each of those cells was asked, on its own, what goes in an inside cell of a mansion that has no staircase yet, and each one gave the same sensible answer. In cell by cell, only the first cell gets that answer, because after it the map has a staircase and code stops offering it. In a batch, no cell knows what the others chose. The same cause explains the scores that fell:
- Rules break between neighbours. 181 cells sit next to a tile that their “never next to” rule forbids, and every one of the 163 broken pairs is inside one request. In cell by cell, 11 cells do.
- Unique things appear many times. Seven maps have more than one of a type that should appear once. Cell by cell has none.
- Paths break. A path continues because the cell before it is a path. In a batch, that cell is still empty. Path continuity fell from 0.58 to 0.26.
- The map gets emptier. Similar questions get similar answers, so the most common tile wins nearly everywhere. The village below has its cottages, and nothing else.
What code decides did not change, because code applies it to every cell in both modes: the buildings and their doors come out the same. The maps even look tidier, but only because they repeat the same tile.
On speed and cost, batched is a clear win:
- Time: 26 times less waiting. The whole run took 40 seconds instead of 1,046, because 64 requests became one.
- Money: a third of the price. The run cost $0.086 instead of about $0.26. This saving is smaller than the time saving, because you pay for tokens, not for requests. Only the shared part (the instruction, the world, the map and its balance sheet) is now sent once; each question still carries its own options and rules.
The full comparison
| What the run did | v11, cell by cell | v11-batch, batched |
|---|---|---|
| Decisions | 3,008 | 3,008 |
| Requests to Jev | 3,008 | 49 answered, 1 refused for size |
| Input tokens | about 6.2 million | 2.04 million |
| Cost of the 38 maps | about $0.26 | $0.086 |
| Time drawing the 38 maps | 1,046 s | 40 s |
| One 8 × 8 map | 64 requests, 22.2 s, $0.005 | 1 request (one map took 2), 0.85 s, $0.0018 |
| One 16 × 16 map | 256 requests, 90.2 s, $0.023 | 4.3 requests, 3.4 s, $0.0072 |
| Fallback cells | 0 | 0 |
So Jev answers a whole map of typed questions in one request, much faster and cheaper. But the maps get worse in exactly the way the design predicts: every decision that needed a neighbour lost it. The judge gives 0.57, against 0.69 for cell by cell.
A better prompt would not fix this. To make batched work here, the fix would be to move more of what cells now learn from their neighbours into the blueprint, which is drawn before any cell is asked. That is the lesson of version 8 again.
LLM and batched, at every stage
Both experiments above used the code of version 11. So one question is left: did the LLM and batched improve along with cell by cell, or were they this good from the start? To find out, I ran both again with the code of four earlier versions: v1 (before any quality work), v4 (the lines), v7 (the typed rules) and v8 (the blueprint). These are the dashed lines in the charts at the top.
Each run gets only what its version knew: its own code, its own vocabulary, and only the rules that the version had. One rule is in every LLM prompt, because the v11 prompt has it: every walkable cell must be reachable.
A difference smaller than 0.05 is noise: a repeat of the v11 batched run landed within 0.05 of the first run on every score.
The numbers for every run
| Run | Structures | Paths | Reachable | Reads as a place | Rules hold | Routes to doors | Landmarks | Judge (GPT-5.2) |
|---|---|---|---|---|---|---|---|---|
| v1, cell by cell | 0.18 | 0.74 | 0.69 | 0.48 | – | 1.00 | 0.07 | 0.09 |
| v1-llm, LLM | 0.92 | 0.75 | 0.87 | 0.65 | – | 0.54 | 0.91 | 0.68 |
| v1-batch, batched | 0.58 | 0.61 | 0.79 | 0.45 | – | 1.00 | 0.21 | 0.20 |
| v4, cell by cell | 0.82 | 0.80 | 0.77 | 0.50 | – | 0.66 | 0.66 | 0.48 |
| v4-llm, LLM | 0.91 | 0.79 | 0.89 | 0.66 | – | 0.51 | 0.93 | 0.68 |
| v4-batch, batched | 0.78 | 0.55 | 0.79 | 0.38 | – | 0.79 | 0.70 | 0.32 |
| v7, cell by cell | 0.95 | 0.80 | 0.82 | 0.52 | 0.99 | 0.61 | 0.78 | 0.48 |
| v7-llm, LLM | 0.92 | 0.76 | 0.90 | 0.62 | 0.80 | 0.65 | 0.90 | 0.68 |
| v7-batch, batched | 0.64 | 0.57 | 0.80 | 0.39 | 0.94 | 0.64 | 0.47 | 0.25 |
| v8, cell by cell | 0.99 | 0.75 | 0.84 | 0.89 | 0.99 | 0.78 | 0.77 | 0.65 |
| v8-llm, LLM | 0.94 | 0.76 | 0.82 | 0.78 | 0.76 | 0.65 | 0.84 | 0.65 |
| v8-batch, batched | 0.98 | 0.66 | 0.85 | 0.87 | 0.99 | 0.82 | 0.71 | 0.57 |
| v11, cell by cell | 0.98 | 0.77 | 0.88 | 0.88 | 0.99 | 0.82 | 0.83 | 0.69 |
| v11-llm, LLM | 0.94 | 0.75 | 0.85 | 0.74 | 0.76 | 0.61 | 0.90 | 0.65 |
| v11-batch, batched | 0.97 | 0.47 | 0.87 | 0.91 | 0.83 | 0.78 | 0.66 | 0.57 |
What the runs show:
- The LLM was good from the start. With only what version 1 knew, the judge gives its maps 0.68; with what version 11 knew, 0.65. Cell by cell first reached that level at version 9. A model that writes the whole map at once sees what one cell cannot: every type used, no type flooding the map, the landmarks in place.
- Only rules about where things go helped the LLM. The typed rules (v7) and the blueprint (v8) moved some of its scores, but the judge barely moved (0.68 to 0.65).
- Batched jumped at one version. The judge gives it 0.20 to 0.32 before version 8, then 0.57. The fixes of versions 2 to 7 all read cells that are already filled, and in a batch those cells are empty. The blueprint is code applied to each cell, so a batch gets all of it.
- At version 1, batched beat cell by cell (0.20 against 0.09). Cell by cell failed there because each cell copied its neighbours, and in a batch there are none to copy. It was still not a good map.
The numbers behind the chart
| Run | Seconds per map | Judge (GPT-5.2) |
|---|---|---|
| v1, cell by cell | 27.2 s | 0.09 |
| v2, cell by cell | 26.8 s | 0.24 |
| v3, cell by cell | 27.2 s | 0.45 |
| v4, cell by cell | 26.3 s | 0.48 |
| v5, cell by cell | 27.1 s | 0.54 |
| v6, cell by cell | 26.8 s | 0.48 |
| v7, cell by cell | 24.9 s | 0.48 |
| v8, cell by cell | 24.6 s | 0.65 |
| v9, cell by cell | 23.7 s | 0.69 |
| v10, cell by cell | 25.4 s | 0.71 |
| v11, cell by cell | 27.5 s | 0.69 |
| v1-llm, LLM | 6.2 s | 0.68 |
| v4-llm, LLM | 5.3 s | 0.68 |
| v7-llm, LLM | 4.9 s | 0.68 |
| v8-llm, LLM | 5.6 s | 0.65 |
| v11-llm, LLM | 3.9 s | 0.65 |
| v1-batch, batched | 1.4 s | 0.20 |
| v4-batch, batched | 1.7 s | 0.32 |
| v7-batch, batched | 1.5 s | 0.25 |
| v8-batch, batched | 0.9 s | 0.57 |
| v11-batch, batched | 1.1 s | 0.57 |
For the per-cell design, the honest lesson is this. Many small decisions that do not see each other are harder to get right than one decision about the whole map. They need more help from code, and much more tuning. What cell by cell gets in return is control: code can check and limit every single cell. These are five runs of each, on one generator, so read them as measurements, not laws.
Running the loop with an agent
Most of the versions were written by an agent, Claude Code, in a loop that I set up and then mostly watched. One turn went like this:
1. Read the last version's scores, and the maps where it lost.
2. Write down one theory, such as:
- "the model does not see enough of the map"
- "the sampling adds noise"
- "the rules are prose, so the model treats them as advice"
3. Change the code, with tests.
4. Commit, so the version in Galtea records exactly what was measured.
5. Run the 38 test cases.
6. Compare with the last version.
7. Write the result into the project's README, one paragraph per
version, so the next turn starts with the memory of this one.
My part was smaller, and I think it was the part that mattered:
- I chose which theory to test first.
- I added specifications. When I saw a failure that the numbers missed, I asked for a specification for it. That is how three specifications became seven.
- I set the budget. 8 by 8 grids for most test cases, and only three 16 by 16 maps, because a room needs space, but a large map costs about four times a small one.
- I said when to stop: when the agent could no longer name a change that it was sure would move a number, the loop had done its work, for now.
The agent also broke things:
- It committed a version with a failing test, because a shell pipe hid the exit code. The continuous integration check caught it.
- It ran two evaluations while it was editing the modules that they used. The wrong maps showed it.
Each one became a written rule for the next turn. The loop does not promise that every version is better. It promises that you will know.
What I have not tried
Both ways of using Jev in this post, cell by cell and batched, are flat. There is one level of decisions, and every decision is about one cell:
world → cell, cell, cell, cell, … 3,008 times.
Every question has the same size and the same shape, and they go either one at a time or all at once. There is a third shape that I have not built. In it, the decisions form a tree instead of a list, and each layer asks about what the layer above it decided:
world → regions → structures → rooms → objects.
- Decide the broad areas first. This corner is the docks, that one is the old quarter, this strip between them is farmland. A handful of typed decisions for the whole map.
- Then, for each area, decide what stands in it. The docks get three warehouses and a crane; the old quarter gets a chapel and six houses. The question is asked once per area, and the options depend on which area it is.
- Then, for each structure, decide its inside. A warehouse is one long hall; a house is a kitchen and two rooms.
- Then, for each room, decide what is in it. Only here does the question get down to one tile, and by then it is about a room that already knows its purpose.
The point is not fewer decisions. The point is that later decisions depend on earlier ones by design, instead of hoping that the model reads a map sketch and works it out. A cell in the kitchen of a house in the old quarter is a much narrower question than a cell at (7, 12). And the model decided that narrowing itself; nobody hard-coded it. This project keeps pointing the same way: the blueprint in version 8 is one layer of such a tree, written by hand in code, and it moved more numbers than any prompt change.
It could suit a decision model well. Each step is still a small typed choice with a few options, which is what these models are fast and cheap at. The structure and the dependencies live in the shape of the tree, not inside one long prompt. It may also fail for reasons that I cannot see yet. One wrong early branch would ruin a whole district, while the flat design fails one cell at a time.
I have not built it, and I have not measured it. There are no numbers for it in this post, in the version table, or in Galtea. It is the next experiment I want to run. Until it goes through the same 38 seeds, it is only an idea that sounds good, and this project taught me not to trust those.
I have not compared the fill orders either. Cell by cell can visit the cells in five orders: a spiral from the centre, frontier growth, clustered patches, a branching walk, or pure random. Every version ran all five, because each test case uses one: 12 spiral, 12 frontier, 8 clustered, 3 branching and 3 random. But each order ran on different places, so the scores cannot tell the orders apart. A fair comparison would draw the same places in every order.
What is still wrong
The corridors are stubs. A route between two doors is a shape, and shapes are the plan's job, so the next version should draw the paths the way it draws the rooms. Rooms are often empty, because the state says what the world lacks, but not what a kitchen holds. Free-standing barriers, such as trees, dunes and pipes, still split some maps in two. And rooms never share a wall, which real streets do all the time. Each of these problems has a number now, and that is the point.
You can try the generator with your own OpenRouter key, read the code and the roadmap, or look at every map of every version, which are committed to the repository next to their scores.
The numbers
Mean score per specification over the same 38 seeds, with every metric as it stands today. A specification cell is highlighted when it moved by 0.05 or more against the previous version; the judge column is not highlighted. The judge column is GPT-5.2 answering "how much does this map read as the setting?" on every map, from 0 to 1.
| Version | Barriers form structures | Paths form routes | Reachable and playable | A place reads as a place | The vocabulary's rules hold | Routes lead to doors | Landmarks and story | Judge |
|---|---|---|---|---|---|---|---|---|
| v1 | 0.18 | 0.74 | 0.69 | 0.48 | – | 1.00 | 0.07 | 0.09 |
| v2 | 0.39 | 0.53 | 0.75 | 0.43 | – | 0.77 | 0.61 | 0.24 |
| v3 | 0.68 | 0.66 | 0.82 | 0.50 | – | 0.62 | 0.79 | 0.45 |
| v4 | 0.82 | 0.80 | 0.77 | 0.50 | – | 0.66 | 0.66 | 0.48 |
| v5 | 0.93 | 0.82 | 0.81 | 0.53 | – | 0.70 | 0.76 | 0.54 |
| v6 | 0.95 | 0.81 | 0.81 | 0.51 | – | 0.60 | 0.68 | 0.48 |
| v7 | 0.95 | 0.80 | 0.82 | 0.52 | 0.99 | 0.61 | 0.78 | 0.48 |
| v8 | 0.99 | 0.75 | 0.84 | 0.89 | 0.99 | 0.78 | 0.77 | 0.65 |
| v9 | 0.98 | 0.80 | 0.87 | 0.87 | 0.98 | 0.87 | 0.84 | 0.69 |
| v10 | 0.98 | 0.78 | 0.87 | 0.88 | 0.99 | 0.86 | 0.83 | 0.71 |
| v11 | 0.98 | 0.77 | 0.88 | 0.88 | 0.99 | 0.82 | 0.83 | 0.69 |
“The vocabulary’s rules hold” has no score before v7: the typed rules it checks came with v7, and the earlier numbers were a later rescore of 5 to 23 maps against rules those versions never had. v1's routes score is a mean over few maps, because v1 drew no doors, so only “a route reaches the edge” was scored, on the maps that had a route. The judge was run on the older versions after the fact, at the time of writing, so that this column is complete.















































































































































































































































































































































































































































































































































































































































































































































































































