Section 6 of 14

Training judges

Four learned judges run the pipeline. The location judge rates a place before any color is chosen. The wallpaper judge rates a finished candidate picture. The gallery judge orders pictures the wallpaper judge has already passed. The palette judge chooses palettes rather than rating quality, and it's covered in color palettes. The first two answer on the same four-point scale and are trained the same general way, but they're separate networks fitted to separate ratings.

The rating scale

Every judge learns from ratings I made by hand on one four-point scale. A 1 is junk: empty, black, or noise. A 2 is unremarkable. A 3 is good wallpaper material, and a 4 is exceptional. I rated according to my own taste in desktop wallpapers. Fine-tuning a judge toward someone else's taste would be straightforward, and the labeled datasets are published for that.

A location rating asks whether a place could make a good wallpaper. So I made it from two renders of the same frame, one in a neutral palette and one in a vivid palette that shows what the geometry can become, and the rating answers for the geometry. A wallpaper rating asks whether this exact picture is worth keeping: its mode, its settings, and its palette. The corpus holds ratings for more than twelve thousand locations, about six thousand smooth renders, and about five thousand renders in the other modes.

One location drawn twice side by side: the same frame in the neutral palette, and in a vivid one.
One location as it is rated: the same frame in the neutral palette, and in a vivid one. The rating answers for the geometry the two share rather than for either picture.

What the judge sees

The location judge scores the small render the walk's gates already examined, so no extra render is drawn for scoring. It's trained across that size and larger ones, so its score means the same thing at each. Training also varies the palette. Each rated location is cached with renders in the neutral palette the judge sees in production and in palettes drawn from a tracked pool, and a further set of palettes is held out of that pool entirely so a finished judge can be read on colorings it has never seen. The goal isn't a judge that's blind to color, but one that doesn't mistake a palette for the place.

The same location at three sizes, drawn to scale: a small thumbnail, a mid-size render, and a large one over three times the width of the thumbnail.
One location at three sizes, drawn to scale. The smallest is what the walk scores, and the largest is what a person rated.

The wallpaper judge scores the candidate picture at 640 by 360. Nothing is judged at wallpaper size; a full-resolution render is what a passing score buys. Its colors are never varied in training, because for this judge the coloring is part of what's rated. One network covers the smooth rendering and every other mode in rendering modes. Sharing the ratings helps the scarcer collection, and giving each kind its own output layer bought nothing.

Network design and training

Each judge is a MobileNetV4 network pretrained on ImageNet and fine-tuned on my ratings. The location judge uses a medium variant, since larger networks bought nothing. The wallpaper judge uses a much smaller one, because the non-smooth modes have fewer ratings and a larger network trained on them came out overconfident.

A rating is ordered, and a 3 is closer to a 4 than a 1 is, so the judges are trained by ordinal regression instead of as plain classifiers. The network answers is this at least a 2, then at least a 3, then at least a 4, each conditional on the one before. That guarantees nothing comes out more likely to be a 4 than to be at least a 3. The pipeline reads two of the answers: P(≥3), the chance I'd call it good, and P(≥4), the chance I'd call it exceptional.

From a score to a decision

A judge returns a probability, and the pipeline acts on bars. Where a bar comes from matters, because a judge's numbers shift every time it's retrained.

The location judge has three bars. Above the admission bar a frame is a find. Between the expansion bar and the admission bar it's a stepping stone the walk may explore through. Above the exceptional bar, read on P(≥4), a find counts as exceptional.

Four frames from one fractal family, in score order, each labeled with the decision its score buys: refused, expandable, a find, and an exceptional find.
One family, one frame per outcome the location judge's score decides. Below the expansion bar the frame is refused; between the two bars the walk may explore through it but does not book it as a find; above the admission bar it is booked as a find; above the exceptional bar it counts as exceptional.

After a retrain, I restate each location bar by volume: the new bar passes the same fraction of a fixed reference pool as the old one did. No ratings are read, so it keeps the same amount flowing through, not the same frames.

A candidate picture enters the pool the gallery is chosen from only if the wallpaper judge gives it at least even odds of a 4. Above that, the wallpaper judge is a gate and not a ranker: it says a picture is worth keeping, but not which of two keepers is better. That ordering comes from the gallery judge, trained almost entirely on pictures that had already cleared the gate, and a gallery seats nothing below its bar. The gallery judge had to be a whole new network. Reusing the wallpaper judge's features and retraining only its output layer performed at chance, so what that judge learned about telling keepers from junk says little about telling a good keeper from a great one. Gallery curation covers how the seats are chosen.

Evaluation and its limits

The judges learn my ratings, so the test is whether a judge agrees with me on pictures it has never seen. I hold those out by neighborhood rather than by frame, because nearby frames are near-duplicates and a judge could otherwise pass by memorizing a neighbor.

What I want is a judge that learns what makes a frame good, in composition, framing, and density, well enough to recognize it somewhere new. The judges do this well enough for the search to work. The limit comes from how I gathered the ratings: rate a batch, retrain, use the new judge to find more, and rate that. A judge trained on what earlier judges found is surest where the search has already been, and a high score on material it picked out is partly the judge agreeing with itself. The walk reserves part of every batch for untouched roots, and I rated batches drawn on purpose from under-seen modes, colors, and score bands. Those help, but the ratings are still a selected population. I think of the judges as instruments tied to my ratings and one score scale, not as measures of whether a fractal is beautiful.