Method Unchanged since volume four Three abandoned measures listed
The Map Room emblem The Map Room Games and getting lost
How the scoring works

Every figure on this site, and how it is produced

The scores here measure one narrow thing badly rather than everything vaguely. This page explains exactly what is counted, how, and the four places where the method is weakest.

48 games logged · six volumes · method frozen since volume four

What the score covers

The number under each review is a navigation score and nothing else. It does not account for combat, writing, performance, art direction or value. A game can be excellent and score badly here, and Marker Fields at 6.1 is exactly that case: a good game that fails the single test this publication applies.

The score is composed from four inputs, weighted as below. The weights have not changed since volume four and will not change mid-volume, because comparing games scored under different weights is worthless.

InputWeightHow it is produced
Recall45%Drawing test one week after finishing, scored against the real layout
Unguided travel share25%Distance travelled without a waypoint active, logged per journey
Landmark usability20%Count of structures actually used as bearings, per region
Recovery quality10%What happens when lost: whether the world offers a way back

The drawing test

Seven days after finishing a game, with no reference material and no return to the game, I draw the world from memory on a single sheet. The sketch is then compared against the actual map and scored coarsely: full, partial, or none.

Full means the major regions are in the right relationship to each other and the connections between them are correct. Partial means the regions are recognisable but the connections are guessed. None means individual places are remembered without any layout holding them together — four dots on a page, in the worst case logged so far.

The scale is deliberately coarse because a finer one would imply a precision the test does not have. There is no meaningful difference between a sketch scored 71 and one scored 76, and pretending otherwise would make the figures look more scientific than they are.

Every sketch is kept. None are redrawn, and none are improved after comparison. The one I most wanted to redo was Lantern Deep, whose upper galleries I had genuinely learned and simply failed to recall on the day.

Logging travel

Sessions are timed with a stopwatch rather than trusted to a platform counter, because platform counters include menus, alt-tabs and the twenty minutes a game spends open while I make tea. Each journey — any deliberate movement from one place to another — is recorded as guided or unguided at the moment it begins.

Guided means a waypoint, path line, compass marker or escort was active for the majority of the journey. Unguided means I set off knowing or guessing the way. Mixed journeys are split at the point the marker went on or off, which is fiddly and the main reason a forty-hour game takes three weeks to review here.

Rules I hold to

  • Every game is bought at retail with my own money, at full price, on release or later.
  • No guides, wikis, maps, route videos or community resources during a playthrough.
  • Default settings for the first playthrough, so the score reflects what ships.
  • If a setting materially changes navigation, it is tested separately and reported, as with the buried bearing option in Marker Fields.
  • A game is only scored if it was played to completion or to the point where the world stopped teaching anything new.

Where this method is weak

Four weaknesses, all real and none of them fixable without abandoning the approach.

It is one person's memory. Recall varies between people enormously, and a test built on a single memory measures that memory as much as it measures the game. The defence is consistency rather than validity: the same memory, tested the same way, across forty-eight games.

It rewards small worlds indirectly. A world with four regions is easier to recall correctly than one with twelve, even if the twelve are better designed. I have not found a fair correction for this, and the abandoned size-versus-recall note on the notes page is where the attempt failed.

It cannot separate navigation from interest. A world I enjoyed being in gets more of my attention, and attention improves recall. Some part of every recall score is enthusiasm rather than design quality, and I cannot say how much.

It punishes worlds designed to be disorienting. A game whose intent is to make you lost forever will score badly here despite succeeding, which is a category error the scale cannot express. Where this applies the review says so in the text, which is a patch rather than a solution.

Abandoned measures

Three things were counted for a while and then dropped, all listed here because a method page that only describes what survived is a sales pitch.

Fast travel usage was logged for three volumes and predicted nothing once fast travel restricted to landmarks was separated from unrestricted fast travel. World area in square kilometres was logged for two volumes and predicted nothing at all. Time-to-first-map-screen was an early favourite and turned out to be a proxy for genre rather than for anything about navigation.