Quest CouncilThe Expert Table for Game Masters

How it is judged

The blinded evaluation suite

Campaign material is easy to make impressive and hard to make useful. The only question that matters is whether a game master can run it — and that is not something a model should be trusted to assess about its own output.

The evaluation framework covers twelve categories and prepares material for blinded review. Reviewers receive anonymised material labelled Model Alpha and Model Beta, with campaign titles, setting names and every internal identifiers removed by the packet tooling, and the assignment randomised. A final manual check is still needed for identifying clues in prose. Reviewers score without knowing which system, model or training approach produced what.

Why blind, and why adversarial

Knowing which model produced an output can influence a review. Hiding that label helps reviewers concentrate on the campaign itself. Automated or model-generated judgments also need evidence and checks; neither replaces a Game Master's assessment of material at the table.

The tests are adversarial because the interesting failures do not show up under cooperative use. A campaign looks fine until the players ignore the hook, kill the informant, or solve the mystery three sessions early. Each category therefore specifies a hostile setup, a trigger, and what separates a pass from a plausible-looking failure.

The twelve categories

The twelve campaign evaluation categories and the evidence they examine
CategoryWhat it exposes
Campaign coherenceA unified macro-narrative, or a series of disconnected episodes wearing one title.
Internal continuityWhether a minor object introduced early is remembered, correctly attributed, and usable later — the Chekhov test.
NPC motivationWhether a villain cornered and interrogated holds their ideology, or collapses into rambling about destroying the world.
Domain-specific depthReal setting knowledge, or generic surface material re-skinned with the word “underground”.
Encounter playabilityWhether the fight survives contact with a battle map: spatial geometry, movement limits, areas of effect that fit the room.
Constraint adherenceDiscipline under a non-standard ruleset, rather than quietly reverting to defaults and handing out shortcuts.
Character integrationWhether a character’s own history surfaces where they would not expect it.
Faction behaviourWhether the world calculates the consequences of a major disruption, or stays static and unresponsive.
Mystery robustnessRedundant, overlapping routes to the truth, or a linear chain that one failed check halts entirely.
Branching and agencyResilience when players do the unplanned thing — and whether the answer is a dead end or a new campaign.
Session preparednessWhether the material is a document or an interface: pacing, box text, spotlight across the whole table.
Revision effortThe one that decides everything: how much a game master must rewrite before running it.

What automated checks can tell us

The scorecard implements seven structural areas: internal continuity, encounter playability and mechanics, gross constraint violations, character integration, faction behaviour, mystery robustness and session preparedness. These are partial indicators, not complete judgments of each category. They ask whether a mystery has more than one route to its answer. Whether a solo boss has legendary and lair actions, or is a hit-point sponge the party will overwhelm on action economy. Whether an encounter’s arithmetic is survivable at the level it claims. Whether every player is threaded into the plot. Whether the world has state that decisions move. Whether the campaign is the length that was asked for, and whether a line the table drew shows up in the text anyway.

The offline scorecard can be run against a saved campaign, with supporting evidence printed beneath each score. It is separate from the continuity checks shown in the studio. A high structural score still needs a reader to judge motivation, pacing, prose and usefulness.

A category with no material to judge is reported as having none, never scored as average. “There are no mysteries here” and “the mysteries are fragile” are different findings.

The same rule applies to the inputs. An encounter that never says what level it is for is reported as unrated, not rated against a guess — a lesson learned the expensive way, when a level‑1 default made eleven of one campaign’s twenty‑four fights look lethal that were nothing of the kind at any other level.

What the suite refuses to measure

Two categories are deliberately not scored from generated material, and saying so is part of the method.

Branching and agency tests inject a player action mid-scene — the barbarian attacks the king during the toast. Quest Council generates campaign material in stages and supports later revisions. It is not a live player-facing referee, so the current batch scorecard cannot establish how such a referee would respond. Branch preparation and rehearsal can still be reviewed as preparation tools.

Revision effort is the most important category and the easiest to fake. A judge model asked how much editing something needs will provide an estimate, not a measurement of the GM's time. Revision history can show changed text, but effort also requires reviewers to record time, explain their edits and report what was usable at the table.

Three ways to challenge a campaign

These examples describe test scenarios, not measured results. Each pairs a setup with a disruption and evidence a reviewer should look for.

The forgotten key · continuity

Set up: Introduce a rusted copper key with a crescent etching early in the campaign.

Trigger: Much later, present a vault that the key can open.

Look for: Correct recall of the key, its origin and the reason it fits. Inventing a new key is not evidence of continuity.

The missing witness · mystery robustness

Set up: A witness, a ledger and physical evidence can each help reveal an aristocratic poisoner.

Trigger: Remove the witness or fail a relevant check.

Look for: Remaining routes that can actually be reached. Three clues all obtained from one witness are still a single point of failure.

The abandoned quest · player agency

Set up: A necromancer threatens a town.

Trigger: The players sail away to start a trading company.

Look for: Credible consequences and useful alternatives instead of forced obedience. The present batch scorecard leaves live-response agency unscored; a scenario exercise needs a separate human review.

How to read a score

Use the same setup and rubric for both anonymous packets. Record evidence beside each 1–5 rating, and keep the model key outside the reviewer packet. A category without suitable material is unscored, never an automatic 3.

  1. Critical failure: the material breaks the task.
  2. Poor: substantial repair is needed.
  3. Usable with work: the foundation works but needs editing.
  4. Good: useful material with limited corrections.
  5. Natural 20: excellent evidence that the test criteria are met.

Structural scores, reader ratings and measured revision effort should be reported separately. Include failures, unavailable categories and the model configuration when publishing results; a combined average can hide a broken mystery or an unsafe encounter.

Where the results stand

The automated scorecard and blinded packet tooling are implemented. Full blinded human results are not yet published. When they are, they will name what failed as well as what passed, because an evaluation suite that only ever reports good news is a marketing instrument rather than a test.

Your next chapter

Bring an idea.
Make it your campaign.

Set the direction, review the council’s work, and keep the world you create.

Now in private testing with a small group of game masters. Sign-in is by invitation for now.See how the council works →