A Simulation Engine, Productized as a Diagnostic Service

A simulation engine that models bounded-rationality play across tabletop games and reports which single change to make, rather than telling a designer to rebalance everything. One engine serves several games through a shared measurement contract. The part worth showing is not the engine: it was packaged into a sellable service with intake, qualification, scope, standardized reporting and pricing, instead of being left as an internal tool. The service is launching now, under its own brand. It sharpens human playtesting; it does not replace it.

Why a game feeling off is hard to act on

A designer usually knows something is wrong long before they can say what. The gap between that feeling and a specific change is where projects stall, and it is a measurement problem before it is a design problem.

Playtest data is thin and expensive

Getting a group around a table produces a handful of data points and costs an evening. Any question needing many plays to answer, which is most balance questions, gets answered by intuition instead, because the alternative is a schedule nobody has.

Rebalancing everything is not a plan

Tools that flag every anomaly hand a designer a list of forty things. Changing forty things at once means the next round of testing cannot attribute an improvement to any of them. The useful output is not the list; it is which item on it to touch first.

Perfect play is not how anyone plays

A simulation assuming optimal decisions measures a game nobody is sitting down to. Real players miss lines, misjudge timing, and follow habits. Modelling bounded rationality rather than perfect play is what makes the output resemble the table it describes.

An internal tool is not a service

The engine existed and worked before any of this was sellable. What was missing was everything around it: who qualifies, what a client sends, what they get back, what it costs, and what happens when the answer is that we cannot help.

The delivery sequence, and what changes hands at each step

This is the service as a client experiences it. Each step names what they provide and what they receive, because the common failure in productizing a capability is a delivery process where the client cannot tell what is expected of them.

What a client provides and receives at each stepThe service delivery sequence, from qualifying the question through to a single recommended change and what to watch for.1Qualify thequestionA no costs nothing2Scope and fixedpriceFrom a publishedrate3MaterialscollectedChecked completefirst4The engine runsNothing partial isreported5StandardizedreportSame sections, sameorder6One changerecommendedWith what to watchfor
A re-run measures the revised game against the same contract.
StepWhat the client providesWhat happensWhat the client receives
QualifyA short description of the game, its state, and the specific complaintWe check whether the game is in scope and whether the question is one a simulation can answerA yes or a no with the reason. A no costs nothing and happens often enough to be worth stating.
Scope and quoteConfirmation of the question to be answeredThe run is scoped against standard packages and priced from a published rate, not estimated per clientA written scope and a fixed price before any work begins
Collect materialsRules, components, and card or unit data in a defined submission formatMaterials are checked for completeness before anything runs, since a partial data set produces a confident wrong answerA list of what is missing, if anything, before the clock starts
RunNothing; this step is oursThe engine plays the game repeatedly under bounded-rationality assumptions, measured against the shared contractNothing yet, deliberately. Partial results in flight are more misleading than none.
ReportNothingResults are written into a standardized report with the same sections in the same order every timeA report a designer can read without a statistics background
RecommendA conversation, and their own judgement about the designWe identify the single change with the largest effect on the complaint brought to usOne recommended change, what it is expected to do, and what to watch for when it goes to the table
Re-runThe revised game data, if they want a second passThe same measurement run again against the same contractA before and after on identical terms, which is the only comparison worth making
The honest boundary, stated plainly: this sharpens human playtesting. It does not replace it. The output is a hypothesis with evidence behind it and a specific thing to watch for, which is what a playtest group should be handed before their evening rather than after.

One engine, several games, and the contract that lets that work

A shared measurement contract

Each game is mapped onto the same set of measurements before anything runs. That contract is what lets one engine serve several different games rather than becoming a bespoke build per title, and it is why a re-run is comparable to the original.

Play modelled the way people actually play

The engine models bounded-rationality play rather than optimal play. Decisions are made under limited information and limited lookahead, which is what produces results resembling a real table. How that is modelled is not published.

One change, not a rebalance

The reporting is built to end in a single recommendation. A designer receives the change with the largest expected effect on the complaint they brought, along with what to watch for. That constraint shaped the analysis more than any other decision.

The service around the engine

Intake, qualification, scoping, a fixed price from a published rate, a defined submission format, a standardized report, and a stated boundary. This is what turned a working tool into something a stranger can buy, and the part most technical teams never build.

Qualification includes saying no

Some games are out of scope and some questions are not answerable by simulation at all. Those are declined at the qualifying step, at no cost. A service that qualifies nobody out will eventually deliver a report to somebody it could not help.

It runs under its own brand

The service operates under a separate brand with its own language rules and its own audience, rather than as a line item under a consultancy. Game designers are not buying automation consulting.

The boundary, kept in front

It does not guarantee a balanced game, it does not replace playtesting, and there is no one-click anything. Balance is a design judgement made by a designer. What the service supplies is measurement and a prioritised hypothesis, which is a smaller claim and a more useful one.

What productizing this actually required

Is this a fit for your business?

A good fit when

  • You have an internal tool that works and no idea how to sell access to it
  • The technical capability is proven and the delivery process around it does not exist
  • You need standardized reporting because bespoke write-ups per client will not scale
  • You want a service with a stated boundary rather than an offer that promises everything

Probably not a fit when

  • You want the engine itself licensed as software you run yourself
  • You are looking for a tool that will decide your design questions for you
  • You want marketing for an offer whose delivery process is not defined yet

What to have ready

  • An honest account of what your tool does well and where it stops
  • The questions clients actually ask you, in their words rather than yours
  • A willingness to define who you will decline, the hardest part of productizing

Questions we get asked

Does this replace playtesting?

No, and the service says so in its own materials. Simulation measures what a rules system does across many more plays than a group can sit through. It cannot tell you whether the game is fun, whether the theme lands, or whether a rule reads clearly. It is meant to tell a playtest group what to look at.

Will it guarantee my game is balanced?

No. Nothing can, and an offer that says otherwise is selling certainty it does not have. What you get is a measurement of how the game behaves under modelled play, and one prioritised change with a stated expected effect. Whether it is right for your design is your call.

How does one engine handle different games?

Every game is mapped onto a shared measurement contract before anything runs. That mapping is the work, and it is what keeps the engine general instead of a bespoke build per title. The details of the contract, the behavioural model, and the scoring are not published.

What if my game is not a fit?

You find out at the qualifying step and it costs you nothing. Some games are outside what the engine handles and some questions cannot be answered by simulation. Declining those quickly is better than producing a report that hedges.

Related


Tell us what the process looks like now and we will map what a system would need to do. No obligation, and you keep the map either way.

Scroll to Top