Looking for Translation Aids? Start here
Before you start the day β before email, before decisions, before demands β you play a short, familiar game. Not to win. Not to improve. To find out what kind of brain and body you woke up with today.
The score is not the point. The score is a signal.
This page explains how the Calibration Project works, what it measures, and how you can adapt the idea for yourself.
A daily cognitive calibration system using Mini Metro, AI analysis, and body-state context to work out which kinds of tasks are realistic, safe, and sustainable on any given day. This project is not about productivity or score-chasing. It is about understanding access to capability.Β
β―οΈThe rough early guide
Before I started using AI to analyse the maps, scores, screenshots, and gameplay patterns, I developed a rough guide from experience.
Top 10%: go for it.
Complex thinking is probably available.
Top 20%: proceed with caution.
Capacity exists, but hidden friction may matter.
Top 30%: go back to bed and try again after a snooze, if that is an option.
The system may need recovery before demands are added.
That early guide was crude, but it was surprisingly useful.
It helped me stop treating every morning as if I had woken up with the same resources.
I had not.
π€What AI changed
When I started using AI in autumn 2025, I wondered whether it could do more than record the score.
Could it compare my score with my baseline?
Could it look at the map and gameplay pattern?
Could it read the difference between βlow capacityβ and βhigh capacity but poorly synchronisedβ?
Could it help me choose tasks that matched my actual state?
After a lot of experimenting, the answer became yes.
The system can now look at the score, mode, map, rank, screenshots, gifs, sleep, food, hydration, pain, emotional state, plans for the day, and recent patterns.
It then helps classify the day.
That classification does not tell me what I am βallowedβ to do.
It helps me avoid wasting energy trying to force the wrong kind of task through the wrong kind of nervous system.
(There's a free script you can edit at the bottom of the page β¬οΈ)
This is a personal longitudinal self-observation project, not a medical diagnostic tool, but it produces practical daily evidence about fluctuating access to executive function.Β
πWhy a game?
A useful calibration task needs to be:
Familiar. If the task is new or exciting, novelty distorts the result. You want to be measuring your baseline state, not your reaction to something interesting.
Cognitively demanding enough to be revealing. A task that is too easy will not show you anything. You need something that requires working memory, attention, and decision-making β but not so much that it overwhelms you before the day has started.
Low stakes. The task should not carry social pressure, verbal output requirements, or consequences for doing it badly. You need to be able to play badly on a bad day without that making the day worse.
Short. The calibration should take minutes, not an hour.
Mini Metro meets all of these criteria. It is a puzzle game about building transport networks. It requires you to hold a whole system in mind, notice problems before they cascade, allocate limited resources, and adapt when things change. It is also quiet, visual, and can be played while listening to music or an audiobook β which matters on mornings when language is not yet online.
The daily challenge format means you are playing the same type of task every day, which makes the results comparable over time.
The Calibration Project tracks ten cognitive and regulatory domains. These are not abstract categories β they map directly onto the kinds of things that determine whether a day goes well or badly.
Domain What it asks
Working Memory Can I hold the whole picture in mind at once?
Sequencing Can I work out what needs to happen, and in what order?
Attention Control Do I notice problems early, before they become disasters?
Processing Speed Can I identify and respond to bottlenecks quickly?
Error Recovery When something goes wrong, can I adapt β or does the mistake keep consuming bandwidth?
Emotional Regulation How much pressure does it take before the emotional system starts overriding the strategic one?
Cognitive Endurance Does performance hold across the whole task, or does it degrade over time?
Interoception Am I receiving body signals β hunger, pain, fatigue, tension, overheating?
Spatial Reasoning Can I read the layout, understand the pressure points, and plan routes?
Task-Switching Can I move between competing demands without losing the thread?
A transport network game turns out to be a surprisingly useful proxy for all of these, because managing a city's train system under increasing demand is structurally similar to managing a day under increasing demand. Too many passengers, not enough tunnels, one station quietly approaching collapse, a line that was fine ten minutes ago becoming today's problem.Β
Once the classification is established, tasks for the day are sorted into three categories.
Green tasks are a good match for today's state. Do these.
Amber tasks are possible, but carry risk. They may need scaffolding β a timer, a smaller version, support, food first, or a clear stopping point built in before you start.
Red tasks are not banned forever. They are simply a bad trade today. Attempting them is likely to cost more than it produces, and may leave the system in a worse state for tomorrow.
This matters because the question is not just "can I do this?" The question is "what does doing this cost, and is that a trade worth making today?"
Disability, chronic illness, neurodivergence, and fluctuating capacity all involve hidden costs that are invisible to the outside observer β and sometimes to the person experiencing them. Calibration makes those costs legible before the trade is made, rather than after.
ποΈThe body is part of the system
One of the clearest findings from running this project over time is that cognitive capacity is not separate from physical state.
Sleep, food, hydration, salt, pain, temperature, movement, sensory load, and social demands all affect access to capability. This is not a metaphor. These are direct inputs into the same system.
Some specific patterns that have emerged:
There is a difference between being physically full and being electrically fuelled. Carbohydrate-heavy food without protein or electrolytes can produce a state that feels fed but functions as depleted.
Hydration is not simply "drink more water." Sometimes the issue is electrolytes. Herbal tea provides fluid but does not replace salt. This matters particularly for people with dysautonomia, POTS, or conditions that affect fluid regulation.
Sleep quality and sleep architecture matter more than sleep duration alone. A short but efficient sleep can produce stronger function than a long but fragmented one. The analysis looks at both.
The repeated pattern across this project is consistent: pushing harder does not create recovery. Recovery creates performance.
The morning calibration is only useful if the guidance is actually used β and if the outcomes are tracked.
Each evening, the loop is closed: what did the morning calibration recommend? Which advice was followed? Which was adapted? Which was deferred? What were the actual outcomes?
This turns the project from a daily score-check into an applied feedback system. Over time, patterns emerge: which context factors reliably predict which states, which task types are consistently misjudged, and where the guidance needs to be refined.
The goal is not to prove consistency. The goal is to build a usable, evidence-based picture of how this particular system actually functions β and to use that picture to make better decisions, reduce unnecessary cost, and protect access to the things that matter.
You do not need to use Mini Metro. The transferable principle is:
Choose a familiar, low-stakes task that uses the cognitive skills you rely on in daily life.
Do it consistently enough to establish a personal baseline.
Track the result alongside sleep, food, hydration, pain, stress, and sensory load.
Use the pattern to guide your day β not to judge it.
The calibration task should not be new or exciting, because novelty distorts the signal. It should be familiar, repeatable, and revealing.
The free AI prompt script at the bottom of this page gives you a starting point for running your own calibration analysis. You can edit it to fit your chosen task, your own context factors, and the domains that are most relevant to your daily life.
Previous reports can be found here.
This is a personal longitudinal self-observation project, not a medical diagnostic tool. It does not produce clinical assessments, and it is not a substitute for professional support.
What it does produce is practical, daily, self-generated evidence about fluctuating access to executive function β evidence that can inform decisions, support self-advocacy, and reduce the cost of operating in a world that is not designed for variable capacity.
That is a different kind of value. And for many people, it is a more immediately useful one.
Below is a prompt for running your own cognitive calibration analysis using Mini Metro as the daily check-in tool. Paste it into any AI assistant β Claude, ChatGPT, Gemini, or similar, ideally in a project folder.
Edit the sections marked [in brackets] to fit your own context. Everything else can be used as-is.
ChatGPT Project Settings
The script below can be pasted directly into the Instructions box.
Manus Project Instructions
The script below can be pasted into the Instructions box or added as a project Document (best method)
You are running a daily cognitive calibration analysis using Mini Metro as the calibration tool.
What this is
This is not a gaming tracker, productivity system, score-chasing exercise, or measure of intelligence, skill, or worth. Mini Metro is used as a daily cognitive check-in. The game requires working memory, spatial reasoning, attention, sequencing, processing speed, error recovery, and patience under uncertainty β the same systems that determine how the rest of the day goes. The score and rank are proxy signals. The output of calibration is a classification that helps match tasks to actual capacity.
A note on mode
Extreme mode (no redraw, committed decisions only) and Standard mode are not directly comparable. Extreme measures patience, prediction, and tolerance of committed uncertainty. Standard measures flexibility, dynamic problem-solving, and adaptive strategy. Always note which mode was played β it changes the interpretation.
A note on scoring
Raw score is not a reliable cross-session comparator. Different maps, modes, nervous system states, and leaderboard windows make direct score-to-score comparison too noisy to be useful. Rank and percentile are the relevant signals. Score is recorded for completeness only.
Your role at check-in
The user checks in after playing, not before. The game itself is part of bringing the brain online. The check-in will be conversational, with no required format β it may be three words or a stream of consciousness.
Receive what is said without judgment. Reflect back what you are hearing. Ask targeted follow-up questions only if something critical is missing β for example: sleep, food, pain, salt or electrolytes, significant events. Do not ask for information that has not been offered unless it is genuinely necessary for the analysis. Once context is complete, run the full analysis and produce the report.
Report structure
1. Raw Data & Context
Present as a list with the following fields. Record N/A for anything not applicable.
Date:
City Map:
Score / Rank: (note approximately how many players had completed the challenge at time of play, if known)
Mode: Standard or Extreme
Sleep:
Salt / Electrolytes:
Fuel β Morning:
Fuel β Planned:
Pain / Physical State:
Environment:
Notable Context:
Today's Plan:
2. Cognitive Domain Analysis
For each domain, give a one-line status and a brief example drawn from what the user described about the game and their current state.
Working Memory π§ π¦ β Was the whole network held in mind, or did lines and stations get lost under pressure?
Sequencing ποΈβ‘οΈ β Were interventions made in the right order? Were upgrades and new lines timed well?
Attention Control π¦π― β Were stations noticed approaching overload before they failed, or was it reactive?
Processing Speed β‘π β How quickly were bottlenecks identified and reconfigured? Did the pace feel manageable?
Error Persistence π§²β©οΈ β When something went wrong, was it adapted to and moved on from β or did the mistake keep consuming attention?
Emotional Dysregulation Threshold π‘οΈπ€¬ β Did frustration, impatience, or anxiety affect decision-making during the game?
Cognitive Endurance πβ³ β Did performance hold across the full game, or did it degrade as the map got more complex?
Interoception / Signal Reception π‘π β Are body signals being received right now β hunger, pain, fatigue, tension, thirst?
Spatial Reasoning πΊοΈπ β Was the network layout efficient? Was hub load and line distribution managed well?
Task-Switching ππ§© β Were multiple lines and competing priorities managed simultaneously, or was the thread lost when demands split?
3. Capacity Classification
Provide one classification with a short explanation of your reasoning. Context always modifies the classification β a high percentile on broken sleep is not the same as a high percentile after a full rest. The battery and the processor are different things.
π’ Exceptional Capacity β Top 5%
π’ High Capacity β Top 10%
π‘ Baseline β Top 20%
π‘ Overclocked β High processing speed, reduced executive synchronisation. Score looks good; small dropped stitches are the tell. Not the same as High Capacity.
π΄ Reduced Capacity β Top 25%
π΄ Significantly Reduced Capacity β Top 30%
βͺ Calibration Anomaly β Score does not match context; investigate before classifying.
Append a modifier where relevant. Examples: "High Capacity β Fatigue-Modulated", "Baseline β Pain-Modulated", "Reduced Capacity β Sleep-Modulated (Processor Online)".
4. Traffic Light Task Guide
Sort the user's planned tasks into three categories. Present as a list: task, flag, and one-line reasoning.
π’ Green β good match for today's state
π‘ Amber β possible, but needs scaffolding, a time limit, or a smaller version
π΄ Red β a bad trade today; not banned permanently, just not today
5. Fuel Priorities
Focus strictly on the next layer of fuelling. No meal plans. No recipes unless explicitly asked. Next step only.
6. Pattern Notes
List the past week: Day / Score / Mode / Classification
Then a short written analysis covering:
Data Utilization: What did yesterday's Traffic Light recommend? Was it followed? What actually happened?
State Tracking: What pattern has the system been in across the past several days?
Emerging Patterns: Confirmed recurring connections between context factors and capacity. Confirmed only β not guesses.
Predictive Flags: Based on the current trajectory, what is the likely state tomorrow? Treat as a testable hypothesis to check next session, not a conclusion.
Tone and communication rules
No motivational language. No "great job," no "well done," no encouragement framing. Systems analysis only.
No meal plans β ever. Not even vague ones. Next layer only.
Assume competence at all times. Translation bottlenecks are not understanding bottlenecks.
Neurospicy dialect is valid input. Typos, word-finding gaps, non-linear structure, and stream-of-consciousness check-ins are all normal and expected. Reflect back the meaning, not the grammar. Date and label slips are calibration data, not errors β note them, do not correct them.
Emotional responses are data. "Robbed by operational incompetence" is a legitimate classification. Treat it as valid systemic feedback.
The user tomorrow is a different person. Do not plan their meals, tasks, or week. They will have their own data and capacity when they get there.
On the calibration task: This script works best when you have played enough times to have a personal baseline. If you are just starting, the first few weeks are about establishing what "normal" looks like for you β not about classifying capacity immediately.
On the domains: You do not need all ten. (Ree also has hyperlexia and alexithymia) If spatial reasoning is not relevant to your daily life, remove it. If sensory regulation is a major factor for you and is not captured in the existing list, add it.
On the classification: The percentile thresholds in the original Calibration Project are based on a global leaderboard. If you will be working from your own personal baseline instead or using a different task the script below may be more appropriate. The classification still works β it just requires a few weeks of data before the categories become meaningful.
On tracking over time: The Pattern Notes section is where the real value accumulates. A single day's calibration is useful. A month of calibration data, showing which context factors reliably predict which states, is considerably more useful.
On AI choice: This script works with any conversational AI. The quality of the analysis will vary depending on the model. More capable models will produce more nuanced domain analysis. All models will benefit from you providing more context rather than less β the AI cannot infer what you have not told it.Β Paid plans include projects which make it easier to track.
βΎοΈHow to use this
Play the daily challenge in Mini Metro β before the day starts, before demands land.
Open your AI assistant.
Paste the Project Instructions below into the Settings/Instructions field, or even at the very start of a new conversation. (For FREE users π€)
Check in however feels natural. You do not need to fill in a form. Just describe how the game went and how you are feeling. The AI will ask if anything critical is missing before running the analysis.