Three layers
The scenarios sit between two other bodies of work, and each does a different job.
| Layer | What it is | What it does |
|---|---|---|
| Trends | The GINC 250: slow structural forces, each scored against every scenario | The evidence layer. Trends explain why a scenario is more or less likely and who is exposed, and they give the picture: the network and the matrix. They are not an input to any calculation. |
| Scenarios | The Library: shocks specified on a common grammar, each with a loading table | The load. A scenario says which capabilities are stressed, how hard, and through which channel. |
| Capability scores | The Atlas: every nation rated on 254 capabilities in nine domains, each domain placed in a tier | The foundation for calculation. A test takes a nation’s scores, applies a scenario’s load and reports what changes. |
Three ways to run a scenario against a nation
A scenario can be run against a nation’s capability scores in three ways. Each answers a different question. The load test and the capture test have been run for illustration against the 2026 capability scores, with exposure set nation by nation, and the results are published with their settings and summed into a vulnerability index and an opportunity index. The break-point test has not been run. The methods and their settings are proposals for the review committee, not adopted practice.
| Load test | Break-point test | Capture test | |
|---|---|---|---|
| The question | What happens to this nation under this scenario? | How much of this scenario can the nation take? | How much of the upside does this nation capture? |
| Direction | Forward: fix the shock, read the damage | Reverse: fix the failure, find the shock | Forward: fix the opportunity, read the share |
| Fixed | The scenario at Standard Severe | The failure condition | The scenario at Standard Severe |
| Solved for | Stressed scores and tier shifts | The intensity at which the nation fails | Capture rate and net position |
| Used for | Downside scenarios; comparison across nations | Downside scenarios; a nation’s own planning | Upside scenarios; the gainers in downside ones |
| Headline result | ‘Two domains fall a tier; energy fails first’ | ‘Fails at 60 days of closure; Standard Severe is 90’ | ‘Captures 40 per cent of the gain; grid capacity is the limit’ |
All three use the same inputs:
- Baseline. The nation’s capability scores and domain tiers from the Atlas.
- Load. The scenario’s loading table, mapped from working domains to the capabilities beneath them through the domain map. High, Medium and Low become a size of shock.
- Exposure. How directly the scenario reaches this nation, from 0 to 1. It starts from the scenario’s regional exposure and is refined by Atlas metrics that the scenario’s own parameters name: the share of crude arriving through one strait, months of import cover, the share of the budget funded by aid.
- Absorption. How much of the shock the nation’s buffers soak up, from 0 to 1. Buffers are the capabilities the scenario’s stakeholder actions point to: stockpiles, fiscal space, backup power, surge capacity.
- Policy response. The grammar’s ninth slider: none, current plans executed, or best practice.
1. Load test
What happens to this nation under this scenario?
The load test is the forward stress test, and the closest to what bank supervisors do. The scenario is fixed at Standard Severe and applied to every nation alike, so results compare.
For each capability the scenario loads:
stressed score = baseline score − load × exposure × (1 − absorption)
- Map each loaded working domain to its capabilities.
- Turn the load into a shock: High is sized to move a fully exposed, unprotected capability down one tier; Medium half of that; Low a quarter.
- Scale by the nation’s exposure and reduce by its absorption.
- Recompute each domain’s score and tier.
- Run three times, at each policy response. The gap between ‘none’ and ‘current plans’ is the capability finding. The gap between ‘current plans’ and ‘best practice’ is the policy recommendation.
It reports: stressed scores by domain; the number of domains that fall a tier; the capability that falls furthest; the recovery horizon from the scenario’s scorecard.
Strengths. Simple, transparent and the same for every nation, so it can be published as a comparable result. It uses only what the Library and the Atlas already hold.
Limits. It is linear: it does not capture one failed capability pulling down another. The load sizes are judgements. It says what happens at one intensity and nothing about how close the nation is to something worse.
2. Break-point test
How much of this scenario can the nation take?
The break-point test is a reverse stress test. It fixes the outcome and solves for the shock. It does not ask how likely the scenario is, which makes it useful for scenarios whose likelihood is contested.
- Set the failure conditions. A nation ‘breaks’ when a domain falls a full tier, or when an essential capability falls below a floor: days of fuel cover, hospital surge capacity, months of import cover, the share of the population with power.
- Start below Standard Major and raise the scenario’s intensity along one slider at a time: severity, duration, external support, then concurrency with a second scenario taken from the first one’s triggers.
- At each step, apply the load test.
- Stop at the first step where a failure condition is met. That setting is the break point.
It reports: the break point on each slider (‘fails at 60 days of closure’); headroom, which is the break point as a multiple of Standard Severe, so that below 1 means the nation fails before the standard case and above 1 means it has margin; the capability that breaks first; and the slider the nation is most sensitive to.
Strengths. It gives a government a number it can act on and a named weakest point. It finds compound failures that a single scenario misses. It is the form in which stress tolerances can be published, while named reverse stress tests stay confidential, as the governance page sets out.
Limits. Someone has to choose the failure floors, and the choice is political. Results depend on the order in which sliders are raised. It needs the load test’s assumptions and adds its own.
3. Capture test
How much of the upside does this nation capture?
The capture test is the positive test. It is built for the upside scenarios, where the question is not what breaks but who gains, and it also applies to the nations that come out ahead in a downside scenario. Here the loading table is read the other way: the capabilities a scenario loads are the ones a nation needs in order to take the gain.
net position = gain × readiness − loss exposure × (1 − adaptability)
- Readiness. Take the capabilities the scenario loads High. Readiness is the nation’s score on them relative to the frontier, and it is held down by the weakest: a nation with cheap power and no grid to carry it captures little of an energy breakthrough.
- Loss exposure. Measure the nation’s dependence on what the scenario makes obsolete: hydrocarbon revenue as a share of the budget, knowledge-work exports, industries built behind tariffs.
- Adaptability. Score the capabilities that let a nation move out of a losing position: fiscal buffers, retraining, government effectiveness.
- Combine the three into a net position.
It reports: a capture rate from 0 to 100; a net position of gainer, mixed or loser; the binding capability, which is the one whose improvement would raise the capture rate most; and the nation’s position among its peers.
Strengths. It turns a scenario into an investment priority and shows that the same event sorts nations into winners and losers. It gives the upside scenarios a result of their own, since ranking them by risk would misread them.
Limits. The size of the gain in an upside scenario is far less certain than the size of a loss in a downside one. Readiness measured today may not hold over a scenario that runs for five years. It is the least tested of the three.
How the three fit together
The three are complementary, not alternatives. A sensible order of work is: the load test for every nation against every downside scenario, as the comparable public result; the break-point test for each nation against the few scenarios where the load test shows it most exposed; and the capture test for every nation against the upside scenarios.
The run so far
The results pages show the load test on the fourteen downside scenarios and the capture test on the ten upside ones, for 197 nations. All settings are published in the test settings file on the data page.
- Scores. Capability scores for 73 capability groups in nine domains, on a scale of 0 to 20 in seven bands of three points, from the GINC National Capability Index 2026. The index describes them as candidate ratings, not adopted ratings.
- Load. High is three points, which is one band; Medium is 1.5; Low is 0.75. Each scenario singles out about ten capability groups that take its domain’s load in full. Other groups in a loaded domain take the load one level lower.
- Exposure, nation by nation. Each scenario has its own exposure formula: a weighted blend of two or three indicators and the scenario’s regional exposure. The indicators come from the 2025 World Factbook and are scaled from 0 to 1: the share of oil and gas use not met at home, months of imports covered by reserves, public debt, trade relative to GDP, natural hazards, land borders with states in conflict, the share of the population over 65, and others. A chokepoint closure, for example, weighs energy import dependence at 45 per cent, the regional exposure at 35 and trade openness at 20. Where a nation does not report an indicator, its weight passes to the rest. Each scenario’s results page lists its formula.
- Absorption. A nation’s mean score on the scenario’s buffer capabilities, as a share of the top of the scale, times 0.75.
- Policy response. No response sets absorption to zero. Current plans uses the nation’s own buffers. Best practice lifts each buffer to 15, the foot of the Advanced band, where it is lower.
- Capture test. Gain and loss exposure are set from indicators where one fits: energy import dependence for the gain from cheap energy, fuel exports for the loss; ageing for pension strain. Where none fits they are set by region, with named exceptions.
What the results report. For the load test: the change in the nation’s capability index in points and per cent; the rank shift, which is the change in its place among all nations once every index is stressed; the number of capability groups and domains that fall a band; and the weakest point, the singled-out capability left lowest. For the capture test: capture, loss and net position; gainer, mixed or loser; rank; and the binding capability.
Two indices
The scenario results are summed into two indices for each nation, weighting each scenario by its likelihood.
- Likelihood. Each likelihood band is turned into a probability at the midpoint of its range: 0.1, 0.6, 3 and 15 per cent for bands 1 to 4, and 40 per cent for band 5, which is open-ended. The indices are computed at two horizons: ten years, from the ten-year band, and three years, scaled up from the two-year band on a constant yearly chance.
- Idiosyncratic scenarios. For scenarios that strike one nation at a time, a nation’s probability is the scenario’s probability scaled by its exposure relative to the average nation. A state with thin reserves and heavy debt is then more likely to meet a funding crisis, as well as harder hit by one.
- Vulnerability index. Expected loss: the sum over the downside scenarios of probability times the share of the baseline capability index lost. It is scaled so that the most vulnerable nation is 100. The loss is taken as a share of baseline, not in points, because a weak state that loses one point has lost far more of what it has than a strong one.
- Worst case, and tail. The largest single fall across the scenarios, with the scenario that causes it, and the mean of the three largest. These sit beside the index because an average hides the one scenario that breaks a nation.
- Opportunity index. The likelihood-weighted mean net position across the upside scenarios, from −100 to +100, with the best single case beside it.
- Other measures. Breadth is the number of downside scenarios in which a capability domain falls a band. Shock absorbed is the share of the load a nation’s own buffers take, across all downside scenarios. The policy dividend is the expected index points recovered by moving from current plans to best practice. The weakest link is the capability that is the nation’s weakest point in the most scenarios.
- Four groups. Each nation is placed against the median of each index: well placed, high stakes, insulated or exposed.
- Seven archetypes. Nations are also grouped by the whole shape of their results. Each nation’s result in all 24 scenarios is standardised and the nations are clustered into seven groups by k-means, with a fixed seed so the grouping does not change between runs. The groups are found by the method; the names and descriptions are GINC’s reading of them. An archetype describes a pattern of results in this run, not a nation.
How to read them. The run shows the method working on real scores. It is not a finding about any nation, and the indices are not adopted rankings. Three cautions apply. Scenarios are treated as independent when one can trigger another, so the sum misses compounding. The two indices move against each other, because the capabilities that absorb a shock are largely those that capture a gain. And very small states produce odd results, because several capabilities do not apply to them.
Two objects: the Library and the Edition
The Library is the stable set of named scenarios. Each scenario has a permanent identifier and slug, a record on one schema and a rating history. The Library grows over time; scenarios are retired, not deleted.
The Edition is the annual cut, released each January. It ranks ten scenarios from the Library on likelihood and impact, records movement against the previous edition and publishes the paper and the data. Rankings move; scenarios persist.
Version strings follow the two objects. The Library carries a semantic version (this is v0.4). An Edition carries its year (2027). Every scenario page shows both.
Three rules shape every page:
- Likelihood is global; impact is per nation. A scenario’s likelihood is a property of the world. Its impact is a property of each nation’s balance sheet. The Edition rates systemic impact and a typical national impact; nation-specific impact arrives when the Atlas connects.
- The data is the product; the site is a view of it. Every scenario is a data object rendered into a page.
- Scenarios are shocks, not trends. Slow structural forces such as polarisation, inequality, climate change, demographic ageing, antimicrobial resistance and economic stagnation belong to the Trends product. Each scenario names the trends that raise its likelihood.
Upside scenarios
Not every shock is a loss. The Library carries ten upside scenarios beside the fourteen downside ones, two in each family. Five entered at v0.3: superintelligence takeoff, an energy abundance breakthrough, a biomedical breakthrough wave, great-power détente and a coordinated sovereign debt reset. Five entered at v0.4: an emerging-market growth breakout, a Middle East regional settlement, an agricultural yield revolution, a digital public infrastructure leapfrog and a hundred-day pandemic defence.
An upside scenario is a shock that most would count as progress. It is not upside for everyone, so each record names who gains and who loses. It is specified on the same grammar and rated on the same scales, read as follows:
- Likelihood is unchanged: the probability that the Standard Severe event occurs within the horizon.
- Impact measures the scale of change, in either direction: output gained or redistributed, and capability band shifts up or down. Levels 4 and 5 are labelled Major and Transformative for these scenarios, not Severe and Catastrophic.
- Load marks the capabilities a nation needs in order to capture the gain or absorb the loss.
Upside scenarios are not ranked. The ranked ten order downside risk, and a priority score would read a large gain as a large threat. Some of these scenarios are classed elsewhere as risks, superintelligence above all; the record says so and lists the open question.
The grammar
Every scenario is specified on the same ten sliders, so that scenarios can be compared.
| # | Slider | Values |
|---|---|---|
| 1 | Severity | major / severe / extreme |
| 2 | Duration (acute phase) | 30 days / 90 days / 1 year / 3 years / 5 years |
| 3 | Onset | sudden / rapid (weeks) / gradual (years) |
| 4 | Warning time | none / days / months |
| 5 | Scope | national / regional / global |
| 6 | Origin | natural / accidental / adversarial (great power / neighbour / non-state) |
| 7 | External support | full / partial / none |
| 8 | Concurrency | standalone / plus one named scenario / plus two |
| 9 | Policy response assumed | none (pure exposure) / current plans executed / best practice |
| 10 | Recovery horizon | months / years / structural |
Recovery horizon is the time to return to 90 per cent of pre-shock capability. It was split from duration after the 2026 closure of the Strait of Hormuz showed about five weeks of closure followed by months of partial recovery.
There are three presets: Standard Major, Standard Severe and Standard Extreme. Standard Severe is the comparability anchor. Every scenario is rated at it, following the Swiss practice of assessing all hazards at one common intensity. Standard Extreme is the reasonable worst case in the UK National Risk Register’s wording: ‘the worst plausible manifestation of that particular risk (once highly unlikely variations have been discounted)’.
In this release the sliders are displayed at their preset values and are not adjustable. Nothing on the site is computed from them.
Likelihood
Likelihood is published in five bands, following the UK National Risk Register convention, at two horizons: two years (to end-2028 for Edition 2027) and ten years.
| Band | Label | Probability over the horizon |
|---|---|---|
| 1 | Remote | under 0.2 per cent |
| 2 | Unlikely | 0.2 to 1 per cent |
| 3 | Possible | 1 to 5 per cent |
| 4 | Likely | 5 to 25 per cent |
| 5 | Highly likely | over 25 per cent |
Each scenario carries a likelihood type.
- Systemic. The event is global or near-global when it occurs. Likelihood is the probability that an event meeting the Standard Severe definition occurs within the horizon.
- Idiosyncratic. The event strikes one nation at a time. Likelihood is the probability that a given nation experiences the Standard Severe event within the horizon. It is estimated from the base rate across the 197 GINC nations, with clustering noted where exposure is concentrated: hybrid campaigns in Europe, displacement in the neighbours of conflicts.
Impact
Systemic impact is global and uses five-year accounting. It follows the GDP@Risk convention developed by Lloyd’s and the Cambridge Centre for Risk Studies, extended from 107 to 197 nations.
| Level | Label | Global output at risk over five years at Standard Severe | Also |
|---|---|---|---|
| 1 | Limited | under 0.5 per cent | no capability band shifts expected |
| 2 | Moderate | 0.5 to 1 per cent | band shifts in one domain, few nations |
| 3 | Significant | 1 to 2.5 per cent | band shifts across several domains or many nations |
| 4 | Severe | 2.5 to 5 per cent | widespread shifts; some structural |
| 5 | Catastrophic | over 5 per cent, or irreversible structural change | widespread structural shifts |
National impact is rated for a typical highly exposed nation on the same five labels. It is judged on seven dimensions adapted from the National Risk Register’s impact framework and mapped to the National Capability Framework: human welfare, essential services, economic damage, security, international position, social cohesion and environment. The score is the highest dimension reached, not an average.
Calibration points: the five-year, probability-weighted losses published by Lloyd’s (pandemic US$13.6 trillion, geopolitical conflict US$14.5 trillion, food and water US$5 trillion, cyber US$3.5 trillion, space weather US$2.4 trillion) sit between levels 2 and 3 of the systemic scale. Their extreme variants (pandemic up to 6.4 per cent of global GDP, conflict up to US$50 trillion) reach level 5. Supervisory ‘severely adverse’ scenarios (Federal Reserve 2026, Bank of England 2025) are level 5 events for the economies they model.
Capability loading
Each scenario maps to nine working domains, three in each dimension of the National Capability Framework: Hard, Soft and Economic. A domain’s load is High where capability band shifts are expected under current plans, Medium where band shifts occur only with no policy response, and Low where there is strain without a band shift.
The working domain labels are placeholders. Their reconciliation to the canonical framework domains is recorded in domain-map.json on the data page; two of the nine do not yet map one-to-one.
Confidence and priority
Confidence has five levels, from very low to very high, each with a one-line reason. Confidence refers to the rating, not the scenario.
The development ranking uses a priority score: two-year likelihood multiplied by the higher of systemic and national impact. Tier I is 15 and above, Tier II is 9 to 14 and Tier III is under 9. Ties are broken by systemic impact, then confidence, then ten-year likelihood.
The panel, not the arithmetic, sets the Edition. The priority score exists for development and transparency. Where the published ranking departs from it, the departure is recorded as a judgement in the edition’s data file and shown on the edition page.
Stakeholder relevance
Stakeholder views order the ranked scenarios by a relevance score from 1 to 5. In this release the score is a development default derived from the loading table: each group reads two working domains; High counts 2, Medium 1 and Low 0; the score is the sum plus one. From the mid-2027 release it is replaced by the split of the panel’s own ratings by stakeholder group.
Borrowed calibration
Where a supervisor or an established scenario programme has published a calibration, the Standard Severe preset borrows it and cites it. GINC’s own judgement is labelled as such. Each programme is described in full, with its strengths and limitations, on the benchmarks page.
| Source | Used for |
|---|---|
| Federal Reserve 2026 stress test scenarios | Global financial crisis: unemployment, house price and commercial property paths |
| Bank of England 2025 Bank Capital Stress Test | Trade volume shock, energy price multiples, the inflationary variant of the financial crisis |
| EBA 2025 EU-wide stress test | Geoeconomic confrontation: the Standard Severe macro path |
| Lloyd’s systemic risk scenarios: geopolitical conflict, human pandemic, cyber, space weather, food and water | The systemic impact scale and five-year accounting |
| Cambridge Global Risk Index | Family classes; severity convention |
| UK National Risk Register 2025 | Likelihood bands; reasonable worst case; national impact dimensions |
| IMF and World Bank Debt Sustainability Frameworks | Sovereign funding crisis: parameter conventions |
| Switzerland, Disasters and Emergencies in Switzerland 2025 | Assessing all hazards at one common intensity; hazard ranking precedents |
| Norway, Analyses of Crisis Scenarios 2019 | National-register precedents |
| WEF Global Risks Report 2026, Eurasia Group Top Risks 2026, CFR Preventive Priorities Survey 2026, Allianz Risk Barometer 2026 | Expert and business perception of likelihood |
Figures marked ‘about’ or ‘around’ on scenario pages are approximate. Anchors from 2026 rely on press reporting and are to be re-checked against primary data before the paper is typeset.
Elicitation method for the Edition
The scores on this site are provisional GINC desk scores. The 2027 Edition replaces them with the results of a structured expert elicitation:
- 40 to 60 named panellists, drawn across the nine regions and the four stakeholder groups.
- Two rounds: structured individual ratings with a written rationale, then a calibrated discussion round.
- Publication of the distribution of ratings, not only the median.
- Panel tagging by region and stakeholder group, which produces the stakeholder and regional splits.
- Adoption of the final ratings by the review committee.
- Conflicts of interest declared.
Limitations
- Loading tables are judgement-based until the Atlas connects.
- There is no agreed loss model for nations; national impact is a structured judgement.
- Likelihoods for adversarial scenarios are politically sensitive and rest on thin evidence.
- Bands and levels carry a risk of false precision. Each is published with its range and a confidence level for that reason.
Trends
The GINC 250 holds the slow structural forces behind the scenarios: 250 trends, each assigned to the one domain of the National Capability Framework it bears on most, and each carrying a theme as a second way to read the list.
The list was built in three passes. It began from published trend and risk lists used by investors, governments and international affairs institutes, and from the trends the scenarios themselves depend on. It was then reviewed as a whole: trends too narrow to bear on national capability were cut, near-duplicates were merged, broad trends were broken out into the parts the scenarios turn on, and trends were added where a capability domain was thin. The result is GINC’s own list; individual trends no longer carry a source. The lists drawn on are named in the change log.
- Capability domain. Every trend has one primary domain. The domains hold between 25 and 34 trends each.
- Trend by scenario. Every trend is scored against every scenario on a seven-point rubric: not related, very low, low, moderate, high, very high, critical. A trend relates to a scenario when it changes how likely the scenario is, how hard it lands, or is itself sharply changed by it. These are GINC desk judgements. Where trends were merged, the merged trend takes the strongest relation among its parts.
- Trend by trend. Similarity between two trends is derived, not scored by hand: 0.30 from shared tags, 0.30 from how alike their scenario profiles are, 0.25 from how alike their names and descriptions are, 0.10 if they share a capability domain and 0.05 if they share a theme. The result is cut into the same seven levels. Critical similarity is rare by design and marks parent-and-child or sibling trends.
- Sensitivity. A trend’s sensitivity is the sum of its scenario scores as a share of the maximum. Breadth is the number of scenarios it scores High or above. Trends with high sensitivity move the most scenarios.
- Network measures. Trends are linked where similarity is Moderate or above. PageRank restarts at each trend in proportion to its sensitivity, so rank flows from scenario-relevant trends along similarity links. Betweenness marks trends that bridge clusters. Eigenvector centrality gathers in the densest cluster and is reported with that caveat.
- Rank. The composite is 0.4 sensitivity, 0.3 PageRank, 0.2 betweenness and 0.1 eigenvector, each scaled so its highest trend is 100.
The measures are computed from the published scores by a script in the repository and can be reproduced from the downloads on the data page.
Relationship to the Atlas, the Shock Register and Trends
The Atlas is the balance sheet: every nation’s capabilities, metrics and sources. Scenarios is the load. This site links to the Atlas for every capability domain, and will link to it for every nation once nation-level results are published.
The Shock Register holds the historical events that calibrate the scenarios; each scenario’s anchors reference it. Trends holds the slow structural forces; each scenario names the trends that amplify it.