How a metric gets chosen, how its peer set is picked, where the figures come from, and what a verdict does and does not claim.
Every claim on this site is a comparison, and a comparison is only as good as the choices behind it — which countries, which measure, which year, which source. This page sets out those choices, so that a reader can check them and disagree with them specifically rather than in general.
For why the project exists and how the five pillars fit together, see about. What follows is the working method.
How a metric is selected
Not everything worth caring about can be benchmarked, and not everything that can be benchmarked is worth a page. A candidate has to clear five tests before it is built.
It measures performance, not intention. The question is what the country achieves, not what a government announces, budgets, or intends. Spending appears as a metric only where the spending itself is the commitment being measured — defence expenditure against the NATO target, for instance — and not as a stand-in for the outcome it is meant to buy.
A like-for-like comparison exists. Either peer countries publish the same measure on a comparable basis, or Canada publishes a standard of its own that it can be held to. Where neither holds, the metric is not built, however important the subject.
A primary source publishes it, and keeps publishing it. A metric has to be refreshable, or the page freezes at its publication date and quietly becomes wrong. A single study, however good, does not support a maintained benchmark.
It measures the thing itself. A proxy is rejected where it answers a different question from the one asked. Enforcement seizures, for example, measure enforcement activity; they cannot size the production problem they are drawn from. Where both questions matter, they become two metrics rather than one metric doing the work of two badly.
It survives disaggregation. Composite indices bundle someone else's weighting decisions into a single number, so the underlying series is preferred wherever it can be obtained. Where an index is genuinely the only available measure, the page identifies which subcomponents drive Canada's score rather than reporting the headline alone.
Related measures are paired, not collapsed. Where two measures of the same subject move differently — a level and its rate of change, the capital raised and where it goes, the assessment stage and the licensing stage — they are built as separate pages that link to each other. Averaging them into one figure would conceal precisely the disagreement that makes the pair informative.
Choosing the peer set
The peer set is chosen on the merits of the comparison and before the result is known, and it is stated on the page. The default is the G7: countries of comparable income and institutional type, against which Canada already measures itself in ordinary public debate.
Two departures from that default are used, both declared where they apply.
A wider or different comparator group, where the G7 series does not exist for all seven countries, or where the G7 is simply the wrong comparison — a measure that turns on resource endowment, federal structure, or population density may be better read against a group selected on that characteristic. The page names the group and the reason for it.
No international comparison at all, where no other country publishes on a basis that can honestly be set beside Canada's. In that case the benchmark is Canada's own published service standard: the target the state set for itself, and how often it meets it.
Like for like. A cross-country comparison must hold constant everything except the country: the same profession, the same stage of the same process, the same type of institution, the same units. A six-week turnaround at one country's nursing regulator set beside a fifteen-week one at another country's engineering body is not a country comparison — it is a profession comparison wearing a flag. Where the like-for-like series does not exist, the page compares within a single profession or process and states what it therefore cannot conclude.
How to read a verdict
Every metric carries one of three verdicts, describing where Canada sits within that metric's peer set, on that measure, at the vintage stated on the page.
Strong
Canada sits in the upper part of its peer set — a position other comparable countries would want.
Watch
Canada sits in the middle — no crisis, but no advantage either, and often a direction worth watching.
Weak
Canada sits at or near the bottom of its peer set.
The verdict follows Canada's position in the peer set, but it is a judgement rather than the output of a formula, and the page shows the working that produced it. Where a measure's level and its direction disagree — a strong position eroding quickly, a weak one improving steadily — the page says which of the two the verdict reflects, and why. A verdict is a one-word summary of a page, not a substitute for reading it.
Three things a verdict is not. It is not a grade for any government: most of these measures move over decades and across administrations. It is not a claim about cause; establishing that Canada trails its peers is a different exercise from establishing why. And it is not comparable between metrics — each metric has its own peer set, so Weak on one page and Weak on another are not necessarily the same distance from the front.
Pillar-level verdicts on the metrics overview are computed from the metrics currently in view, so filtering the dashboard recomputes them rather than leaving a stale headline in place.
An important exception. For some measures no comparable international series exists — no other country publishes passport processing times or veterans' disability claim backlogs on a basis that can be set beside Canada's. Where that is the case, the benchmark is Canada's own published service standard, and those pages say so explicitly. We would rather compare Canada to its own commitments than manufacture a false international ranking out of mismatched definitions.
Where the figures come from
Sources are used in order of preference, and the page says which level it is drawing on.
The primary series, from the agency that produces it. Statistics Canada, the Bank of Canada, the IMF, the OECD, the World Bank, departmental performance reports and departmental plans, and the equivalent national agencies in peer countries.
The underlying records, assembled here. Where no published series exists but the case-level or administrative records do, the series is constructed from those records and the construction is described on the page, so that someone else could rebuild it.
A named institutional source, cited and flagged. Where neither of the above is obtainable — a research institute, an industry association, a private data holder — the holder is named, the page says the figure depends on them, and the limitation travels with the number.
Not used: figures taken from news reports, values read off someone else's chart, aggregator sites and encyclopaedias, and any number that cannot be traced back to a body prepared to stand behind it. Every chart on every metric page carries its source, the specific table or series identifier, and the vintage of the data.
How figures are checked
Each metric page has a companion workbook holding the figures behind its charts. Before a page is published, and again before any change to it ships, it passes four checks.
Page to workbook
Every number and every chart point on the page traces to a cell in the workbook behind it. Nothing appears on a page that does not exist in its data, and no value is ever invented to close a gap or illustrate a point.
Workbook to source
Every figure in the workbook matches the cited source at the stated vintage. Where a source has revised its series since the figure was taken, the revision is adopted rather than the original quietly retained.
Comparability
Where countries define a measure differently, the difference is disambiguated on the page rather than averaged away. Where a ratio can be computed on more than one basis, the page states which basis it uses.
Internal consistency
No claim contradicts another: a series described as a record high has to exceed the maximum of the series shown. Where the same figure appears on more than one metric page, it agrees across them, or the difference is explained on both.
Charts
Missing data is left missing. A gap in a series is drawn as a gap. No interpolation, no value carried forward from the previous period, no placeholder standing in for a number that was not published.
Every chart states what it plots and in what units, and carries a source line with the citation and the vintage.
Projections are labelled as projections. Where a value is a forecast or an agency outlook rather than an observed figure, the chart and the surrounding text both say so.
Country labelling is consistent across the site, so the same country reads the same way on every page it appears on.
Vintage, updates, and revisions
Figures carry the vintage of the data, not the date of the page. A metric published this month may show a figure from two years ago because that is the most recent complete period its source has released; where that is the case the page says so, rather than substituting something more current and less comparable.
Metrics are updated when their sources publish, not on a fixed calendar. A page is not refreshed with a partial period presented as a complete one. Where a source revises history, the revised series is adopted; where a revision changes a finding, the change is noted on the page rather than absorbed silently.
Corrections
If you believe a figure here is wrong, please say so. Corrections are published rather than quietly absorbed.
A challenge to a specific figure will be answered specifically: the source it came from, the judgement calls made in constructing it, and any revisions since publication. That is the standard this project holds itself to. A number that cannot be defended in that detail should not be on the site in the first place.
How metrics are tagged
Besides its pillar, each metric carries two sets of tags, used by the filters on the metrics overview.
Cross-cutting lenses group metrics by the question a reader arrives with rather than by the structure of the site. Seven are in use: cost of living; economic competitiveness; energy and resources; infrastructure delivery; innovation and technology; national security and sovereignty; and social fabric. Most metrics carry one or two, and three is the practical upper bound — a lens is applied only where the metric is genuinely informative for that reader, not wherever it is loosely relevant.
House of Commons standing committees map each metric to the committee whose mandate would actually study it — usually one, occasionally two. Sixteen committees currently appear in the filter, and the list is self-curating: a committee is added only once a metric is mapped to it, so the dropdown never offers an empty view.
What this method cannot tell you
Position, not cause. A benchmark establishes where Canada sits. It does not establish why, and the causal claim is usually the contested one. The pages set out what the evidence supports and stop there.
Peer sets are not shared between metrics. Each metric has its own comparator group, chosen for that measure. The metrics should not be read as though they share a denominator, added into a composite, or ranked against one another.
Comparability is imperfect even at its best. Official series defined similarly are still collected differently, and national statistical practice varies. Where that materially affects a comparison the page says so; the ordinary caution that applies to any cross-country comparison applies here too.
Coverage is a judgement. The 34 metrics currently published are not the whole of national performance. What is here reflects a judgement about what matters and what can be sourced defensibly. The absence of a metric is not a finding that the area is fine — more often it means no comparable series exists yet.
A benchmark is a starting point. Its value is in telling you which questions are worth asking, and in denying anyone the comfort of an unbenchmarked claim. It is not the answer to the question it raises.
Think something here is wrong? Questions about a figure, corrections, and suggestions for metrics worth adding are all welcome — see contact.