Skip to main content
Quartyl
Results & Analyticsprofessional

Risk and Reliability Scores: What the Numbers Mean

What the benchmark risk and reliability scores actually are: fixed additive point tables over the accepted set’s dispersion, size, loss share, outliers and range tightness.

Quartyl Team

Alongside the range, the statistical engine scores the benchmark itself — two 0–100 numbers computed on the accepted comparables set: a benchmark risk score (lower is better) and a benchmark reliability score (higher is better). They are an index of the pool’s shape, not an opinion about the tested party and not a probability. A study can sit comfortably inside a range and still carry a weak benchmark — these are the numbers that tell you how much weight the conclusion carries.

Both are fixed additive point tables over the accepted PLI distribution (backend/app/services/statistics/statistics.py:219-296, thresholds in backend/app/constants/statistics.py:236-316). There is no statistical significance testing behind them: no p-values, no confidence intervals, no hypothesis tests. The engine’s only inferential statistic is a Spearman rank coefficient computed alongside each Pearson correlation for the correlation chart, and its p-value is discarded.

The benchmark risk score

Risk starts at 0 and accumulates penalties; the total is capped at 100.

Factor Points added
CV above 50 / above 30 / above 15 +30 / +20 / +10 (one band only)
Accepted count below 4 / below 8 / below 15 +30 / +20 / +10 (one band only)
Loss-making share above 20% / above 10% / above 0 +25 / +15 / +5 (one band only)
Outliers beyond the 1.5×IQR fence +5 each, capped at +15

Loss-making is counted generously: a comparable with a PLI of exactly zero counts as a loss. Outliers are scored, not removed — no company leaves the accepted set because of the fence; the fence only adds risk points.

The deterministic insight lines that render on the results view read two bands, not three: above 60 emits an “elevated risk” danger line, above 30 a “moderate risk” warning line, and at or below 30 the engine stays quiet about risk (backend/app/services/statistics/insights.py:227-238).

Benchmark reliability

Reliability starts from a base of 50 and adds credits, with one penalty branch; the result is clamped to 0–100.

Factor Points
Accepted count ≥ 20 / ≥ 10 / ≥ 5 +25 / +15 / +5 (one band only)
CV below 15 / below 30 / below 50 +15 / +10 / +5 (one band only)
IQR under 0.5 × the median (range tightness) +10
Loss-making share above 10% −15
Any loss-making comparables (share above 0) −5

Two consequences of that arithmetic are worth knowing before you argue with a score:

  • The scale is not symmetric. With no bonus triggered and the heavier loss penalty applied, reliability lands at 35, so the bottom of the scale is unreachable in practice; the credits top out at exactly 100.
  • Two tightness thresholds are in play. The reliability credit uses IQR/median below 0.5; the separate “narrow interquartile range” success insight is stricter, at below 0.3 (backend/app/services/statistics/insights.py:252-262). A set can earn the credit and still not print the insight line.

The bands the insights use on reliability:

  • 70 and above — “high confidence in the derived range”.
  • 40–69 — “moderate confidence. Consider additional comparables”.
  • Below 40 — “low confidence. The benchmark may not withstand audit scrutiny”.

Risk and reliability are related but not mirror images — one accumulates penalties from 0, the other dispenses credits from 50 — so a set can read moderate on both, or low risk with only moderate reliability when the set is thin but happens to be tight.

What moves the numbers

The four quality signals behind both scores, and the thresholds the deterministic insights apply to them:

  • Set size significance. Fewer than five accepted comparables trigger a warning against the OECD’s 3–5 reliable comparables guidance; five to nine print as adequate; ten or more print as a robust statistical basis (insights.py:152-174). This is the single most common driver of a low reliability score.
  • Dispersion (CV). Above 50% is high variability — significant heterogeneity among comparables that reduces the reliability of the range. 30–50% is moderate dispersion. Anything above zero but below 30% is read as a homogeneous set.
  • Range tightness. IQR below half the median earns the reliability credit; below 0.3× the median it also earns the success insight.
  • Loss-making comparables. The scores count a PLI of zero as a loss (<= 0), while the written insight triggers only on a strictly negative minimum: loss-making comparables increase benchmark risk and can distort the arm’s-length range. See the loss-making comparables reference for how they are handled in a defensible study.

What is not an input

The acceptance rate (accepted over screened) is reported next to the two scores in the same risk_metrics block, but it enters neither the risk score nor the reliability score (backend/app/services/statistics/statistics_bridge.py:86-87). It does carry weight in the separate defensibility grade — see below.

Low reliability: the re-benchmark signal

A low reliability score is not a verdict on the tested party — it is a diagnosis of the pool. The pattern it usually describes: a thin set (small count, few years of data), a scattered set (high CV), or both. The conventional responses, in order:

  1. Broaden the search — geography, industry classification, size filters. A pool that is thin because the screen is over-tight is a search-design problem.
  2. Add years — multi-year averaging over more fiscal years grows the per-company data and stabilises the quartiles.
  3. Revisit the screens — a screen rejecting at an abnormal rate is the usual cause of thin pools; read it off the screening funnel rather than off the acceptance percentage.

The statistical significance of small samples is covered in the statistical significance glossary entry; the accept/reject record that a re-benchmark reworks is the grid documented in the Accept-Reject matrix reference.

How the scores feed the defensibility view

The risk and reliability scores are computed inside the analysis run and persisted with its statistics. The defensibility grade — an A–F letter score with a 0–100 composite — is a separate weighted computation in the predictive-risk service, reading the same persisted result (backend/app/services/predictive_risk/predictive_risk.py:412-573). It weights its own criteria, including sample size, range tightness, margin placement and the acceptance rate (bands at 70%, 50% and 30%), and it is part of the audit-readiness view, so it exists only for firms holding the predictive_risk plan feature; there the per-study panel on the results view (audit probability, litigation risk, defensibility grade with its criteria bars) sits next to the risk gauges, and the portfolio roll-up lives in the Predictive Risk Engine. Treat the audit-probability and jurisdiction weights as heuristic estimates, not measured enforcement statistics: a jurisdiction absent from the weight table scores a neutral 50.

FAQ

Are the risk and reliability scores the same number in opposite directions? No. Both look at CV, sample size, loss share and range tightness, but risk accumulates penalties from 0 while reliability dispenses credits from a base of 50, so they can diverge on thin-but-tight sets.

Can a within-range study have a high risk score? Yes. A thin, scattered set can produce a range the tested party falls inside; the scores say how much weight that conclusion carries, and that is precisely when a reviewer should ask for more comparables or a narrower search.

Do the scores change when I override a disposition in review? Yes — committing overrides rebuilds the persisted analysis result and recomputes the statistics on the effective accepted set, both scores included (backend/app/services/workflow/workflow.py:1716-1722). The raw counts on the cross-study results list are the run’s own and do not move. Re-review after a material override is not optional ceremony.

Do these scores certify the benchmark? No. They are a deterministic index of the accepted set’s dispersion, size, loss share and range shape. Nothing in them tests a hypothesis, and nothing in them reads the comparables’ business descriptions.

See it working in your workspace

Sign in to run the steps above on a real study — or book a demo and we will walk the workflow end to end.

Related docs

Book a Demo

Tell us what you'd like benchmarked

We'll confirm a 30-minute screen-share slot within one business day.

We reply within one business day. Your details are used only to arrange the demo — never shared or sold.