Skip to main content
Quartyl
Predictive Riskprofessional

Predictive Risk Engine: Audit, Litigation, Defensibility

Three deterministic risk scores per study: an audit exposure index, a dispute index and the A-to-F defensibility grade, read together as a triage set.

Quartyl Team

The Predictive Risk Engine (the Predictive Risk Engine item in the sidebar, under Intelligence & Analytics) scores every study that has a successful analysis on three axes: how exposed the recorded position looks to a tax authority, how likely it is to end up in a dispute, and how well the benchmarking package would defend it.

What these scores are: deterministic rule-based indices over the data the study already recorded — weighted sums over fixed thresholds. Despite the product name, there is no trained model, no fitted coefficients and no backtest behind them. “Audit probability 68” is not a 68% chance of an audit; it is the position of that study in your own queue. Nothing here is validated against observed audit or litigation rates.

The three scores

Plan feature: This capability requires the predictive_risk plan feature (see Plan Features).

Score (UI label) Scale How it is built Bands
Audit Probability 0-100 index Weighted sum of five scored factors — jurisdiction, the tested-party margin’s deviation from the pool median, related-party intensity in the accepted set, the loss-making share of the pool, and CV — with sample size reported unscored. The jurisdiction value is a heuristic enforcement-intensity estimate; a country with no entry scores a neutral 50, i.e. unrated High Risk at 60 and above, Moderate Risk at 35 and above, Low Risk below 35
Litigation Probability 0-100 index Half the audit index plus 30% of the jurisdiction’s litigation value, plus 15 for a below-range conclusion (8 for above-range), plus 10 when the recorded benchmark reliability is under 40 Critical at 65 and above, High at 45 and above, Moderate at 25 and above, Low below 25
Defensibility grade A-F (with a 0-100 score) Weighted mean of five criteria: sample size, dispersion (CV), range tightness (IQR to median), margin placement and acceptance rate A at 85 and above, B at 70 and above, C at 55 and above, D at 35 and above, F below 35

The defensibility grade is the one a tax officer can actually inspect — it grades your file. The other two are internal triage indices that say which files to look at first; treat them as a ranking, not a forecast.

How the scores are read together

The three scores answer different questions, and they are read as a set:

  • Audit Probability — does this file look like the kind of file that attracts attention? Driven by the jurisdiction entry, how far the tested party’s margin sits from the pool median, the related-party intensity recorded on the accepted comparables, the loss-making share of the pool, and the pool’s variability.
  • Litigation Probability — given the audit index, does the rest of the record make a dispute more likely? It inherits half of its value from the audit index, adds the jurisdiction’s litigation value, and takes a fixed uplift for a below-range (or above-range) conclusion and for a low recorded benchmark-reliability figure. It adds nothing new about the pool itself.
  • Defensibility grade — when the file is read, what does it say about itself? A weighted grade over sample size, dispersion (CV), range tightness (IQR to median), margin placement and acceptance rate.

The combination to act on is a high audit index with a weak defensibility grade: the file scores as exposed on the heuristics, and it is the least able to withstand scrutiny on the criteria that can be checked. Low audit with a strong grade is the position to leave alone; high audit with a strong grade is the one to document and keep current.

What each study shows

Selecting a study in the row of study buttons re-renders the whole lower section for it:

  • Audit Probability gauge + “Audit Risk Factor Breakdown” — the score and its band, then each factor with its value out of 100, its weight (shown only where the weight is non-zero), its status colour (high / moderate / low) and a plain-language description line, under the summary sentence.
  • Litigation Probability gauge + “Litigation Probability Analysis” — the score and level, the contributing-factor list and the mitigation-suggestion list. The contributing factors are labels displayed beside the number: the high-CV and small-set items are flagged there but carry no weight in the litigation arithmetic, which only moves on the audit index, the jurisdiction value, the conclusion and the reliability figure.
  • Defensibility gauge + “Defensibility Scorecard” — the 0-100 score and the letter grade with its label, then each criterion with its score, a bar and its detail line, and the range analysis: the 25th-75th figures, the IQR, the CV and the sample size.
  • View Full Study — jumps to the study itself, where the comparables and the dispositions live.

The factors are human-readable — they are the explanation of the number, not a second number.

The portfolio view

The same three scores, aggregated across the studies the firm has completed: the high / moderate / low audit counts, the average audit index and the average defensibility score — the portfolio risk view that ranks which study to fix first.

What the engine does not do

  • It does not write to the study from this page — the risk endpoints derive their scores from the stored analysis data at read time, so the figures you see here never desynchronise from the file. Separately, when the feature is on, the analysis run embeds its own risk snapshot in the study’s results for the results-page insights; that snapshot is taken mid-run, before the final statistics pass, so the insights copy can differ from what this page computes. This page and the API are the live view.
  • It does not adjust a disposition — the lever is always the analysis (the pool, the parameters, the conclusion), never the number.
  • It does not score studies without a successful analysis — drafts, failed runs and studies whose stored result is not marked successful are skipped, not zeroed.
  • It does not invent inputs. The scores read only what the study recorded: the parameters (jurisdiction, TP margin) and the analysis result (the statistics block, the accepted comparables, the conclusion, the recorded counts). A figure the study never stored cannot move a score — and where a rule does have a fallback it is a neutral one (missing margin data scores 50, a jurisdiction with no table entry scores 50, a missing reliability figure reads as 50).
  • It does not see the rejected comparables or the override ledger — no score consumes either of them.
  • It does not output company-level detail. The response ends at the factor and criterion level, plus the accepted-comparable count.

FAQ

Is a score of 68 a 68% chance of an audit? No. The scale is a 0-100 index built by adding weighted factor points, and the bands are fixed cut-offs. It orders your files against each other; it does not estimate a rate of real-world events, and it has not been validated against any observed audit or litigation outcomes.

Are the scores live or stored? Derived, never stored. The portfolio summary is computed on read and then cached per tenant (per owner for a Superadmin), and study transitions clear that cache domain — so a re-run surfaces on the next read once the study has moved state, with an analytics TTL as the backstop. A single study is recomputed on every read of its own endpoint, with no cache in front of it.

Can I adjust a score by hand? No. The scores are derived; the lever is the analysis (the pool, the parameters, the conclusion), not the number.

Why is the sample size informational in audit probability? It is reported without a weight in the audit index, but it is a fully weighted criterion in the defensibility grade — it flags a thin pool in one and grades the documentation quality in the other. See the factor model.

See it working in your workspace

Sign in to run the steps above on a real study — or book a demo and we will walk the workflow end to end.

Related docs

Book a Demo

Tell us what you'd like benchmarked

We'll confirm a 30-minute screen-share slot within one business day.

We reply within one business day. Your details are used only to arrange the demo — never shared or sold.