Skip to main content
Quartyl
Benchmarkingprofessional

AI-Assisted Comparable Screening: How It Works

How AI-assisted comparable screening works in a benchmarking study — what the model evaluates, what it must cite, and the human controls that keep the Accept-Reject matrix defensible.

Quartyl Team

Screening a candidate pool company by company is the most labour-intensive step in a benchmarking study — and the step where the quiet errors hide. AI-assisted screening changes the economics: the review of every candidate company, against the tested party’s profile, at a cost close to zero. It does not change the standard: the Accept-Reject matrix still has to be defensible company by company, and the question the guide answers is where the machine stops and the professional starts.

What the screening problem actually is

For each candidate company, the screen answers: would an unrelated party in this company’s position be an appropriate economic comparator for the tested party? That is a comparison across the comparability factors — products and revenue mix, functions performed, assets and intangibles owned, risks borne, market and economic environment — applied to the evidence that exists for the company: its business description, segment disclosures, product listings, ownership and registration data.

Most of the evidence is textual and heterogeneous: a business description in the annual report, a product page on the website, a segment note, a registry entry. That is precisely the workload a language model handles well — and precisely the workload where an uncontrolled model also fails well: confidently, and at scale.

How it works in a governed study

A defensible AI screening step has four properties:

1. The input is the tested party’s profile, not a prompt

The model is given the tested party’s characterized profile — functions, assets, risks, products, revenue mix (the FAR profile locked in the study setup) — and each candidate company’s evidence set. The comparison is profile versus evidence, run identically for every candidate. A screen that runs on a hand-written prompt per company is not a method; it is a series of opinions with a machine speed.

2. Every decision carries a cited reason

The output per company is one of four dispositions — accept, almost accept, reject, or flag for review — with the specific reason grounded in the evidence: “reject — company derives 60%+ revenue from its own product line; tested party is a pure contract provider”, “almost accept — same distribution function, but it also manufactures, so the scope differs from the tested party’s”, or “flag — business description does not disclose the revenue mix; segment data required before disposition”. Almost accept and flag are recorded as their own outcomes rather than folded into accept or reject: one is a comparable company on a different scope, the other is a question, and neither is a rejection the reviewer did not make. The reason must point at what was seen, not at a general impression. A matrix full of “not comparable” without the substance is the exact matrix the TPO dismisses.

3. The human review is a designed step, not a formality

The professional’s job after the AI pass is not to re-read everything — it is to review the decisions that matter: every reject (is the reason real and grounded?), every accept of a company with material red flags (size gap, related-party revenue, adjacent function), every almost accept (is the scope difference tolerable, or is the function really different?), and every flag (resolve with the missing evidence or reject with the stated reason). The review is logged — who reviewed, what was changed, why.

4. Overrides are governed, and the audit trail is the point

Where the human reverses the AI’s disposition — keep a company the model rejected, or reject one it accepted — the override is recorded with its rationale, separate from the model’s output. The same discipline applies in reverse: the model’s output that the human confirms is an accepted AI disposition, and in a defensible file the ground truth is traceable — for any company, you can state what the model said, what the professional decided, and the reason. That trace is what the examination asks for when the matrix is challenged, and it is what separates AI-assisted screening from AI-replaced screening.

What AI cannot do (and must not be asked to)

  • Verify the numbers. The screen works on descriptions and disclosures; the financial lines come from the filings and the database reconciliation. A model’s “this company has a 14% OP/OC” is not a number the study can use.
  • Judge risk ownership from a website. Who bears the inventory risk, who controls the IP, who funds the working capital — those come from contracts and the functional record, and they are professional judgments.
  • Resolve genuine ambiguity. Where the evidence does not decide comparability, the honest output is flag, not a coin-flip dressed as a confidence score.
  • Replace the comparability standard. The factors are the OECD/Rule 10D factors. The model applies them; it does not define them, and a file that lets the model “decide the standard” has no method left to describe.

The control set, summarized

Control What it catches
Profile-driven input (not per-company prompts) Drift between companies; inconsistent standards
Cited, evidence-grounded reasons Vacuous dispositions; hallucinated facts
Human review of rejects + flagged companies The errors that would kill the matrix
Governed overrides with recorded rationale Silent human capture of the method; untraceable decisions
Traceable ground truth (model output vs final decision) The examination’s core question: what was decided, by whom, on what evidence
Sampling re-verification against primary sources Model facts that do not survive contact with the filing

Where it fits in the study

The screening step sits between the quantitative thresholds and the range: search design → candidate population → quantitative filters → AI-assisted qualitative screen (profile vs evidence, cited dispositions) → human review and governed overrides → final pool → adjustments → range. The matrix the file presents is the post-review matrix — and its defence is the trail behind it, not the model that produced the first draft.

The same discipline is why the platform’s screening surfaces every disposition with its reason in the comparables grid, records each override with its rationale in the ledger, and keeps the model’s ground-truth output attached to the study — so the matrix that goes into the Local File is the matrix the audit can replay, company by company.

See also

Run the screens as a study, not a spreadsheet

Quartyl applies the method, PLI and screening steps above as a pipeline — and keeps a documented reason for every exclusion.

Related docs

Book a Demo

Tell us what you'd like benchmarked

We'll confirm a 30-minute screen-share slot within one business day.

We reply within one business day. Your details are used only to arrange the demo — never shared or sold.