Skip to main content
Quartyl
Benchmarking in Quartylprofessional

AI Screening: How It Scores Companies and Cites Evidence

AI screening in Quartyl: LLM FAR analysis of each comparable against the tested party, verbatim evidence quotes, the outcome categories, human override and the gated deep pass.

Quartyl Team

AI screening is step 5 of the pipeline, shown in the stepper as AI Screening (FAR Analysis): the language-model pass that checks, company by company, whether the surviving comparables actually do the same work as the tested party. It runs after quantitative and qualitative screening, so it sees a population that has already passed the numbers and the description-similarity gate. Its output is not a verdict of record - it is a structured recommendation with evidence, waiting for a human.

Where it runs and what it sees

The input to the step is the union of the qualitative screen’s accepted and almost-accepted companies. Rejected companies and hard flags do not go through the LLM.

This step sends data to a third party. Company names and their description text are submitted to the configured inference provider - an OpenAI-compatible endpoint set by AI_PROVIDER_BASE_URL, with the model slug set by OPENAI_MODEL (default qwen/qwen3.8-max:free). If your client, your jurisdiction or your internal policy restricts where comparable-company data may be processed, point that base URL at a provider you are allowed to use before you run a study. The qualitative step before this one is fully local; this one is not.

Per company, the model is given:

  • The company name.
  • The assembled description from the dump - full overview, business description / oneliner and products-and-services text, with the registry trade description demoted to a last resort because it describes what a company is permitted to do rather than what it does. Truncated at 2,048 characters.
  • An Independence / related-party line where the dump carries it: the company’s related-party transaction share as a percentage of revenue printed next to the study’s own threshold, and the vendor independence indicator with its letter class decoded into prose (“a majority shareholder exists”, “controlled by a shareholder”). Both are fields from the extract you upload, not a registry lookup, and the line simply is absent when the dump carries neither - which the model is told to ignore rather than read as independence either way.

The answer is constrained by a strict JSON schema: inferred functions, assets and risks, an accept/reject decision, a standardized outcome category, a confidence level, a reason, and a verbatim evidence quote. Requests are batched at 20 companies per call, at temperature zero, with a 120-second call budget and two retries; batches run concurrently. Batching is what keeps a large study inside a sane wall-clock time, and the company index the model echoes back is checked positionally, so a verdict cannot be attached to the wrong row.

Plan feature: The deep second pass described below requires the advanced_ai_screening plan feature (see Plan Features).

How each company is scored

The rules are industry-agnostic and relative to the tested party’s profile - the study’s screening profile supplies the domain content, not the prompt. Stated as the step applies them, in order:

  1. Evidence first. If the description does not carry enough to decide, the answer is Flagged, not a guess.
  2. Caller exclusion criteria. Any exclusion criterion you set on the study is injected as labelled bullets; a company that clearly triggers one is rejected against that label.
  3. Independence and related-party pricing. A company whose related-party share sits above the study’s own threshold cannot be accepted automatically: it comes back Flagged at low confidence for you to decide. A margin set by intra-group pricing is not arm’s length. Group membership alone is not a disqualifier - most companies in a database have a shareholder, and a group subsidiary is an ordinary comparable - so the independence indicator is weighed alongside function, product and risk rather than treated as a verdict. Absent data is not a finding either way.
  4. Product category. A wholly different core product or service category is Non Comparable Product. Selling extra categories on top of the same core line is not a mismatch - it is Almost Accepted.
  5. Function and business model. The core business model must be the same type as the tested party’s, judged against the accepted and out-of-scope models, the route to market and the product scope you recorded on the study. A company whose primary activity sits outside them is Non Comparable Function. This is where a wholesale-only tested party screens out consumer-facing comparables, and where a tested party that does R&D is not penalised for it - the direction follows from your profile, not from a built-in rule about B2B, B2C or R&D.
  6. Additional operations. Core function and product match, but broader scope - extra lines, extra channels, extra functions - is Almost Accepted, never a rejection: the extra complexity is exactly what a human should weigh.
  7. Secondary attributes. Geography and IP ownership are compared to the profile. Divergence on these is grounds to flag rather than reject, unless you declared it an exclusion criterion. Owning some brand or software is not a rejection; a business whose value comes from IP it licenses is a different functional profile.

The standardized outcomes

Every company comes back in one of five outcome categories - this is the “standardized reason id” of the AI pass:

Category Meaning Bucket
Accept FAR profile sufficiently matches the tested party Accepted
Almost Accepted Core match, but broader scope than the tested party, or accepted at reduced confidence Almost-accepted (borderline)
Non Comparable Function Business model or function profile fundamentally different Rejected
Non Comparable Product Core product line entirely different Rejected
Flagged Cannot be decided from the evidence, or independence/related-party data bars an automatic accept Flagged for review

Each result also carries a confidence level (high, medium, low) and a reason text naming the specific misalignment or alignment. The reason text lands on the comparable as its rationale, and it is what you read in the grid.

When the pass cannot produce a usable verdict - no description to judge, an API failure, or an answer that cannot be evidenced - the company is flagged for manual review. It is never written up as accepted, and never marked accepted “with low confidence” as a placeholder for a decision the machine did not make. An unavailable screen is visible as unavailable.

Evidence per company

Each screened comparable persists its evidence alongside its row:

  • The inferred functions, assets and risks.
  • The outcome category and the confidence level.
  • The reason text (the rationale shown in the grid and reports).
  • A verbatim evidence_quote - the exact span of that company’s own description the decision rests on. It may be empty only for Flagged.
  • A write-once AI-recommendation event in the study’s evidence timeline, recorded with the system as the actor and the rationale as its content.

The quote is checked, not trusted: it has to be locatable in the description that was sent for that company. A verdict whose quote does not match is discarded and the company is flagged as unevidenced - which is also what makes batching safe, since a model that mixes up two rows in a batch gets caught rather than published.

The reason text also feeds the ground-truth machinery: standardized reason strings carry a deterministic reason id, so when a reviewer later overrides an AI decision, the override is linked to the exact reason being corrected, and the cross-study registry can show which AI reasons reviewers systematically disagree with. See Overrides.

What the AI cannot do

  • No judgment of record. The AI recommendation is an input to the review, not the decision. The comparable’s effective status is whatever your grid says; the engine status stays visible underneath it.
  • No free-form excuses. The schema forces category, confidence, reason and quote - there is no path where the model rejects a company without a stated, structured, traceable cause.
  • No overrides of its own. The model cannot accept a company the qualitative screen flagged, or resurrect one the quantitative screen rejected. It can only refine the population it was handed.
  • No invention of facts. It sees the description text from your dump and, where the extract carries it, the related-party line above - and nothing it cannot find in that text may carry a verdict.

The final decision is made by a human in the comparables grid, and any decision that goes against the recommendation is recorded as an override against the engine’s reason.

The deep pass: advanced AI screening

Gated by the advanced_ai_screening plan feature, this is step 7 - Advanced AI Screening (Deep FAR) - and the final AI phase. A second, separately configured model (OPENAI_ADVANCED_MODEL, default mistralai/mistral-large-2512) takes everything the first pass kept or flagged, and receives the same tested-party profile and the same screening text as Phase 1 - plus, where the gated Web Research step gathered it, the labelled Web research: findings. A second opinion decided on fewer facts would not be a second opinion; that is also why the flagged companies are in the input, since they were flagged for carrying too little to judge and the web may have supplied what the record lacked. Per company it produces:

  • A granular functions/assets/risks breakdown - specific functions (contract R&D, toll manufacturing, distribution, licensing, treasury), specific IP and capacity, specific economic risks (inventory obsolescence, credit, R&D failure, FX, regulatory).
  • An explicit acceptance reason: why this comparable is suitable for this tested party, referencing the specific FAR alignment points.
  • A confidence score from 0.0 to 1.0 on the match.
  • A final verdict in the same five categories, with a rejection reason where applicable.

Unlike the first pass, this phase is called once per company (90-second request budget, two retries) rather than in batches, and it is fail-closed: a company with no description, an API error, or an unevidenced verdict is moved out of the accepted set and flagged - it is never carried through on the strength of an analysis that did not happen. “Could not decide” is kept in the flagged bucket rather than merged into rejections, because those are different findings. A partial failure is not fatal; only a run where every company fails raises, which marks the study FAILED so a misconfigured model surfaces as a failed study rather than a report whose deep screening never ran. Firms without the plan feature simply end the AI phase at step 5.

The method context - where AI-assisted screening sits in the OECD comparability framework and how a reviewer should treat its output - is in the knowledge guide AI-Assisted Comparable Screening.

FAQ

Does the AI see the financials? No PLI is put in front of it - the numbers were already handled by the quantitative step, and this pass is about functional comparability. One exception is worth knowing: the per-company line the model receives includes the related-party transaction share against your threshold, because that is comparability evidence rather than a margin.

Why batch the companies? One structured call for 20 companies amortizes request overhead and keeps per-study cost bounded. Batching does not make failures contagious: when a batch call fails after its retries, every company in it is flagged immediately rather than retried twenty times in sequence; when a batch responds but an individual entry is unusable, only that company is retried on its own.

Can I re-run only the AI step? No - the run is one pipeline. Re-run the study (re-screen); the pipeline is idempotent, verdicts are cached per study so a re-run of the same study does not re-pay for companies whose screening text and profile are unchanged, and the grid keeps your reviewer decisions visible as overrides.

See it working in your workspace

Sign in to run the steps above on a real study — or book a demo and we will walk the workflow end to end.

Related docs

Book a Demo

Tell us what you'd like benchmarked

We'll confirm a 30-minute screen-share slot within one business day.

We reply within one business day. Your details are used only to arrange the demo — never shared or sold.