Quantitative Screening in Quartyl: Filters and Results
Quantitative screening in Quartyl: the exact filter order from multi-year completeness to RPT and revenue thresholds, Golden Rule aggregation, and the recorded rejection reason per company.
The quantitative screen is the first and cheapest gate: it runs on the numbers alone, before any enrichment or language analysis, and it decides which companies from the dump are financially plausible comparables. Its output is an accepted set, a rejected set, and - critically - a stored, specific reason for every rejection. This page documents the filter order, the aggregation rule, and how to read the results.
What the screen does
It runs as step 3 of the
pipeline, immediately after
the raw dump has been persisted to object storage. Input is the standardized
dump DataFrame; output is accepted_companies, rejected_companies, counts,
the column mapping and the detected years. Every rejected row carries a
Rejection Reason string that is specific and factual - not “not
comparable” - and that string is what shows in the grid and the reports.
Filter order
The checks run in three phases, and they behave differently:
- Phase A - multi-year rejections. Vectorised across the whole table, in this order, and each only fires on rows that have no reason yet.
- Phase B - the per-row rejection checker. It does not stop at the first
hit: every check it finds is collected and joined into one string with
"; ", so a company can be out for two stated reasons at once. Rows already carrying a Phase A reason are skipped. - Phase C - the configurable threshold filters. These are applied to the whole column and overwrite the stored reason, so on a company that was already rejected for something else, the last threshold it also breaches is what you end up reading.
| Phase | Check | What it tests | Recorded reason (examples) |
|---|---|---|---|
| - | Column detection and standardization | Required columns present and mappable (Company, Revenue, Cost, Operating Profit minimum) | Fatal at file level - names the missing columns |
| - | Multi-year detection and aggregation | Year-suffixed columns recognized; PLI computed per the Golden Rule | n/a (computation, not rejection) |
| A1 | Multi-year completeness | At least two years with valid operating profit and revenue | “Incomplete Multi-Year Financials: missing data for year(s) …” |
| A2 | Persistent loss | Operating profit negative in every valid year | “Persistent Operating Loss Maker” |
| A3 | Intangible intensity | Intangible assets over 10% of total assets | “High Intangible Asset Intensity (owns IP) (x / y = z%)” |
| B1 | Entity type | Cooperative, non-profit or similar entity types | “Cooperative/Non-profit entity (not comparable)” |
| B2 | Basic information | Company name and industry present | “Insufficient business information: missing …” |
| B3 | Intangible to revenue | Intangible assets over 30% of revenue | “High intangible assets (xM = y% of revenue)” |
| B4 | Status | Insolvency, liquidation or equivalent status | “Company under insolvency/liquidation proceedings” |
| B5 | Financial completeness | Required financial fields present and non-zero; no negative revenue, cost or operating expenses | “Insufficient financial information: missing …”, “Negative revenue (invalid financial data)” and similar |
| B6 | Extreme value on the selected PLI’s own ratio | The ratio the selected PLI’s band was calibrated for, outside that PLI’s band | “Loss making company” / “Extreme margin (x% on Berry Ratio) - unlikely to be comparable” |
| C1 | Related-party transactions | RPT above the study’s max RPT threshold | “High RPT” |
| C2 | Employee cost | Below the study’s minimum employee cost threshold | “Low employee cost” |
| C3 | Revenue floor | Below the study’s minimum revenue threshold | “Below minimum revenue” |
| C4 | Revenue ceiling | Above the study’s maximum revenue threshold | “Above maximum revenue” |
The Phase C rows are the configurable financial filters - minimum and maximum revenue, maximum RPT share, minimum employee cost - taken from the study’s parameters and recorded on the study as part of its configuration. A threshold left at its sentinel default is not applied at all.
There is no industry or geography filter in the run. The filter function does accept an industry argument, but the pipeline never passes one: the country and NACE/industry fields on a comparable are metadata carried through to the grid and the reports, not a row filter. If your screen must be limited to a jurisdiction or a sector, that is decided by what you upload and by your review dispositions - not by a control on this step.
Statistics are not computed here either. The range is built at final analysis, over the accepted set as it stands after review.
The Golden Rule aggregation
When the dump carries multiple fiscal years, Quartyl never averages per-year margins. It aggregates raw financials first:
aggregated PLI = sum(numerator across years) / sum(denominator across years)
Year columns are detected by pattern (for example Revenue_2022,
Revenue FY22, Revenue (2022)), the per-year numerators and denominators
are computed for the selected margin type, and the averaging method chosen in
the wizard (simple or weighted) is applied consistently to the tested party
and the pool. The detected years are recorded in the dump metadata, so the
aggregation window is part of the study’s record.
The multi-year rejections above run on those per-year columns. After they have
spoken, the single-value Revenue, Operating Profit and Cost columns are
replaced with the maximum across the available year columns - deliberately
lenient, so a company that is missing the most recent year but has earlier data
is not thrown out for a gap in the dump. Every size-based check after that
point - the revenue floor and ceiling, financial completeness, the threshold
filters - reads that strongest year.
Which companies survive, and why
- The accepted set is the population that passed every check with no reason recorded. These companies flow on to enrichment and qualitative screening; they are candidates, not yet decisions - the grid is where the accept/reject of record happens.
- The rejected set is everything that failed at least one check. The
stored string names what fired: the per-row checker accumulates all of its
findings separated by
"; ", while the multi-year checks and the threshold filters each write a single reason - and a threshold filter that also matches will replace whatever was already there. So a company that is both under the revenue floor and high-RPT shows one reason, and a checker-stage company can show two. - The multi-year checks are PLI-agnostic. They read only the universal operating-profit and revenue year columns, never the selected margin’s numerator and denominator, so switching PLI does not change which companies are short on years, persistently loss-making or intangible-heavy.
- The extreme-value gate is not PLI-agnostic - and it should not be. Each
PLI has its own band and is measured on its own ratio:
OP/OCfrom -5% to 100%,OP/Sales,TP/Sales,EBIT/Salesand plain operating margin the same,EBITDA/Salesto 150%,Berry Ratioto 3.0, andEBIT/Total Assetsto 50%. Reading a margin against an asset-return band would make the verdict depend on which PLI you happened to select rather than on the company, which is the mistake the earlier version of this gate made. Where the ratio cannot be formed - a required column absent, or a zero denominator - the check declines to evaluate rather than reject, leaving the company to the semantic and AI layers.
How to read a rejected company
Three questions, in order:
- What is the reason? Read the stored string - it names the check and, where relevant, the numbers (the missing years, the intangible ratio, the margin).
- Does the data actually show it? Open the dump - the raw numbers behind the decision are persisted and inspectable, per Raw Dump Data.
- Is the company genuinely out? If you believe the engine is wrong, the correction is a reviewer override with a recorded rationale on the comparables grid - not an edit to the filters after the fact.
The method-side counterpart - how these screens map to OECD comparability discipline and what a TPO expects - is in the knowledge guide Quantitative Screening.
FAQ
Why are persistent loss makers rejected outright? A company that loses money every measured year is not a valid arm’s length benchmark for a profitable tested party, and multi-year data makes the call unambiguous - one bad year is survivable, every year is not.
Do the financial thresholds reject on one side only? Each filter has a one-sided threshold by design: revenue has a floor and a ceiling, RPT has a ceiling, employee cost has a floor. They encode size and structure similarity, not a score.
What happens to a company with missing description text here? Nothing - description quality is not a quantitative check, and the checker deliberately does not reject on it. Missing, placeholder or too-short descriptions flow through to the qualitative screen, which flags them for review with “Missing/Insufficient Business Description” rather than scoring them on absent text. Description quality never rejects a comparable anywhere in the run; it only ever defers it to you.
See it working in your workspace
Sign in to run the steps above on a real study — or book a demo and we will walk the workflow end to end.
Related docs
Raw Dump Data: Inspection and Reuse After Screening
Raw dump data in Quartyl: the parquet persisted after quantitative screening, inspecting the numbers behind every filter decision, reuse across re-runs, and the per-tenant retention window.
Read docThe Screening Pipeline: Quantitative to Qualitative (8 Steps)
The eight-step async analysis pipeline in Quartyl: quantitative, qualitative and AI screening, gated web research and deep screening, final analysis, per-step progress.
Read doc