Skip to main content
Quartyl
Glossary

Search String: The Benchmark Query That Finds Comparables

The search string defined: the database query — industry codes, keywords, size and geography filters — that builds the candidate pool, and how to keep it neither over- nor under-inclusive.

Quartyl Team

Definition

The search string is the study’s first filter, expressed as a database query — the combination of industry classification codes (the NIC code or the NACE family for the tested party’s activity), the keyword terms (the product and activity words: “trading”, “distribution”, “contract manufacturing”, the segment language), and the structural filters (the [size filter] (/docs/glossary/size-filter) — the turnover/asset bounds; the geography — the search regions; the reporting currency and period) that the search design fixes before any result is seen. Its output is the candidate population — the set of companies the screens will work on — and its quality bounds the study: a search that is over-inclusive (the keywords too broad, the codes too high-level) returns companies that cannot be comparable no matter how well the qualitative screening is done (waste, and the [accept-reject matrix] (/docs/glossary/accept-reject-matrix) full of rejects that read as noise); a search that is under-inclusive (the keywords too narrow, one code, one region) returns a population so small that the [comparable set] (/docs/glossary/comparable-set) cannot reach the statistical significance floor. The string is designed and frozen first — the [design] (/docs/benchmarking/search-design) is documented in the study before the results are examined, which is what separates a search from a fishing expedition in the examination’s eyes.

The search string, in one query:
  1. The classification (the NIC/NACE codes for the activity — 1 family, not 1 code)
  2. The keywords (the product/activity terms — the segment language)
  3. The size bounds (the turnover/asset floor and ceiling — the size filter)
  4. The geography (the regions the search covers)
  5. The periods (the financial years the data must cover)
The element The content
The classification The NIC/NACE codes for the tested party’s activity (the 3-digit/4-digit family — broad enough to catch the segment, narrow enough to keep it relevant)
The keywords The product and activity terms that define the tested party’s function (the FAR expressed in search language — “distribution”, “trading”, the product family)
The size filter The turnover/asset bounds that keep the comparables in the tested party’s size band (the size filter — typically a floor and a ceiling around the tested party)
The geography The search regions (the regional vs local decision — the home market, the region, the multi-region search)
The periods The financial years required (the multi-year averaging horizon — 3 years of data for a 3-year study)

The working read (the search design guide): the string is a hypothesis about where the tested party’s comparables live, and like any hypothesis it is checked against its own failure modes — the over-inclusion (too many candidates, the segment diluted) and the under-inclusion (too few, the set short of significance). The discipline: the string is written from the tested party’s FAR, not from the database’s categories (the tested party’s function drives the terms; the codes are the classification’s way of saying it); the string is documented in full (every code, every term, every bound — the Local File benchmarking annex carries it); and the string is frozen before the results (a string that is re-tuned to change the population after the range is seen is the search-version of the set-stability objection). Where the search is run in Quartyl, the search design (the NIC codes via the NIC finder, the keywords, the size and region filters) is captured on the study and the pipeline’s candidate list is the string’s output — the record and the population together.

Example

A routine distributor of packaged foods (Indian tested party). The search string:

Element The value
Classification NIC 4632 (retail sale of food, beverages and tobacco) + 4639 (retail sale of other articles) — the family, not one code
Keywords “trading”, “distribution”, “packaged food”, “FMCG”, “staple” — the tested party’s product language
Size filter Turnover 0.5×–5× of the tested party (the size band — the size filter bounds)
Geography India first; the string records the regional extension (South Asia) as the fallback if the Indian population is thin
Periods FY23–FY25 (the 3-year horizon, matched financial years)

Run: 74 candidates in India. That number is the design’s output, seen for the first time now — the [quantitative screens] (/docs/benchmarking/quantitative-screening) then take over. Had the string used only NIC 4632, the population would have been 31 (under-inclusive for a 3-year set); had it added the generic keyword “food”, it would have pulled in producers (over-inclusive — the manufacturing function is not the tested party’s). The string’s two failure modes, checked in the design itself.

See also

FAQ

Is the search string the same as the screening filters? No — the string is the population query (the classification + keywords + bounds that define which companies are candidates); the [quantitative screens] (/docs/benchmarking/quantitative-screening) are the filters applied to the candidates (the profitability filter, the outlier removal, the period completeness). The string runs once, against the database; the screens run on the string’s output, inside the study. In the file they are documented together (the search design section) but they are different decisions — and a string that smuggles in a screen (a profitability condition in the query itself) is a design error: it hides an exclusion behind the population.

How wide should the classification be — one NIC code or a family? A family (the 2–3 sibling codes that cover the tested party’s activity), unless the activity is narrow enough that one code is exact. The rule: the classification should cover what the tested party does (the function), not what the tested party is called (the legal description). A distributor spanning food and home-care lines needs the family (4632 + 4649); a pure staple trader may be exact on one code. The test is the same as the keyword test: would a genuinely comparable company be missed by the narrower version? If yes, widen — the over-inclusion is cheaper to screen out than the under-inclusion is to fix (the comparable set cannot be built from candidates the search never returned).

Why must the string be frozen before seeing results? Because the examination’s standard reading of a re-tuned string is that the search was aimed at the outcome. The sequence that survives: (1) the FAR says what the tested party does → (2) the string encodes it (codes, terms, bounds, regions) → (3) the population is produced → (4) the screens decide. The sequence that does not: see the range, find it uncomfortable, adjust the keywords to change the population, re-run. The benchmarking mistakes checklist carries the frozen-string discipline as a line item — the string in the Local File is the pre-result version, and its date says so.

Run the screens as a study, not a spreadsheet

Quartyl applies the method, PLI and screening steps above as a pipeline — and keeps a documented reason for every exclusion.

Related docs

Book a Demo

Tell us what you'd like benchmarked

We'll confirm a 30-minute screen-share slot within one business day.

We reply within one business day. Your details are used only to arrange the demo — never shared or sold.