Skip to main content
Quartyl
Benchmarking in Quartylprofessional

The Screening Pipeline: Quantitative to Qualitative (8 Steps)

The eight-step async analysis pipeline in Quartyl: quantitative, qualitative and AI screening, gated web research and deep screening, final analysis, per-step progress.

Quartyl Team

A study run is one background job walking through eight named substates, from QUEUED to FINAL_ANALYSIS. Each substate is what the progress stepper shows you, and each step is scoped so that a failure, a retry, or a re-run produces the same settled record rather than a different one. This page maps the full path and what each step does to the comparable population.

How a run executes

When a study is submitted for analysis, a job is created in the analysis queue and a worker picks it up. The worker runs inside a single database session, mutates one in-memory results structure step by step, and commits at each major transition so the progress view never lies about where the run is. DB persistence of the comparable set happens once, at the end.

Plan feature: This capability requires the advanced_ai_screening plan feature (see Plan Features).

Plan feature: This capability requires the predictive_risk plan feature (see Plan Features).

The two blockquotes above mark the steps in the table that are plan-gated. For a tenant whose plan does not hold the feature, the step does not run and the run continues; the stepper does not hide it, either - the step stays in place and is labelled “(Not in your plan)”, so the record always shows that a stage was available and not used rather than leaving you to infer it. A tenant with no assignable plan has all gated steps denied; tenant-less platform accounts run every gated step.

# Substate Progress What happens Failure policy
1 QUEUED 0% Job sits in the Redis analysis queue waiting for a worker N/A
2 PREPARING_DATA 0-15% File downloaded from storage, size/sheet/parse validated, required columns detected Fatal
3 QUANTITATIVE_ANALYSIS 15-35% Raw dump persisted as parquet, then column detection, multi-year aggregation, hard rules, financial filters Fatal
4 QUALITATIVE_SCREENING 35-65% Company-by-company comparability scoring on a local embedding model, no LLM tokens, with recorded reasons Fatal
5 AI_SCREENING 65-70% LLM FAR analysis of accepted and almost-accepted comparables against the tested party profile Retried once; a company whose call fails is flagged, not dropped
6 WEB_RESEARCH 70-76% Each candidate’s own website and public records of its group structure are read as labelled web evidence (gated by advanced_ai_screening, the same flag as the deep pass below) Never fatal - a site it cannot read simply yields no findings
7 ADVANCED_AI_SCREENING 76-82% Deep FAR pass over the Phase-1 accepted set, one call per company, with acceptance reason and confidence score (gated by advanced_ai_screening) Fail-closed per company; fatal only if every company fails
8 FINAL_ANALYSIS 82-100% Inline predictive-risk analytics (gated by predictive_risk, best-effort), statistics, comparables persistence Fatal for statistics; risk analytics non-fatal

Steps 1-2: Queued and preparing data

The queue is capped: when the analysis queue is full, submissions are rejected with a 429 rather than silently queued forever. Once a worker takes the job, PREPARING_DATA validates the input before anything is computed:

  • File size within the 50 MB limit.
  • A resolvable worksheet (for Excel) and a non-empty file.
  • The required columns detectable - at minimum Company, Revenue, Cost and Operating Profit. A file that fails this is rejected with the missing columns named.

Validation failures fail the study immediately; they are not retried.

Step 3: Quantitative analysis and dump persistence

At the start of this step the full raw dump is written as a normalized parquet to object storage under dumps/{tenant}/{study_id}/dump.parquet, with the storage key, row count, column mapping and detected years recorded on the study. Persisting the dump before screening runs is deliberate: later steps and the report rebuild can always retrieve the exact numbers, and a failure in a later step cannot destroy the source data. See Raw Dump Data.

The screen itself then does column detection and standardization, multi-year PLI aggregation under the Golden Rule (sum of numerators over sum of denominators - never averaging of per-year margins), multi-year rejection checks, rule-based hard rejects, and the configurable financial filters. Every rejected company gets a specific stored reason. The full filter order and the acceptance set semantics are in Quantitative Screening.

Step 4: Qualitative screening

Each accepted comparable’s business description is encoded with a local, quantized embedding model - sentence-transformers mxbai-embed-large-v1, int8 ONNX - and scored against the tested party’s description. This step is entirely on your own infrastructure: it spends no LLM tokens and sends no company data to any third party, which is why it sits before the pass that does.

The score is a composite of semantic similarity (0.60 weight) and keyword overlap (0.35 weight), plus small bonuses for company-name and industry-column alignment. The composite maps to ACCEPT (0.42 and above), ALMOST_ACCEPT (0.36 and above), FLAG (0.34 and above) or REJECT (below that). Layered on top, a function or product keyword mismatch rejects only if the raw semantic similarity is also below 0.55 - a company whose text genuinely resembles the tested party’s is not thrown out over a missing keyword. Companies with no usable description are flagged for human review - never rejected on that alone - and cooperative or charity entities are rejected with a recorded reason, because their disqualification is about legal form rather than dissimilarity. What leaves the step is four buckets - accepted, almost-accepted, flagged and rejected - each company with its specific reason. The scoring detail is in the knowledge guide Qualitative Screening.

Step 5: AI screening (FAR analysis)

Shown in the stepper as AI Screening (FAR Analysis). The accepted and almost-accepted comparables go through FAR analysis - Functions, Assets, Risks inferred from the description and compared against the tested party profile - through the OpenAI-compatible gateway you have configured, at temperature zero, 20 companies per call. This is the step where data leaves the platform: company names and their description text are submitted to that third-party provider, so set the base URL and model to an endpoint your engagements permit.

The text the model judges is the dump’s own description fields - overview, oneliner, products and services, with registry trade description as a last resort - plus, where the upload carries it, an Independence / related-party line built from your own figures: the company’s related-party transaction share printed beside the study’s threshold, and the vendor independence indicator with its letter class decoded. Where neither field is present the line is absent, and absence is never read as independence. Each verdict is constrained to a strict JSON schema and one of five outcome categories: Accept, Almost Accepted, Non Comparable Function, Non Comparable Product, or Flagged, and it must quote the company’s own description verbatim. A quote that cannot be found in that company’s text, a failed API call or a description too thin to decide on all produce Flagged - never a quiet accept. Per-company evidence is captured and persisted with each comparable. See AI Screening.

Step 6: Web research

Shown as Web Research, and gated on the advanced_ai_screening entitlement - the same single flag that gates the deep pass after it, since one capability covers both. Where it runs it reads each candidate’s own website and public records of its group structure, and stores what it finds as a labelled Web research: record on the company - for the Web Analysis sheet and for the deep pass below, which can then decide on evidence the workbook never carried. It adds evidence and can reject nothing: a company whose site does not resolve, or whose pages yield no readable profile, is recorded as such rather than filled in. Detail in Website Enrichment.

Step 7: Advanced AI screening (deep FAR)

Shown as Advanced AI Screening (Deep FAR), and gated by advanced_ai_screening. A second, separately configured model (OPENAI_ADVANCED_MODEL) re-examines only the companies the first pass accepted, one structured call per company: granular functions/assets/risks, an explicit acceptance reason, and a confidence score from 0 to 1, against the same profile and the same evidence standard. It is fail-closed - a company it cannot analyse or evidence is pulled out of the accepted set and flagged - and it is the final AI phase of the run; without the plan feature, step 5 is the last.

Step 8: Final analysis

Inline, best-effort predictive-risk analytics run first when the firm holds predictive_risk: audit probability, litigation risk and a defensibility score from the settled comparable set. A failure here is logged and the run continues without risk data - it never fails the study. Then the arm’s length statistics are computed over the accepted set, the comparable rows are written (old rows replaced, fresh evidence events appended), and the results structure is persisted on the study. The study moves to In Review.

The full outcome - counts, conclusion, the tested party’s position in the range - is summarized on the job, which is what the frontend polls through GET /jobs/{id}. The defensibility score lands in the defensibility grade of the finished study.

Progress, retries and idempotency

  • Progress is per-step. The job carries processing_step and a percentage; the frontend renders the stepper from them, not from a bare percentage guess.
  • Transient failures retry. The Celery task retries up to twice on transient errors (network, storage); the AI screening pass retries its provider call once, and web research is best-effort - it can add evidence but never fails the study. File-not-found and validation errors are fatal and never retried.
  • Re-runs are safe. Re-submitting a study replaces the comparable set wholesale (delete old, insert new) and appends a fresh audit event, so a re-run converges to one settled record instead of accumulating ghosts. Study state only ever moves forward along the workflow; a run never advances a study it does not own.
  • Timeouts are sized for large dumps. Soft limit 7200 s (two hours), hard limit 8100 s (two hours fifteen minutes), which is why a very large dump can sit in qualitative screening for a while without being broken.

FAQ

Why is qualitative screening before AI screening? The embedding pass is cheap, local and deterministic; it trims the population so the LLM pass runs on comparables that already look functionally plausible, and every LLM verdict has clean description input to cite.

What if the Web Research step does not run? The run continues - it is gated on the same advanced_ai_screening entitlement as the deep pass, so a tenant without that feature skips both steps. Nothing about the screens changes: they read the description fields in your dump. What you do not get is any web evidence - the report’s External Sources column stays empty, the structure and profile columns carry only what your workbook supplied (plus the vendor independence indicator where the dump has one), and the deep pass decides on the workbook alone. A blank there is “not found”, never a gap filled in by inference.

Can I watch a step run in detail? The stepper shows the current substate and progress. The job’s result summary, the per-comparable reasons and the study’s audit chronology are the detailed record once the run settles.

See it working in your workspace

Sign in to run the steps above on a real study — or book a demo and we will walk the workflow end to end.

Related docs

Book a Demo

Tell us what you'd like benchmarked

We'll confirm a 30-minute screen-share slot within one business day.

We reply within one business day. Your details are used only to arrange the demo — never shared or sold.