Raw Dump Data: Inspection and Reuse After Screening
Raw dump data in Quartyl: the parquet persisted after quantitative screening, inspecting the numbers behind every filter decision, reuse across re-runs, and the per-tenant retention window.
The dump is the study’s raw data, preserved whole. At the start of the quantitative step of the pipeline, before any screening changes anything, the full uploaded file - normalized to standardized column names - is written as a parquet file to object storage. It is the numbers behind every filter decision, the rebuild source for the report, and a bounded, retained artifact rather than an in-memory byproduct.
What the dump is
- The content - the complete comparable population as uploaded: every company row, all detected columns (the standardized financials, the year-suffixed multi-year columns, the description fields, the industry and website fields), nothing filtered out. Rejections are decisions made on top of this data; the data itself keeps everything.
- The format and location - a normalized parquet at
dumps/{tenant}/{study_id}/dump.parquetin the storage provider, with the storage key, the file size and a checksum recorded on the study. The raw rows are never stored as database JSON - the database keeps the key and the metadata, storage keeps the data. - The metadata - the row count, the column mapping (which original column mapped to which standard field, the output of the detector) and the detected fiscal years. That metadata is what makes the dump inspectable: you can see what the pipeline understood the file to be.
Plan feature: This capability requires the
dump_retention_configplan feature (see Plan Features).
Inspecting the dump
The dump exists so a reviewer can answer “what number was that decision made on?” For a rejected comparable - “Below minimum revenue,” “High RPT,” “Incomplete Multi-Year Financials: missing data for year(s) 2023” - the dump holds the exact revenue, the exact RPT, the exact year columns, so the reason can be checked against the source instead of trusted on faith. The same is true in the other direction: an accepted company’s PLI, its per-year figures and its description fields are all in the dump, which is why the screen’s reasons are auditable rather than opaque.
The report rebuild reads the same artifact: when the study is approved and the Excel workbook is generated, the builder downloads the dump and rebuilds the full search-result and web-analysis sheets from it - on any worker host, because the key is a storage reference, not a local file path.
Reuse across re-runs
- A re-screen is a fresh run on the same source file. Re-submitting a study (re-screen) re-reads the uploaded file and produces a new dump for that run, so each run’s numbers trace to that run’s artifact.
- The dump outlives the run’s in-memory work. Pipeline steps mutate an in-memory results structure; the dump is the persistent source that survives step failures, worker crashes and retries, and that report generation later depends on.
- Deleting a study deletes its dump with it. Orphan dumps - storage objects whose study row no longer exists or is soft-deleted - are reclaimed by the retention sweep, so storage does not accumulate data for studies that are gone.
Retention
The dump is retained on a per-tenant window, and the window is configurable:
| Parameter | Value |
|---|---|
| Default window | 7 days |
| Configurable range | 1 to 31 days, per tenant |
| Who sets it | The platform superadmin, on the Add Firms console; assigning a plan carries the plan’s retention value to the tenant |
| When the clock starts | When the study enters In Review - a study still awaiting review never loses its dump to the calendar |
| Enforcement | A daily scheduled sweep lists the dump prefix, resolves each key’s tenant, and deletes objects older than that tenant’s window |
Two properties matter in practice:
- A deleted dump only blocks new rebuilds. The sweep removes the parquet; it never touches the study’s persisted analysis data or any report already generated. What changes is that a new Excel report can no longer be rebuilt from the full dump.
- The per-tenant window is why a bucket lifecycle rule is not enough. Storage lifecycle rules are bucket-level and cannot read a per-tenant setting, which is why the sweep - not the bucket - is the enforcement mechanism.
FAQ
Can I download the dump? The dump is a system artifact referenced by storage key, not a user-facing download. Its contents surface where they matter: the per-company reasons in the grid, the details panels, and the report sheets rebuilt from it.
Why persist the dump before screening instead of after? So that whatever the screen decides - and whatever fails later in the pipeline - the complete pre-decision data is already safe. A dump saved after screening would only prove the decisions the screen made.
What if the retention window expires mid-review? It does not, by construction: the clock starts at In Review, and the default 7 days (or the tenant’s configured window) runs from there. A study in review keeps its dump until the window genuinely elapses, and the sweep is what enforces the boundary.
See it working in your workspace
Sign in to run the steps above on a real study — or book a demo and we will walk the workflow end to end.
Related docs
Quantitative Screening in Quartyl: Filters and Results
Quantitative screening in Quartyl: the exact filter order from multi-year completeness to RPT and revenue thresholds, Golden Rule aggregation, and the recorded rejection reason per company.
Read docThe Screening Pipeline: Quantitative to Qualitative (8 Steps)
The eight-step async analysis pipeline in Quartyl: quantitative, qualitative and AI screening, gated web research and deep screening, final analysis, per-step progress.
Read doc