Skip to main content
Quartyl
Studies & Workflowbeginner

Uploading Your Data File: Formats, Column Detection and Mapping

Upload a comparables data file to Quartyl: the accepted Excel formats, the upload and preview flow, how column detection recognizes the financial lines, and the mapping review.

Quartyl Team

One file carries a study’s benchmarking: the comparables population’s financials, one row per company, in a spreadsheet. The wizard takes the file on step one, the pipeline validates it again before screening, and the normalized result is persisted so the report can be rebuilt later. This page is the file’s life — format, upload, detection, mapping, and what happens after confirmation.

Accepted formats and limits

Constraint Value
Formats .xlsx, .xls, .xlsm — Excel workbooks only; CSV is not accepted
Maximum size 50 MB
Stored files per user 25 — delete unused uploads before the next one
Layout One sheet of company rows; the sheet is resolved as preferred name → a recognised dump-sheet name → the first sheet in the workbook
Required lines Company, Revenue, Cost, Operating Profit — the run fails with the missing list named if any is absent

The upload flow: post, store, preview

The wizard takes the file in one trip, and a file record exists from the moment you drop it:

  1. Post the bytes. The browser sends the file to the backend’s proxy upload endpoint. The server checks the extension, the size and your stored-file allowance, registers the file record in pending status with its storage key, and writes the bytes to object storage itself. The record flips to completed and the file metadata — id, name, size, content type — comes back in the response. There is no client-side put to the object store and no separate confirm call: the old presigned request-upload-then-confirm sequence was removed, and every caller now uploads through the server.
  2. Preview the stored file. The wizard immediately asks for a bounded preview by file id — the sheet names, the resolved sheet, the detected column mapping with its candidates and confidence, the detected years, and any required column the file cannot supply. Changing the sheet re-previews by id, so you never re-send up to 50 MB of bytes to look at a different sheet.
  3. Submit carries the id. The wizard stores the file id as the file drops, so submission simply attaches it to the study — and a draft can carry the id without starting a run.

Downloads work the other way: a completed file is fetched through an expiring download URL, never a raw storage path.

Column detection

The file’s headers are rarely the standard names — “Total Revenue (USD)” versus “Revenue”, “EBIT” versus “Operating Profit”. The detector maps each column to the standard financials by scoring, not by guessing:

  • Exact pattern match against the known column patterns adds 0.5.
  • Regex match against the pattern family adds 0.3.
  • Content validation — the column’s values look like the metric (a numeric range check, on financial fields only) — adds 0.2. Scores cap at 1.0.
  • A column is accepted at a confidence threshold of 0.3; below it, the column is left for the mapping review rather than silently mapped.
  • For every field the preview also reports up to three plausible candidate columns, each with its score and the reasons behind it, so you can see what the detector weighed.

The recognized fields — around thirty of them — are Company, Industry, Status, Entity Type, Function, Website, NACE Rev. 2, Primary code (the national industry classification), BvD Independence Indicator, Business Description, Trade Description, Full Overview, Revenue, Sales, Cost, Operating Profit, Operating Expenses, Employee Cost, Number of Employee, RPT, EBITDA, Gross Profit, Intangible Assets, Total Assets, Stock, Debtors, Other current, Cash, Current assets, Non-current assets, and Net Cost Plus. Aliases ride along with each: turnover, net sales and total income land on Revenue; EBIT, PBDIT and profit from operations on Operating Profit; headcount on Number of Employee; goodwill and patents on Intangible Assets. There is no database or source selector anywhere in the product: an export from any provider is read on its headers alone, and it works exactly as far as those headers are recognised.

The more of the balance-sheet and cost lines that are present, the more the run can compute — the Berry Ratio needs gross profit, the intangibles rejection check needs intangible and total assets.

Multi-year files carry the year in the column name — Revenue_2022, Revenue FY22, Revenue (2022) — and the years are detected and recorded alongside the mapping. How those years are consolidated is in Multi-Year Data.

Review the mapping

The wizard surfaces the detected mapping for review. Detection is high-confidence, and the most common first-run surprise is a mis-mapped column — a Cost column read as Revenue, or a year column read as a plain metric. Before you continue:

  • Check the four required lines first: Company, Revenue, Cost, Operating Profit. A wrong one here poisons every number downstream.
  • Re-map any column the detection got wrong, or mark a field ignored so the run stops trying to read it; your choices are saved as a mapping override on the study and merged with auto-detection at screening time.
  • Read the candidate list with it. Each field’s preview entry carries up to three plausible columns with a score and the reasons behind them (an exact name match, a specific pattern, a numeric content check), so a mapping that looks wrong can be checked against what the detector actually saw.

The mapping and the detected years are stored with the study’s parameters block — part of the record, not a transient step — while the study row keeps the dump pointer, its size, and a SHA-256 checksum.

What the pipeline does with the file after submission

  1. Preparing Data — the file is downloaded from storage and re-validated: size, sheet, non-empty, and the required columns through the same detector. Validation errors name the problem: the size over the limit, the missing sheet, the empty file, the missing columns.
  2. Quantitative Analysis — the columns are standardized, the PLI is computed per company (with the multi-year consolidation where years are detected), the multi-year rejection checks run, the financial filters are applied, and the rule-based rejections populate each rejected company’s reason. See Quantitative Screening.
  3. The dump is persisted — after screening, the full normalized dump is written as a parquet file to object storage, and the study stores the pointer, the byte size and the checksum. The raw file itself is never copied into the database; the dump is what a later report rebuild reads. The parquet is kept for the tenant’s retention window — seven days by default, settable between one and thirty-one — measured from the study’s entry into review, then swept by a scheduled cleanup job.

The errors that actually happen

Error Meaning
422 — unsupported type The extension is not .xls, .xlsx, or .xlsm; CSV is rejected
413 / size error The file exceeds 50 MB
429 — file limit You already hold 25 stored uploads; delete unused ones
422 — no data rows The file parsed but the sheet is empty
422 — sheet not found The named sheet does not exist in the workbook
422 — required columns missing Company, Revenue, Cost, or Operating Profit could not be detected — the response names the missing ones
429 — queue full Three analyses are already running for your tenant, or the global ten are full; retry in a few minutes

The creation flow that hosts the upload is in Creating a Study.

See it working in your workspace

Sign in to run the steps above on a real study — or book a demo and we will walk the workflow end to end.

Related docs

Book a Demo

Tell us what you'd like benchmarked

We'll confirm a 30-minute screen-share slot within one business day.

We reply within one business day. Your details are used only to arrange the demo — never shared or sold.