Skip to main content
Quartyl
TP Software & AIprofessional

Transfer Pricing Software: What to Look for (2026 Buyer's Guide)

A capability matrix for buying TP software: database access, workflow, evidence, documentation, governance and security — and the questions to ask before you sign.

Quartyl Team

Transfer pricing software is two products sold as one: a data asset (the comparable company database) and a workflow system (the study, the screens, the documentation, the review). Most buying failures come from evaluating only the data — “how many comparables, how many jurisdictions” — and discovering later that the workflow does not match how the team actually produces a defensible study. This guide is the evaluation framework: the capability matrix, the questions per capability, and the signals that separate a benchmarking database from a transfer pricing system.

The capability matrix

Evaluate against the lifecycle of the study, not against feature lists:

Stage What to evaluate The failure mode it prevents
Scope & setup Tested party profile capture (functions, assets, risks), PLI and parameter setup, multi-year data handling The study starts in a spreadsheet and the system is bolted on at the end
Search & screening Quantitative filters (size, sector, geography), qualitative review workflow, AI-assisted screening with evidence The comparables set is a black box nobody can defend
Adjustments Working capital and other adjustments computed from the data, with the method recorded Adjustments live in side sheets and the range does not reconcile
Range & statistics Arm’s length range (IQR / full range), percentiles, outlier handling, significance The range is a number without a documented derivation
Documentation Local File / Master File generation, benchmarking annex, override rationale captured in the document The report and the working file diverge
Review & governance Role-based review (analyst → manager → partner), state machine (draft → review → sign-off → archived), audit trail of every decision The partner signs off a number no one can walk back to its inputs
Evidence & defense Source capture per comparable, immutable ledger, export for audit “Where is the evidence?” in the TPO proceeding
Data access The database itself: coverage, fields, refresh cadence, API The tool’s data is weaker than the public filings it claims to wrap
Security Residency, encryption, access control, AI data policy Sensitive pricing data leaves the firm’s control

Stage by stage: the questions to ask

Scope & setup

  • Does the system capture the tested party profile as data (functions, assets, risks — the FAR), or as free text? Free text is lost at review time; structured FAR is what makes the tested party selection auditable.
  • Are the benchmarking parameters (PLI, years, method, range convention) set per study and visible on every output? A study whose parameters live in someone’s head is not a study the next reviewer can check.
  • How is multi-year data handled — is the averaging method (simple, weighted, per-period) a recorded choice, or a hidden setting?

Search & screening

  • What quantitative filters exist, in what order, and are the thresholds recorded per study? The quantitative screening sequence is audit evidence — a system that can’t show which filter removed which company has a documentation gap by design.
  • Is the qualitative screen a structured, company-by-company review with accept/reject and a reason — or a checkbox? The qualitative screen is where defensibility is won; the reason field is the difference between a screen and a formality.
  • Where AI is used in screening, what does the system show — the data the AI read, the reason for each score, and a human decision on top? An AI screening feature with no per-company evidence is a confidence machine, not a screening feature (see AI in transfer pricing).

Adjustments, range and statistics

  • Are working capital adjustments computed from the comparable data (ratios, days) with the method and the tested party base recorded — or pasted in from Excel? See the working capital adjustment guide for what “computed” should mean.
  • Does the range output show the IQR bounds, the mid-point, the tested party position, and the outlier treatment — with the IQR convention chosen, not assumed?
  • Is the comparability set size and its statistical adequacy shown (the benchmark reliability question the TPO asks first)?

Documentation

  • Is the report generated from the study data — so the benchmarking annex in the Local File and the working file are the same numbers — or is the report a separate artefact?
  • Is every override (a comparable kept despite a red flag, a rejected company reinstated) captured with the reason, and does that reason appear in the document? The defending the accept-reject matrix exercise starts with that ledger.
  • For India: does the documentation output follow the Rule 10D / 10E structure (Rule 10D, Rule 10E)?

Review & governance

  • Is there a state machine — draft, manager review, partner sign-off, archived — with role-based permissions, so a study cannot be “final” without the sign-off that owns it?
  • Is there a complete chronology: who did what, when, and what changed? When the TPO asks “who approved this comparables set and on what basis,” the answer should be a screen, not a memory.

Evidence & defense

  • Is the source evidence per comparable (the filing, the website extract, the database record) captured at screening time, not reconstructed at audit time?
  • Is the evidence ledger immutable — additions and changes logged, not overwritten?
  • Can the whole evidence file be exported for a proceeding (CSV/Excel for the working set, a certified export for the record)?

Data access

  • What is the database actually — proprietary filings, a licensed feed, or a scrape? What are the coverage (companies, jurisdictions, years), the field depth (segment data, balance sheet detail), and the refresh cadence?
  • Is there API access for the firm’s own tooling, and under what terms?
  • How does the vendor treat your own data — the tested party file, the parameters, the outputs? (The security section covers the questions.)

Database vs tooling: the honest split

The two halves fail differently, and the buyer should price them separately:

Dimension The database The workflow system
What it is Comparable company financials, classification, coverage The study lifecycle on top of that data
How it’s compared Coverage, field depth, refresh, jurisdiction mix Screen fidelity, evidence, documentation, governance
Where teams get burned Assuming coverage depth from a headcount (“50,000 companies” without the field list) Assuming the report matches the work, then discovering it doesn’t
The right question “Show me three comparables from our search, with every field you hold on them” “Walk me through one study from upload to signed report, with the audit trail open”

A strong database in a weak workflow produces studies the team has to re-document by hand; a strong workflow on a weak database produces defensible process around thin data. The TP databases guide has the landscape and the comparison view for the data side specifically.

The signals (and the anti-signals)

Signals of a real system:

  • The demo runs a real study end to end — upload, screens, overrides, range, report — not three pre-built slides.
  • The evidence is shown during the screening, per company, with the source.
  • The review flow has roles and states, and the sign-off is a recorded event.
  • The security answers are specific (residency, encryption, AI data policy, access control) rather than “we’re secure.”

Anti-signals:

  • “The AI does the screening” with no per-company evidence or human decision layer — the AI guide has the control framework.
  • The report is a PDF export of the database search results, with the benchmarking annex assembled separately in Excel.
  • No state machine: anyone with a link can edit the “final” study.
  • The pricing page is a per-seat number with the database as an undifferentiated line item.

FAQ

Do we need both a database and a workflow tool, or is one product enough? One product is fine if it is genuinely both — the test is the capability matrix above. The common failure is buying a database (the comparable company feed) and discovering the workflow is a thin UI, so the real work happens in Excel and the “system of record” is a spreadsheet.

How important is AI-assisted screening in a 2026 evaluation? Important as acceleration with evidence, not as a replacement for the screen. The questions are: what does the AI read, what does it show per company, and where is the human decision recorded? A screening feature that cannot answer those three is not a control, it is a risk the firm has adopted (see AI in transfer pricing and manual vs AI-assisted).

What is the realistic evaluation effort? Two weeks: week one — the capability matrix against the vendor’s demo, one real study run end to end with the audit trail open; week two — the data deep-dive (three comparables, every field), the security review, and the reference calls. Firms that evaluate on the feature list alone tend to discover the gaps in their first audit.

How should total cost be compared? Cost per defensible study, not cost per seat: the seat price plus the internal hours the workflow saves (or adds) per study, plus the data access terms. A cheaper seat that doubles the documentation hours is the more expensive option.

Run the screens as a study, not a spreadsheet

Quartyl applies the method, PLI and screening steps above as a pipeline — and keeps a documented reason for every exclusion.

Related docs

Book a Demo

Tell us what you'd like benchmarked

We'll confirm a 30-minute screen-share slot within one business day.

We reply within one business day. Your details are used only to arrange the demo — never shared or sold.