Transfer Pricing Benchmarking: Methodology, Data & Worked Study (2026)
The end-to-end benchmarking study: scoping, search design, quantitative and qualitative screening, adjustments, the arm's length range and refresh — with a worked Indian case.
A benchmarking study is an argument, not a calculation. Two analysts with the same tested party, the same method and the same database can produce materially different ranges — and the difference is almost always in the decisions, not the arithmetic. The purpose of a study is to produce two things at once: the range (the conclusion) and the record (every decision that produced it, in an order a reviewer can re-run). This guide is the end-to-end methodology, with a worked Indian case at the end.
The six-stage anatomy of a study
| Stage | Core question | Output | Typical failure mode |
|---|---|---|---|
| 1. Scoping | Who are we benchmarking, and for which transaction? | Tested party, method, PLI, period | Inherited last year’s choices; wrong tested party |
| 2. Search design | Which population could contain comparables? | Search criteria: codes, size band, geography | Over-inclusive or over-exclusive search |
| 3. Quantitative screening | Which candidates survive the data filters? | Screened pool + a log of every exclusion | Unexplained deletions; thresholds that flatter the range |
| 4. Qualitative screening | Are the survivors actually comparable? | Accept-Reject matrix with per-company reasons | Boilerplate notes written after the range was known |
| 5. Adjustments | How far can the comparison go? | Adjusted PLI values for each kept comparable | Working capital adjusted without measuring the difference |
| 6. Range & conclusion | Is the price arm’s length? | IQR or full range + the tested party’s position | The range choice undocumented |
Each stage has exactly one defensible output. If you cannot point at the document that stage produced, the stage has not happened.
Stage 1 — Scoping
Scoping fixes four things, in this order:
- The transaction(s) — what is actually controlled, split into components (services vs goods vs finance) because one intercompany relationship often carries several transactions, each with its own method.
- The tested party — the entity whose result gets compared against the pool. This is the “least complex” decision, and it is the one a Transfer Pricing Officer re-litigates first. See tested party selection.
- The method — chosen by the best method rule from the decision framework, not by habit.
- The PLI — the indicator that isolates the tested party’s routine contribution, with a denominator you can actually measure for the pool.
All four rest on the functional analysis of the parties involved. If the FAR is wrong, every downstream choice is wrong in a way the data will not show you.
Stage 2 — Search design
The search defines the population from which candidates are drawn. Three knobs, each of which silently shapes the range:
- Industry classification — NIC-2008 (India) or NACE (elsewhere) codes, or a defined search string over product lines. A code that is too broad (“business services”) drags in comparables you will spend the whole qualitative screen rejecting; a code that is too narrow produces a pool of four and a range with no statistical meaning.
- Size band — a revenue or asset band around the tested party, commonly expressed as a ratio (e.g. 0.25x–4x the tested party’s revenue). Size matters because scale drives cost structure.
- Geography — the countries the search covers. OECD practice allows multi-regional searches; Indian practice is most defensible when the pool is geographically coherent or the tax/cost differential is addressed.
The rule for all three: write the criteria down before you run the search and keep them verbatim in the file. A search you cannot re-run is a search a reviewer will not accept.
Search all 1,297 NIC-2008 sub-classes
Free tool covering the full official NIC-2008 classification. Nothing you type leaves your browser.
Stage 3 — Quantitative screening
Quantitative screening applies the data-driven filters — size, profitability, sector, geography, PLI denominator availability — in a fixed order. The full filter discipline, including the log you must keep, is in quantitative screening. The one non-negotiable: every company the filters drop gets a written reason at that moment. “Removed for size” is a reason; “removed” is not.
Stage 4 — Qualitative screening
Quantitative survival is necessary, not sufficient. The qualitative review checks what the numbers cannot: product lines, revenue mix, customer base, assets, and whether the company is actually doing the tested party’s job. This is where the Accept-Reject matrix is built — and it is the part of the study an auditor reads most closely, because it is the part that is either genuine or fabricated. See qualitative screening.
Stage 5 — Adjustments
Comparables are never identical, and the adjustments bridge the gap that matters for the PLI:
- Working capital first — differences in current-asset and current-liability ratios change the cost base or revenue the PLI sits on. Measure the difference; do not assume it is immaterial. See working capital adjustment.
- Everything else — country premium, scale, accounting differences, extraordinary events. Adjust where you can measure; exclude where you cannot, and record why. See comparability adjustments.
An adjustment you cannot compute is an exclusion with a better story.
Stage 6 — The range and the conclusion
The final step is the arm’s length range — most defensibly the interquartile range, with the full range as the Indian practice fallback — and the tested party’s position inside it. The mechanics and the decision rules are in IQR vs full range. “Inside the range” is the conclusion; “inside which range, computed how, from which pool” is the defence. Write both.
Worked example: a routine Indian services provider
Setup. Company S is an Indian IT-enabled services company providing back-office and process services to a non-resident parent. Routine profile: no owned intangibles, no material price risk, a service-level agreement with a fixed scope. One controlled transaction: the services fee. Method: TNMM. Tested party: Company S (the least complex party). PLI: operating profit on operating costs (OP/OC).
Step 1 — three-year averages (FY 2023-24 to 2025-26).
| Year | Revenue (₹ cr) | Operating costs (₹ cr) | Operating profit (₹ cr) | OP/OC |
|---|---|---|---|---|
| 2023-24 | 182.4 | 175.1 | 7.18 | 4.10% |
| 2024-25 | 197.8 | 189.6 | 8.34 | 4.40% |
| 2025-26 | 204.5 | 197.3 | 7.89 | 4.00% |
| Average | 194.9 | 187.3 | 7.80 | 4.17% |
Step 2 — search. India; NIC codes covering information services, offshore process and business services; revenue band ₹50–2,000 cr (≈0.25x–10x is too wide — the final band used was 0.5x–5x of Company S); three financial years per company.
Step 3 — quantitative screen (12 candidates found).
| Company | OP/OC (3-yr avg) | Disposition | Reason |
|---|---|---|---|
| A | 3.2% | Keep | — |
| B | 3.6% | Keep | — |
| C | 4.0% | Keep | — |
| D | 4.5% | Keep | — |
| E | 4.8% | Keep | — |
| F | 5.1% | Keep | — |
| G | 5.4% | Keep | — |
| H | 5.9% | Keep | — |
| I | 6.3% | Keep | — |
| J | 2.1% | Reject | Revenue ₹38 cr — below the 0.5x size band |
| K | (1.4)% | Reject | Loss-making in all three years; no documented turnaround |
| L | 4.7% | — | Survives quantitative screen; fails qualitative (below) |
Step 4 — qualitative screen. L is a diversified conglomerate: 62% of its revenue is outside the service line and its service segment carries the group head office. Rejected: “service line is a minor segment; entity-level PLI reflects group-level activities, not a standalone service provider.” The other nine pass — product lines verified against filings and websites, revenue mix ≥80% in-scope for all of them.
Step 5 — working capital. Company S’s current asset/operating cost ratio is 0.14; the pool median is 0.12. The 0.02 gap, at the pool’s median operating cost margin on working capital, supports a +0.2pp adjustment to Company S’s OP/OC. Adjusted: 4.37%. No other material differences; no further adjustments.
Step 6 — the range. Nine adjusted values, sorted: 3.2, 3.6, 4.0, 4.5, 4.8, 5.1, 5.4, 5.9, 6.3. Median 4.8. Q1 = median of the lower half (3.2, 3.6, 4.0, 4.5) = 3.8. Q3 = median of the upper half (5.1, 5.4, 5.9, 6.3) = 5.65. IQR: 3.80% to 5.65%, mid-point 4.73%.
Conclusion. Company S at 4.37% (adjusted) sits inside the IQR. No transfer pricing adjustment is supported by this evidence. The file contains the search criteria, the twelve-row disposition log, the FAR comparison for L, the working capital workings, and the range computation — each stage’s output, in order.
Refresh and roll-forward
A study is a point-in-time artefact. The discipline:
- Re-run annually. The pool drifts (companies enter and leave, segments change) and the tested party changes too. Rolling last year’s study forward with only the new year’s tested-party numbers is the practice that gets caught, because the TPO is re-running the pool for the current year anyway.
- Watch the trend. A tested party that lands inside the range for three years but drifts toward one edge each year is a pre-emptive re-benchmark, not a problem — if you see it coming.
- Know what invalidates a study. A reorganisation, a new intangible, a changed business model, a new intercompany product — any of these changes the FAR and voids the tested party selection. The study must be re-scoped, not re-run.
- Keep it contemporaneous. The documentation must exist by the due date of the return, not be reconstructed after a notice. See the compliance calendar.
FAQ
How many comparables is enough? There is no statutory number. Practice settles around five to fifteen companies after screening, enough for the quartiles to mean something. A pool of three is a range you are showing, not one you are defending; a pool of forty with thin notes is worse than a pool of ten with strong ones. The count is a symptom of search design, not a target.
Can I roll forward last year’s study? Only if nothing changed — same tested party, same FAR, same transaction, same pool. In practice the pool almost always changes, and the TPO re-runs it, so the defensible answer is a full re-run with the prior year’s results shown as the trend line.
Is being outside the range an automatic adjustment? No. It is a signal that the comparability analysis needs a second look: differences that should have been adjusted, a pool that was too narrow, or a tested party that was never routine to begin with. “Outside the range” starts a review; it is not a finding.
Which database should I use? Commercial benchmarking databases are the default in Indian practice; public filings plus GLEIF/LEI data is a viable supplement that strengthens (never replaces) the search. Whatever the source, the criteria, the snapshot date and the exclusion reasons belong in your file — the database is the library, the study is the argument.
Run the screens as a study, not a spreadsheet
Quartyl applies the method, PLI and screening steps above as a pipeline — and keeps a documented reason for every exclusion.
Related docs
Tested Party Selection: The "Less Complex" Principle, Applied
How to pick the tested party in a TNMM study: the OECD "least complex" logic, the six-question test, the documentation trail, and the objections transfer pricing officers raise most.
Read docIQR vs Full Range: Choosing the Arm's Length Range
OECD's interquartile range versus the full range: quartile math step-by-step, when each is defensible, the mid-point argument, and how to document the choice in India.
Read doc