Ground-Truth Registry: The AI Reason’s Outcome Record
The ground-truth registry in Quartyl: the cross-study record that links each AI screening reason to its ground-truth outcome — the measure of whether the AI layer earns its keep.
Definition
The ground-truth registry is Quartyl’s cross-study record that links
each AI screening reason to its ground-truth outcome — the registry
(per tenant, across the studies) that tracks the AI reason (the
standardized reason text, the ai_reason_id — the reason the AI gave for
the score, the assessment, the recommendation) and the ground-truth
outcome (the verified result — the outcome that validates or
overturns the reason: the TPO’s acceptance of the exclusion, the
appeal’s confirmation, the re-benchmark’s result). The registry
aggregates the comparables per tenant, per reason (the count, the
outcome, the validation) — the measure of whether the AI layer is
earning its keep (the reasons that are validated by the ground truth,
against the reasons that are overturned). The override
micro-ledger carries the
per-study record (the override, the rationale, the ai_reason_id
linkage); the registry is the cross-study aggregation (the reasons, the
outcomes, the validation — the firm-level view).
The ground-truth registry, in one record:
1. The AI reason (the standardized reason text, the ai_reason_id — the
reason the AI gave for the score / the assessment / the
recommendation — per comparable, per study)
2. The ground-truth outcome (the verified result — the TPO’s acceptance
of the exclusion, the appeal’s confirmation, the re-benchmark’s
result — the outcome that validates / overturns the reason)
3. The aggregation (per tenant, per reason — the count of the
comparables, the outcomes, the validation rate)
→ the measure (the AI layer’s performance — the reasons validated,
against the reasons overturned)
→ the methodology review (the sampling against the outcomes, the
[AI governance](/docs/tp-software/ai-in-transfer-pricing) cadence)
| The element | The content |
|---|---|
| The AI reason | The standardized reason text (the ai_reason_id — the reason the AI gave for the score / the assessment / the recommendation — the reason, not the rationale prose) — per comparable, per study, the link (the ai_reason_id in the override micro-ledger) |
| The ground-truth outcome | The verified result (the outcome that validates / overturns the reason — the TPO’s acceptance of the exclusion, the appeal’s confirmation, the re-benchmark’s result, the human decision’s validation) |
| The linkage | The ai_reason_id (the link — the AI reason in the registry, the override in the micro-ledger, the same reason, the cross-reference) — the connection the registry runs on |
| The aggregation | The per-tenant, per-reason (the count of the comparables, the outcomes, the validation rate — the firm-level view, the cross-study aggregation) |
| The use | The measure (the AI layer’s performance — the reasons validated against the reasons overturned) and the methodology review (the sampling against the outcomes, the AI governance cadence) |
The working read (the AI controls guide and the AI-assisted screening): the ground-truth registry is the Layer 4 (the record control)’s measurement — the audit trail records the events (the captures, the decisions, the overrides); the registry measures the AI layer (the reasons, the outcomes, the validation — the performance). The governance (the AI governance guide’s review cadence — the periodic sampling of the AI-assisted screens against the human outcomes) runs on the registry: the sampling (the reasons, the outcomes) is the registry’s aggregation, the methodology review (the AI layer’s adjustment, the threshold’s tuning) is the registry’s use. The registry is the difference between the AI layer as the black box (the reasons, unmeasured) and the AI layer as the measured tool (the reasons, the outcomes, the validation — the performance known).
Example
The tenant (the firm, the studies): the ground-truth registry
aggregates, per reason — (1) the AI reason “different product mix”
(the ai_reason_id, the standardized text — the reason the AI gave for
the exclusion recommendation, per comparable, per study); (2) the
ground-truth outcome (the verified result — the TPO’s acceptance
of the exclusion (the validation), the appeal’s confirmation (the
validation), the re-benchmark’s result (the validation) — the
outcomes, per comparable); (3) the aggregation (per tenant, per
reason — the count of the comparables (the 48), the outcomes (the 45
validated, the 3 overturned), the validation rate (the 94%)). The
measure: the AI layer’s performance on the reason “different product
mix” (the 94% validation — the reason earns its keep). The methodology
review (the AI governance
guide cadence): the sampling
(the reasons, the outcomes), the adjustment (the threshold’s tuning, the
reason’s refinement) — the registry’s use, the governance loop. The
override micro-ledger
carries the per-study record (the override, the rationale, the
ai_reason_id linkage) — the registry is the cross-study aggregation,
the micro-ledger is the per-study record.
See also
- Evidence Ledger (the per-study record)
- AI-Assisted Comparable Screening
- AI in Transfer Pricing (the governance)
FAQ
What is the “ground truth,” in the registry? The ground truth is
the verified outcome — the result that validates or overturns
the AI reason (the TPO’s acceptance of the exclusion, the appeal’s
confirmation, the re-benchmark’s result, the human decision’s
validation). The ground truth is the external check (the authority’s
outcome, the appeal’s result, the re-benchmark’s number) — against the
AI reason (the internal recommendation). The registry links the two
(the ai_reason_id, the outcome) — the validation (the ground truth
confirms the reason) or the overturn (the ground truth contradicts the
reason) — the measure (the validation rate, the performance).
How does the registry differ from the evidence ledger? The evidence ledger is the per-study record (the captures, the snapshots, the source, the decisions, the overrides — the events, the sequence) — the forensic record, the work as it happened. The ground-truth registry is the cross-study aggregation (the AI reasons, the ground-truth outcomes, the validation rate — the measure, the performance) — the firm-level view, the AI layer’s performance. The two are complementary: the ledger is the record (the events, the sequence), the registry is the measure (the reasons, the outcomes, the validation). The audit packet carries the ledger (the evidence file); the registry is the governance tool (the performance, the methodology review).
Does the registry change the AI’s scoring? No — the registry is the measure (the performance, the validation rate) and the methodology review (the sampling, the adjustment) — it does not change the scoring (the model, the thresholds, the per-study decision) in real time. The registry’s use is the governance loop (the AI governance guide cadence): the sampling (the reasons, the outcomes) → the methodology review (the threshold’s tuning, the reason’s refinement) → the next study (the adjusted methodology, the refined reason) — the loop, not the real-time change. The per-study scoring (the model, the thresholds) is the methodology (the firm’s, the versioned) — the registry measures the methodology’s performance, the governance adjusts it (the cadence, the record).
Run the screens as a study, not a spreadsheet
Quartyl applies the method, PLI and screening steps above as a pipeline — and keeps a documented reason for every exclusion.
Related docs
Evidence Ledger: The Immutable Record of the Study’s Evidence
The evidence ledger in Quartyl: the immutable, append-only record of every evidence event — the captures, the snippets, the snapshots and the source, per study.
Read docAI-Assisted Comparable Screening: How It Works
How AI-assisted comparable screening works in a benchmarking study — what the model evaluates, what it must cite, and the human controls that keep the Accept-Reject matrix defensible.
Read doc