Research · Published:
AP exception reviewer agreement research
A blinded review design for testing whether two reviewers identify the same evidence gaps, owners, and stop conditions.
Methodology
This study asks whether independent reviewers reach the same factual classification and routing decision from the same frozen evidence. It examines representative AP exceptions across invoice, receipt, credit, vendor, tax, and payment workflows. The research is designed for finance leaders who need measurable process evidence without transferring approval, accounting judgment, vendor acceptance, access administration, or payment authority to an outsourced preparation role. The output is a reproducible register and owner-facing decision packet, not a declaration that a transaction is valid. GAO's Green Book provides the framework for responsibility, quality information, control activities, and monitoring. NIST SP 800-53 and the least-privilege definition support attributable events and restricted permissions. IRS recordkeeping guidance reinforces retaining records that support business activity. Company policy and accountable owners still determine the actual treatment.
Evidence and scope
The population register retains every in-scope item, including completed, held, rejected, corrected, superseded, and reopened cases. Core fields include entity, stable case ID, supplier or employee identity, transaction reference, amount and currency where relevant, intake time, source versions, queue state, assigned owner, reviewer, decision event, implementation event, and final disposition. The six primary observations are cause classification, missing source, risk flag, named owner, stop condition, requested decision. Excluding difficult or unresolved cases would bias the study toward clean outcomes, so selection rules are frozen before measurement and every exclusion receives a reason. Sensitive bank, tax, identity, and personnel values are masked in analytical files and remain in approved systems.
Key Stats
Time is treated as evidence. Each event retains its original display value, timezone, actor or process, system, and stable event identifier. A normalized UTC time is added for comparison without replacing the original. The protocol separates intake, packet-ready, routing, owner acknowledgement, evidence request, response, decision, implementation, and verification. That distinction prevents waiting for missing source evidence from being attributed to an owner and prevents post-decision implementation time from being attributed to the preparer. If system clocks or exports conflict, the variance is recorded as a limitation and routed to the system owner rather than silently resolved.
Research-to-practice
Source preservation is tested before analysis. Original documents, portal exports, workflow histories, approval records, master-data versions, and relevant correspondence receive stable identifiers or hashes where the approved platform supports them. Extracted values remain linked to the source and version from which they came. A current screen cannot prove what a reviewer saw earlier, and an edited attachment cannot retroactively replace the file used for a prior decision. The reconstruction packet therefore keeps both the historical and later record, explains the event connecting them, and marks evidence that was unavailable at the tested point in time.
Implementation
The analytical method uses explicit classifications. For each signal, reviewers choose aligned, different, missing, superseded, not applicable, or owner interpretation required and cite the supporting record. Calculations show inputs, precision, rounding, and result. Identity comparisons retain legal entity, internal account, source-document name, and any approved relationship record. The protocol does not treat matching totals, familiar display names, a successful status, or a prior transaction as proof. Those observations can guide review, but an authorized owner must resolve substantive differences involving scope, validity, accounting, compliance, access, or release.
Key Takeaways
Quality is tested with a blinded second review. Two reviewers receive the same frozen packet and independently record cause classification, missing source, risk flag, named owner, stop condition, requested decision. Agreement is measured separately for factual fields, classification, missing evidence, named owner, and stop condition. A disagreement is adjudicated against the source and taxonomy; it is not settled by choosing the faster answer. The taxonomy and instructions are revised when wording, missing source structure, or system ambiguity causes recurring disagreement. The original reviewer responses remain in the research record so improvement does not erase evidence of the earlier measurement weakness.
Findings
Challenge cases are deliberately included: records spanning two entities; documents replaced under the same visible name; timestamps near a cutoff; an approval followed by a value change; an automated state without an attributable actor; overlapping owners; transferred queues; a supplier message introducing sensitive changes; a rejected action later retried; and closure followed by new evidence. The method passes when reviewers reconstruct the same sources, chronology, factual differences, and accountable decision point. It fails when current state substitutes for historical evidence, when missing records are treated as negative proof, or when the support role resolves a controlled decision to improve a metric.
Findings
Role boundaries are part of the study design. AP support may inventory sources, normalize timestamps, prepare comparisons, maintain an evidence-based queue, calculate documented measures, and route bounded questions. It may not approve invoices or vendors, choose accounting or tax treatment, verify sensitive changes through an unapproved channel, grant or remove privileged access unless assigned as the authorized administrator, override controls, approve payment, or release funds. Finance, purchasing, receiving, tax, compliance, treasury, security, human resources, and system owners retain their designated decisions. The register records those assignments so a missing owner is visible as a design defect rather than quietly absorbed by the analyst.
Findings
Useful measures include population completeness, source completeness, packet-ready rate, owner acknowledgement, decision time, implementation time, missing-evidence cycles, classification agreement, reopened items, repeated causes, and verified closure. Report medians and distributions rather than only averages, and stratify by cause, entity, owner role, risk, and workflow where sample sizes permit. Do not rank individual employees by raw speed or reward hold removal. A controlled stop can be the correct result. Counts are accompanied by definitions, numerator, denominator, exclusions, observation window, and known system gaps so management can distinguish a process signal from an attractive but unsupported benchmark.
Findings
Limitations include overwritten histories, inaccessible external systems, shared identities, clock drift, policy ambiguity, informal decisions, unrecorded supplier events, and small samples. The study cannot prove fraud, legal ownership, tax status, contract validity, or ultimate receipt of funds. It also cannot establish causation from a workflow correlation alone. Its defensible conclusion is narrower: whether independent reviewers reach the same factual classification and routing decision from the same frozen evidence can be evaluated when the population, source versions, chronology, roles, classifications, decision, implementation, and residual limitations are preserved. Management may use recurring evidence gaps to prioritize system or workflow changes, but the research packet does not authorize those changes.
Findings
The final handoff contains a methods note, frozen population, data dictionary, source register, exception taxonomy, event chronology, reviewer-agreement table, findings with denominators, challenge-case results, limitations, and an action register. Each proposed action names the accountable owner, required evidence, target review date, and verification method. A compact case summary shows cause classification, missing source, risk flag, named owner, stop condition, requested decision and links them to original records. Later corrections append a new version rather than rewriting the measured history. This structure lets another reviewer reproduce the result, lets leaders see where delay or recurrence actually arises, and keeps the outsourced support role focused on evidence preparation instead of policy interpretation or transaction approval.
Turn evidence into an owner-ready AP handoff
Our support team can preserve sources, prepare comparisons, and maintain the review queue while your authorized owners retain policy, approval, access, and payment decisions.
Review AP support servicesSources
These primary sources support the control principles and evidence boundaries in this report.
FAQs
Are the planning numbers benchmarks?
No. They describe a testable workflow shape and are not promises, market averages, or production targets.
What should an outsourced AP assistant own?
Repeatable preparation, documentation, status tracking, and follow-up within least-privilege access. Named finance owners retain approval and payment decisions.
When should an item be escalated?
When evidence is missing, a request changes payment details, a duplicate or fraud signal appears, or the item falls outside the written rule.