Agentic AI in accounts payable: six adoption blockers
A guide for AP and GBS leaders evaluating agentic AI in accounts payable
Having been in the AP automation game for over 20 years, we’ve seen a lot of our contemporaries publishing the same number without a lot of context, if any: 80% touchless processing, often higher. It appears in sales decks, on websites, and even in analyst briefings, and it has become the benchmark the industry uses to signal that automation is working.
The SSON State of Accounts Payable Market Report 2026 pits a number against the reality: 57% of organizations report touchless processing rates below 25%, and a further 18% sit between 25% and 50%. That means roughly three quarters of AP teams below 50% touchless, against an advertised benchmark of 80%, more than a decade into serious investment in automation technology.
SSON State of Accounts Payable Market Report 2026
Touchless invoice processing rates
57%
of organizations report touchless rates below 25%
18%
sit between 25% and 50%
7%
exceed 75%
76%
of practitioners rate agentic AI as the most transformative technology for AP in 2026, yet production adoption remains early-stage across the market.
This is not a technology gap because we know the platforms exist – we are one, after all – with document parsing extracting accurately, and three-way matching running automatically at global scale. The gap instead, between what vendors promise and what organizations actually achieve, is an architecture gap, and it has specific, identifiable causes that we've observed demo after demo failing to address.
What is agentic AI in accounts payable?
Agentic AI in accounts payable refers to AI systems that act on invoices rather than only reading them. An agent can resolve an exception, apply GL coding, route an approval, or post to the ERP without a human completing each step, working from vendor master data, open purchase order lines, and goods receipt records.
This is a different category from OCR and rules-based workflow automation. Those tools extract data and move an invoice along a predefined path. An agent decides what to do when the path is not predefined.
The metric agentic AI is employed to improve is the touchless processing rate: the share of invoices that complete the cycle from receipt to posting with no human intervention at any step.
The legacy of the decisions you make now, regarding the ambitious touchless invoice processing rates you’ve committed to deliver, remains to be seen. There will be a lot of too-good-to-be true claims or well-intentioned promises put in front of you about the role of agentic AI in your AP function.
Key takeaways:
- Roughly three quarters of AP teams process below 50% touchless, against an advertised benchmark of 80%
- The performance gap is mainly architectural, not a limitation of the technology
- Six blockers account for most of it: data readiness, ERP write access, audit evidence, fraud surface, team adoption, and ROI modeling
- ERP integration ranks as the single biggest barrier at 38%, ahead of data quality at 32%
- Deterministic rules can handle high-volume structured decisions at fixed cost while the unit economics of agent-heavy workflows is unstable at best
- The requirement to evidence an automated financial decision comes from tax and audit law, and applies now, regardless of EU AI Act timelines
1. Your data is not ready for agents, and agents will tell you that the hard way
We’ve heard it so many times: “Garbage in. Garbage out.”
Put in the context for your organization: agents make decisions. Bad data makes bad decisions at scale and at speed, and in AP the consequences are not theoretical. An agent resolving a price mismatch exception needs accurate vendor master data, current PO data, GL coding history, and contract terms – all accessible, all consistent, and drawn from the same source. Most enterprise AP environments, particularly those running multiple ERP instances across entities and regions, do not have that foundation in place.
SSON State of Accounts Payable Market Report 2026
16%
of finance leaders fully trust their enterprise data. For most organizations deploying AI agents, that's where they start.
You can see how the problem compounds in multi-entity operations. Vendor master records maintained across regional ERP instances typically contain duplicates, inconsistent naming conventions, and missing tax codes that no one has had the budget or mandate to clean. An agent working from that foundation does not produce better exception management than a human would. In fact, the implications in this scenario of AI vs human are significantly more dangerous because the agent produces errors with confidence. Mistakes become harder to catch precisely because they arrive without the uncertainty signals a human reviewer would naturally flag.
EY practitioners described the pattern in a June 2026 webinar. A global retail and fashion company with 800,000 invoices across 78 countries was already automating at a high rate a year into leveraging an agentic AI AP strategy.
The move to 90% touchless came from master data cleanup, not from more agents or more sophisticated orchestration of the existing agents. A investigation of the system revealed duplicate vendor records and missing fields in a reporting dashboard. The client then corrected the underlying data and therefore the automation improved. Ultimately, the agent architecture remained completely unchanged. The technology did not get smarter; the conditions it operated in simply got cleaner.
How we approach the "agents meet bad data" scenario at Springtime
Before deploying agents, we recommend auditing the specific data fields they will act on. Vendor master completeness, PO coverage rates, and GL coding consistency are the three that matter most. Rather than treating data quality as a prerequisite that must be fully resolved before the project begins, define a measurable readiness threshold for each field, this means a specific, testable condition that must be met before autonomous resolution goes live on that exception type.
Many of our customers find success by deploying agents in a narrow pilot, with human review still in place, and together, Springtime and customer alike will use the error outputs the agents surface to prioritize which data problems to fix first.
With this approach, the agents we deploy become a diagnostic tool as much as an automation tool in this phase, and the data cleanup that follows is grounded in actual failure patterns rather than assumptions about where the gaps are.
2. ERP integration is the most underestimated blocker
An agent that can only read invoice data and flag exceptions is not really an agent in any operationally meaningful sense. The value of agentic AP comes from agents that can resolve exceptions, update GL coding, trigger approval routing, and post to the ERP without human intervention at each step. That requires write access to core financial systems, and that is where enterprise IT architecture, security governance, and SAP clean-core frameworks create friction that we’ve seen standard vendor demos consistently underrepresent.
SSON State of Accounts Payable Market Report 2026
Biggest barriers to AI implementation
This critical distinction between read and write access matters. Agents reading ERP data for context – think vendor master records, open PO lines, goods receipts – are the type that carry minimal operational risk because nothing is committed to the financial system until a human acts on what the agent has surfaced and recommended.
But agents writing to ERP systems – completing tasks like posting invoices, updating GL codes, or modifying vendor bank details – is where the risk profile fundamentally changes. Governance added after deployment cannot be trusted to hold. In an agentic environment where agents can execute write actions across financial systems, control must be an architectural constraint from the start, not a layer on top applied once the system is already running.
The challenge intensifies for organizations mid-migration on S/4HANA, or running multiple ERPs across entities. Deploying agents into a transitional architecture creates dependencies that are expensive to maintain, difficult to version-control, and compound in complexity with every configuration change to the underlying ERP.
How we approach the ERP integration challenge at Springtime
We treat ERP integration as a precondition, not a parallel workstream. It’s important to define the specific write operations each agent requires before scoping its decision logic and work backwards from what the ERP governance framework will permit rather than forwards from what the agent could theoretically do. For organizations mid-migration, we prefer to consider deploying agents as a layer above the existing ERP first. This approach reduces dependency on migration timelines and allows automation to deliver measurable value before the ERP landscape stabilizes.
3. Governance and auditability: what the auditors will ask that the pilot or demo didn’t cover
When an agent makes an autonomous decision that turns out to be wrong, an auto-approved duplicate invoice or a misrouted exception because a vendor master record was stale, someone must explain what happened to the CFO, to internal audit, and potentially even to a regulator. Therefore, the requirement is not that the agent made the right decision most of the time but that every decision is traceable, explicable, and reviewable on demand, with enough detail to satisfy external scrutiny a year after the fact.
SSON State of Accounts Payable Market Report 2026
Top concerns about giving AI agents decision-making authority
Most organizations’ current AP systems log exceptions that humans resolve. Agentic systems need to log every decision the agent made, the data it acted on, the model output that informed the decision, and the outcome, at transaction level rather than in aggregate.
This isn’t a new standard because Article 233 of the EU VAT Directive already requires a reliable audit trail between every invoice and the supply it relates to, held from the point of issue until the end of the mandatory storage period. The requirements are also stricter in Germany where the GoBD adds traceability and immutability standards for electronic bookkeeping. In cases where a human matched an invoice to a goods receipt, this trail is the approval record. In instances where an agent makes that same determination autonomously, the agent’s reasoning becomes the evidence and it has to be captured to the same standard.
AP invoice automation currently falls outside the Annex III high-risk categories of the EU AI Act, with obligations deferred to December 2nd, 2027 in any case. Only the AI literacy duty under Article 4 and the transparency duty under Article 50 apply to AP today. The requirement to evidence an automated financial decision comes from tax and audit law, and it is in force now.
How we approach the governance challenge at Springtime
Define audit requirements before scoping agent capabilities, and work with compliance and internal audit teams to establish what a defensible audit trail looks like for autonomous AP decisions at the level of granularity your external auditors and regulatory framework require. If a platform cannot produce that trail by design, it is not ready for production in a regulated finance environment. Invoicetrack is built to produce it and is governance first by design.
Invoicetrack’s Accountable AI methodology is designed around the concept of “retrieval” not “reconstruction”. In practice, this means every AI decision traces to a named rule, specific evidence, an accountable party and a confidence score captured at the moment the decision was made. The immutable record covers user actions, AI decisions, approval decisions, and configuration changes, each with a timestamp and named actor. A compliance team interrogating a specific invoice 14 months later can query the log directly, without having to perform a time-consuming forensic exercise.
There’s also an additional two controls that will matter to auditors that are present in Invoicetrack’s Accountable AI architecture: autonomy thresholds and gated model retrains. Autonomy thresholds are configured per process step and per company code by your process owners. Below the threshold, human approval is mandatory. Model retrains are gated: queued, reviewed in a test environment, and require explicit approval before reaching production, with the model version active at the time of any decision forming part of that decision’s permanent record.
4. The fraud surface expands when agents get write access
Agentic systems with write access to payment workflows and vendor master data are high-value targets for fraud tactics that are growing in sophistication faster than most AP controls are evolving. The risk is not hypothetical, and data from a recent SSON market report reflects this point plainly.
SSON State of Accounts Payable Market Report 2026
2/3+
of organizations report attempted supplier fraud in the last two years.
15%
suffered financial loss
53%
detected and prevented an attempt
3%
are confident in their current controls
AI-enabled social engineering now produces vendor change requests that closely mirror existing approval styles and communication patterns. Tactics including deepfake voice and video calls are being used to pressure AP teams into bypassing standard payment controls. Synthetic vendor identities and manipulated invoices are designed to exploit high-volume processing environments where individual transactions receive less scrutiny.
Let’s put the impact in numbers, where the scale is clearest in the US data: the 2026 AFP Payments Fraud and Control Survey (based on data from 465 US corporate finance practitioners) stated that 76% of organizations reported attempted or actual payments fraud in 2025. 74% were affected by business email compromise specifically, representing a significant increase from 2023 and 2024. Financial losses were reported by 48% of organizations with revenue under one billion dollars, and 66% of organizations exceeding one billion dollars.
An agentic system operating autonomously within these patterns is not just a target. It can be the mechanism through which fraud succeeds, if an attacker understands the agent’s decision boundaries and engineers a request that falls within them. Put a different way, a flaw in a manual process may result in a typo that a human catches on review but a logic flaw in an autonomous payment agent triggers an erroneous financial transaction before any human sees it.
How we approach the question of fraud at Springtime
Invoicetrack operates on a “SAP is master” principle with things like vendor bank details or payment terms originating the ERP and flowing into the AP automation platform as read-only reference data. That would mean that a person will ill intentions who tries to manipulate what Invoicetrack sees could still not overwrite the authoritative record.
Also worth mentioning are the confidence thresholds that govern autonomous decisions. These are defined by the customer’s team, not Springtime. Therefore, the boundary a fraudster would need to exploit is explicit, auditable and controlled by the organization.
Finally, every AI-influenced decision is logged against our rule and evidence trail system. Sensitive operations like changes to a vendor bank account would trigger specific anomaly-detection routing. The audit record would also distinguish between what the system decided autonomously and what a human approved, which is exactly what the investigators need when a transaction looks wrong after the fact.
5. Change management: the people problem that keeps being underestimated
AP teams who have spent years managing around broken automation are not naturally trusting of a system that promises to resolve exceptions autonomously, and their skepticism is rational rather than oppositional. They have seen legacy OCR tools that misread invoices and created more manual correction work than the paper process they replaced. They have seen workflow automation that only shifted bottlenecks rather than removing them, and vendor promises that did not survive contact with their actual invoice mix, with non-PO volumes higher than expected and exception rates that the sales presentation positioned as merely edge cases.
RPA followed the same pattern at greater cost. It was faster than manual processing but brittle, expensive to maintain, and unable to handle anything outside the narrow path it was built for. Agentic AI is more capable, but the sales cycle looks familiar enough that skepticism is a reasonable starting position.
SSON State of Accounts Payable Market Report 2026
1/3
of AP organizations aren't yet engaged with AI adoption internally.
17%
describe their team as skeptical, with real resistance to work through
15%
haven't had the AI conversation with their team yet
That's in a year when 76% of practitioners call agentic AI the most transformative technology in AP.
The gap between strategic recognition and operational enthusiasm is real, and it matters because agentic AI implementations that do not leverage the process intelligence and practical insights of the AP team tend to produce one of two outcomes: the technology gets abandoned after the pilot, or it gets used in ways that route around the governance it was designed to enforce.
The same SSON report offers some encouragement: 68% of practitioners are either excited and ready or cautiously optimistic about AI in AP. So we know that the appetite exists, but there’s an understandable lack of confidence that the technology will deliver the freed capacity it promises rather than creating new supervisory workload on top of the existing manual burden. But this confidence can only be earned through demonstrated results in a controlled environment, not through a compelling demo.
How Springtime approaches the people problem
Because confidence can’t be just conjured up in a demo, we design an implementation where the AP team’s process knowledge shapes how the system get configured from the very start.
Before go-live, Invoicetrack pre-trains on 18 months of the customer’s historic invoice data. The automation AP teams see from day one reflects the reality of their own invoice mix and their own supplier base, not a generic demo dataset. Exception rates in the pilot are visible, trackable and explained by cause.
Confidence thresholds that govern when the system acts autonomously and when it routes to a human are set by the customer’s process owners, not by Springtime. When a team member corrects an automated decision, that correction feeds back into the next model retrain. The system learns from the AP team’s expertise, which means practitioners are contributors to the automation, not passengers in it.
6. Proving ROI before the investment committee
Traditional ROI models for AP automation are built around cost per invoice. Agentic AI creates value across dimensions that cost-per-invoice does not capture, and it introduces costs that cost-per-invoice models routinely underestimate. Both gaps tend to surface at the same time, roughly six months into production, which is not a comfortable moment to be revising the business case.
SSON State of Accounts Payable Market Report 2026
Metrics that would move the needle most with leadership
ROI stories that lead with straight-through processing rates land better with executives than ones that lead with headcount.
On the cost side, EY’s June 2026 analysis splits an agent’s total cost into seven categories and finds that most business cases include only three: tokens, licenses and platform infrastructure. The four that surface later are governance, organizational change, failure recovery, and regulatory cost. An organization that scopes ROI based on API costs alone will consistently underestimate what production-grade agentic AP really costs to run at enterprise invoice volumes.
Vendor pricing models carry a second, less visible cost risk. Some vendors price per invoice and absorb the underlying token cost themselves rather than passing it along to the customer. That structure depends on token prices continuing to fall and on the vendor reaching enough scale to make the economics work. If either assumption isn’t realized, the risk does not stay exclusively with the vendor. It reaches your operation, in the form of price increases, service changes, or a vendor that cannot sustain the model it sold to you. Any ROI case should account for this by asking a vendor directly how their pricing model holds up if token costs plateau or increase, and what contingency exists if it doesn’t.
In a recent webinar, EY practitioners noted that establishing cost-per-invoice baselines is one of the most common gaps they encounter at the start of implementation engagements, and one of the most consequential for the credibility of the eventual ROI demonstration.
How Springtime approaches the ROI conversation
The business case for AP automation works better when it leads with touchless processing rates rather than cost per invoice. According to the SSON Market Report 2026, 35% of shared services leaders say achieving 80% or higher touchless processing is the metric that moves the needle most with leadership. Invoicetrack’s analytics layer tracks touchless rate, cycle time, and exception backlog continuously, so the metrics that matter to investment committees are available in production, not reconstructed for a quarterly review.
On the cost side, Invoicetrack’s pricing is volume-aligned and fully itemized: transaction fees, country fees, and connector fees are visible upfront, with rates that step down automatically as invoice volume grows. There are no hidden inference costs passed through separately and no structural dependency on token pricing trends. The per-invoice economics are predictable across the contract term, which matters when the investment committee asks what happens to the business case if AI infrastructure costs move.
The agent ceiling: why more agents is not the path to 80% and beyond
Agentic AI of course, has a place in your AP automation architecture, but it should not be the whole of it. Deploying agents to supervise other agents, or to make routine decisions a coded rule could have made instantly, and you’re creating a problem where none existed. The agent applies reasoning to a decision that never needed it, and that reasoning has to be paid for, monitored, and explained after the fact. A coded rule would have made the same call in an instant, for a fraction of the cost, with an audit trail no one will question.
If you scale that approach across enterprise volume you will see the bill due twice: a cost line that climbs with every invoice, and a chain of handoffs your auditors have to untangle a year later. That is not a sustainable position, financially or from a governance standpoint and it is the opposite of accountable.
Invoicetrack is built the other way round. Policy and rules handle the high-volume, structured decisions at a fixed cost per transaction. AI-validated logic is built and tested against your historic data before it goes near production. Targeted reasoning is reserved for the genuine exceptions where judgment earns its keep, surfaced to a human who owns the call. Every automated outcome traces to a named rule, specific evidence, an accountable party, and a confidence score, stated at the moment of the decision rather than pieced together long after.
That combination is what moves an AP operation past 50% touchless and keeps it there without processing costs that rise with volume. It is also what makes the result defensible to a CFO, an internal auditor, and a regulator.
High touchless rates are of course, achievable but they are not reached by adding agents alone. 80% touchless and beyond is possible, and sustainable, when its supported by a foundation of clean, remediated data, write access that is scoped before it is granted and an architecture that reserves reasoning for the decisions that need them most. Invoicetrack is built exactly for that, by design.
FAQs
According to the SSON State of Accounts Payable Market Report 2026, 57% of organizations report touchless rates below 25%, and a further 18% sit between 25% and 50%. That puts roughly three-quarters of AP teams below 50% touchless, against an advertised benchmark of 80%, more than a decade into serious investment in automation technology.
The shortfall is usually architectural rather than a limitation of the technology itself. Six conditions account for most of it: enterprise data that agents cannot reliably act on, ERP write access that governance will not permit, audit evidence that does not meet the standard auditors apply, an expanded fraud surface once agents can write to payment workflows, AP teams that were not brought into the design, and ROI models that underestimate the full running cost. Platforms where AI was added on top of an existing workflow engine rather than designed into it tend to hit these limits earlier.
Agentic AI has a role in AP, but it should not be the entire architecture. For high-volume, structured decisions like three-way matching, coded rules outperform agents on speed, cost, auditability, and reliability. Agent reasoning adds most value in genuine exception handling where contextual judgment is required and rules cannot be defined in advance.
Vendor master completeness, PO coverage rates, and GL coding consistency are the three fields that matter most with good and ready-to-use process documentation in place. A measurable readiness threshold should be defined for each before autonomous resolution goes live on any exception type. Consider manual effort to investigate errors, review and implement corrections.
Agentic systems with write access to payment workflows are high-value fraud targets. Over two-thirds of organizations reported attempted supplier fraud in the last two years and it’s forecasted to become a daily challenge for all organizations. Agent scope should never include autonomous payment execution without very strong controls such as pattern fraud detection within the AP process and human-in-the-loop validations at the payment authorization step, regardless of invoice value.
Token pricing represents only a portion of actual deployment cost. A fully loaded cost model should include platform infrastructure, governance and oversight, organizational change, human review capacity for escalations, and the cost of remediating failures. ROI framing that leads with touchless rate lands better with leadership than those that lead with headcount reduction.
Auditable AI reconstructs what happened after the fact. That reconstruction takes time, often several hours spent interrogating an agent about why it made a particular mistake. Multiply that by enterprise invoice volume and the problem becomes clear. Some providers run agents that supervise the work of other agents, adding layers of complexity, and each layer lengthens the time it takes to establish what happened and why. The scope for hallucination is large, and so is the risk that comes with it.
Accountable AI is broader. It is about who owns the outcome. An accountable AP operation tracks whether the AI works inside policy boundaries and learns from manual corrections, and it keeps every automated decision stated and available for review at any point rather than reconstructed on demand. The AP team decides what gets automated and how, which closes the door on agent misinterpretation. Agents run continuously on the underlying data, read it, identify patterns, enrich decisions against policy, and produce coded automation rules that a human tests and approves before anything goes live.