AI agents are the busiest category in enterprise software right now, which is exactly why most of the money spent on them is wasted. The useful question is not "should we use AI agents" but "which parts of this are already a $99-per-month product, and which parts will no vendor ever build for us?" Getting that line wrong in either direction is expensive.
The failure numbers, and what actually causes them
In June 2025, based on a poll of more than 3,400 organizations, Gartner predicted that over 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls. The same firm expects agentic AI to make around 15% of day-to-day work decisions by 2028, up from essentially zero in 2024 - so the technology is not the problem.
The MIT report The GenAI Divide: State of AI in Business 2025 put it more bluntly: roughly 95% of pilots delivered no measurable profit-and-loss impact, drawn from 52 executive interviews, 153 leader surveys, and 300 public deployments. That figure got repeated everywhere, usually without its two most useful findings:
- Tools built with external vendors succeeded about twice as often as internal builds. The instinct to "just have our own team try it" is, statistically, the worse bet.
- Budgets concentrated in sales and marketing, which is where measured ROI was lowest. The agent work that pays is unglamorous back-office cost removal, not campaign generation.
Underneath all three of Gartner's stated causes sits one root problem: the project began as a capability demonstration rather than as a named metric with a baseline. Nobody agreed in advance what success would look like, so nobody could prove it happened, so the budget got cut. That is the specific failure our pricing model is designed to make impossible - if we cannot define and measure the metric, we do not get paid, so we will not start.
Buy this. Do not hire us for it.
We would rather lose the project than sell you a build you can replace with a subscription. Where products already win:
- Front-desk call answering for home services. This is a solved, funded category - Jobber, Goodcall, Numa, Smith.ai and others sit roughly in the $49 to $300 per month range, and Avoca reached unicorn valuation building voice AI specifically for HVAC and plumbing. A custom build cannot compete with that on cost or on time-to-value. (Note that the widely quoted missed-call statistics in this space - 27% of calls unanswered, around $1,200 of revenue lost per missed call, 30-50% booking lifts - are vendor-published marketing figures, not independent research. Use your own call logs.)
- SMB messaging in Brazil. Meta released WhatsApp Business AI to Brazilian SMEs in early 2026, after testing in Mexico. It answers customers around the clock from your catalog, website, and stored policies, with no programming, no integration, and no additional cost. Paying anyone to rebuild that is money set on fire. Kantar research cited by Meta found 88% of Brazilian adults consider messaging a fast way to reach companies and 75% are likelier to buy from brands they can message - the demand is real, but the tool to meet it is now free.
- Generic support deflection. Horizontal platforms have industrialized this. If your tickets are mostly standard questions answerable from a knowledge base, buy the platform and spend your effort on the knowledge base and escalation design, which is where the outcome actually lives.
- Early-stage receivables chasing. Collections is already a contingency market - traditional agencies take roughly 25-50% of recovered amounts, and AI-native platforms advertise success-only fees in the 5-15% range. A consultancy building you a bespoke dunning agent is unlikely to beat that economically. (Recovery-rate claims from these vendors are self-published; treat them as marketing.)
The commoditization pressure here is real and it is aimed squarely at consultancies. A Google Cloud executive said publicly in February 2026 that thin-layer LLM wrapper companies face existential risk as foundation model providers absorb their functionality, and Gartner has been reported as expecting around 40% of enterprise applications to embed vertical agents by the end of 2026. Any firm still selling "we'll build you a chatbot" in that environment is selling you a depreciating asset.
Where custom agents still win
Products are built for the workflow that most customers share. The money a product cannot reach sits in the parts of your operation that are specific to you:
- Work that spans systems no single vendor covers. When a task requires reading your TMS, checking the ERP, cross-referencing a spreadsheet somebody maintains by hand, and updating a homegrown database, there is no product for that - because that combination exists only at your company. This is where the integration depth argument actually holds.
- The exceptions products treat as out of scope. Vendor agents handle the happy path and hand everything else to a human. In most operations, the happy path was never the expensive part. The unbilled accessorial, the mismatched invoice, the order stuck between two systems - that residue is where the recoverable money concentrates.
- Legacy software with no integration market. If your core system is old enough that nobody sells a connector for it, no vertical SaaS roadmap is coming to rescue you.
- Recovery work, not deflection work. Agents that find money someone else then pays you - unbilled exceptions, under-invoiced services, expired claim windows - produce a number verified outside your own estimate. That is both better business and the only kind of agent work we can honestly price on results.
How we measure an agent, and why deflection is a bad metric
Most agent reporting optimizes for the wrong thing. Deflection rate rewards an agent for ending conversations, not for resolving them - a system that frustrates people into giving up scores beautifully on it. What we will actually agree to be paid against, in order of how verifiable they are:
| Metric | Who verifies it | Strength |
|---|---|---|
| Recovered money (receivables collected, unbilled exceptions invoiced) | The paying counterparty, then your AR | Strongest - we cannot influence or estimate it |
| Fully resolved contacts with no human handoff | Your ticketing system, against a pre-deployment baseline | Good, if resolution is defined before launch and re-contact within a window counts as unresolved |
| Staff hours returned on a named process | Your own time and volume records over matched periods | Workable, and the number most often inflated - so we hold it to the same baseline discipline |
| Deflection rate, containment rate, messages handled | Nobody meaningful | We will not price on these, and neither should you |
Guardrails we build in by default
- Escalation over guessing. Below a confidence threshold the agent hands to a human with the full conversation context, and every escalation is logged as training signal rather than hidden as a failure.
- No autonomous financial or contractual commitments. An agent may prepare, draft, and recommend; a human authorizes anything that binds you.
- Read-only until proven. Agents start observing and drafting before they are ever allowed to write to a production system.
- Complete audit log of actions taken and data read. Gartner named inadequate risk controls as a top cancellation cause, and the same logging that satisfies risk review is what makes the results measurable enough to price on.
The honest summary
AI agents are a genuinely enormous market and most of it is not ours to sell. If your need is a front desk, an SMB inbox, or standard ticket deflection, buy the product - we will name it on the call and charge you nothing. If your need is the multi-system, exception-ridden, legacy-bound operational work that no vendor will ever productize, that is a real project, and we will take it on with no upfront fee and get paid from what it recovers.