Most AI budgets go to sales and marketing. Most AI returns don't.
MIT went through 300 enterprise AI deployments and found 95% delivered no measurable profit. The money is going to the half of the business where AI demos best, not the half where it pays.
Search "enterprise AI ROI" right now and you get a lot of confident numbers. 171% average returns. 80% of deployments profitable. Follow almost any of them back and you land on a vendor's lead generation page. The primary research says something less flattering and a lot more useful.
MIT's 2025 GenAI Divide report worked through 52 executive interviews, 153 surveys, and 300 public deployments. 95% of pilots delivered no measurable P&L impact. Not 95% of side experiments. 95% of the pilots companies were funding and reporting on.
The reason is the part people skip past. It was not model quality. The models were fine. The failures came from workflow integration, tools that sat next to the work instead of inside it.
But the number I keep coming back to is the budget one. Between 50 and 70% of enterprise AI spend goes to sales and marketing. Back office automation returns the most. So the money is going to the half of the company where AI demos well, and the returns are sitting in the half nobody wants to put on a board slide.
I have run this from the inside. Running QA at scale on LLM training data pipelines, the thing that moved numbers was never the model. It was writing the rubric, training the people, and sitting in office hours until the workflow actually changed. Ramp time dropped about 20% and error rates about 15%. None of that came from a better model.
Klarna is the clearest public version of the same lesson. Their AI support agent handled 2.3 million conversations, cut resolution time from 11 minutes to under 2, and was tracking toward roughly $40 million in savings. Then in mid 2025 they started rehiring human agents. The CEO said they had focused too much on efficiency and cost.
That second half is worth more than the first. Volume metrics and quality metrics drift apart for months before anyone notices, because volume reports weekly and quality reports in churn. And nobody models the cost of unwinding a replacement that did not work. The rehiring ran past what the savings case ever showed.
The work that holds up is boring on purpose. Back office automation. Reconciliation. Document processing. Contract review. None of it is impressive in a demo. All of it is measurable, auditable, and reversible, which is most of why it survived.
That said, this is a product decision before it is a technology decision. Pick the use case by whether you can measure it and undo it, not by whether it looks good at the all hands. I have written here before that evals are the new acceptance criteria, and that every agent needs a blast radius before it ships. This is the same argument one level up. Choose the work where being wrong is cheap.
95% of pilots showing no profit is not a story about models. It is a story about where the money went. Put it in the boring half of the business and the number starts moving.