Press "Enter" to skip to content

Deloitte AI Infrastructure Leader: From ‘Tokenmaxxing’ to ‘Valuemaxxing’

As enterprise AI usage climbs, companies are starting to ask a more basic question: What are they getting for all those tokens?

Nicholas Merizzi, a principal and U.S. Silicon2Service and AI infrastructure leader at Deloitte, says the conversation is shifting from “tokenmaxxing” – maximizing AI consumption – to “valuemaxxing,” or tying that consumption to measurable business outcomes. That change is pushing companies to scrutinize cost per outcome, rethink how they deploy AI agents and, at high enough volumes, consider owning AI infrastructure rather than relying exclusively on cloud APIs.

In this email Q&A with The AI Innovator, Merizzi discusses how to control the cost of multi-agent systems and what KPIs CFOs and CIOs should be watching.

The AI Innovator: What does the shift from “tokenmaxxing” to “valuemaxxing” mean in practical terms?

Nicholas Merizzi: Earlier this year, the notion of tokenmaxxing was all the focus. The reality settled in when costs jumped significantly almost overnight. This is when we pivoted to focusing on value.

Tokenmaxxing treats tokens generated as a sign of engagement and usage (similar to measuring “lines of code” in a program from decades ago), while valuemaxxing is focused on deriving value and business outcome, forcing enterprises to balance design decisions.

Organizations are increasingly focused on the total cost of AI and maximum value per token rather than just token counting. What it means in practical terms is that leaders have set up the right dashboards, established financial governance, included it in the 2027 budgeting process, and reviewed their spend to ensure their token usage is tied to the products that make the most sense.

What KPIs should CFOs and CIOs be watching to decide whether the value created by an AI application actually justifies its token, compute and infrastructure costs?

Executives need a business value frame typically centered on five key metrics: revenue, cost, client experience, speed, or quality. For CFOs this often translates to measures like realized revenue, margin improvement, and cost avoidance, in addition to quality, reliability, security, engineering efficiency, and visibility into what drives cost.

Organizations are increasingly looking to tie cost to outcomes such as resolved tickets and the numbers of summarized documents, served customers, and lines of code checked-in and accepted. Although this is not an exhaustive list, you can see the shift from just creating dashboards that monitor cost and token count to measuring cost per business outcome.

Multi-agent systems could dramatically increase token consumption. How should enterprises design and govern these systems so the cost justifies the value gained?

Multi-agent systems are complex by nature. Because of this, being clear about the purpose of each agent in that system is vital. The data and tools agents require, and the actions they are approved to take, are the most important factors. For design and governance practices enterprises should focus on four main areas: understanding where generation adds value, controlling context and route intelligently, enforcing budgets and observability, and managing governance control. 

Other leading practices include utilizing deterministic code wherever consistency matters, applying agents selectively, and tightly controlling the context and models they use. Governing agents can be done by enforcing budgets, permissions, observability, reliable fallbacks, and centralized oversight.

At what point should rising AI consumption change an enterprise’s infrastructure strategy – for example, using smaller models, moving some inference on-prem or investing in dedicated AI infrastructure rather than continuing to pay for cloud APIs?

Deloitte’s paper on the Economics of AI measured the inflection point for when private AI becomes the most economic option. I have personally been designing and building out AI factories for enterprises in 2026 at a much greater velocity given that owning your own AI factory can be the most economical path for heavy token volumes.

This enables enterprises to shift to a “no token architecture,” bringing more predictable spend and being deliberate in their design and consumption of external frontier models. Combining the use of an AI factory with technologies such as model routing, the use of open weight models, and embracing small language models are all areas that CIOs should be exploring.

How long should a business wait for ROI from AI before deciding it should move on to another deployment?

The answer will vary by use case, with ROI for some agentic workflows being more immediate or visible and others will have a longer runway as we redesign the process to be agent-centric. Businesses should treat observability as a first-class citizen in their systems and measure the cost, performance, behavior, and delivered value of the entire system.

Also understanding the importance of addressing skill gaps and marketplace adoption of the product is very valuable. You must have the right people building the systems and solutions, and the right message in the marketplace to help customers realize the value you see.

That said – if repeated iterations don’t provide a credible path to adoption or quality, or the workflow or use case itself is evaluated as not in alignment with business objectives or doesn’t provide a path to potential profit, I would recommend measuring and capturing lessons learned from that work to improve the chances of success for other use cases.

Author

Get the latest insights about enterprise AI.

Subscribe to our newsletter. Thank you.

Be First to Comment

Leave a Reply

Your email address will not be published. Required fields are marked *

×