
This article first appeared in Forum, The Edge Malaysia Weekly on August 24, 2026 - August 30, 2026
The unit cost of artificial intelligence tokens, the smallest unit of data processed by an artificial intelligence (AI) model, has plummeted since 2023. Yet across the corporate landscape, chief financial officers and engineering leaders are facing unexpected AI invoice shocks.
Over the past year, organisations have aggressively nudged employees to integrate AI into every workflow. Developers were given access to tools like Claude Code, Cursor and Copilot, who outfitted AI coding assistants and autonomous AI agents. In many companies, raw usage was actively encouraged and linked directly to team performance metrics.
The result? A cultural and operational phenomenon known as “tokenmaxxing”, the assumption that higher token consumption directly correlates with higher productivity. The economic reality, however, is telling a very different story.
Historically, when the unit cost of a resource drops, total spending drops too. In enterprise AI, the opposite has occurred. This is a classic case of Jevons paradox, an economic rule where making a resource cheaper and more efficient actually causes overall demand for it to skyrocket, driving total spending up instead of down.
Organisations pushed AI adoption without establishing clear guardrails or tying consumption to tangible business metrics and the results were:
● Budget exhaustion: Uber reportedly burnt through its entire annual AI budget in the first four months of the year.
● Licence revocations: Enterprise leaders, including Microsoft, made headlines by revoking developer access to high-tier AI capabilities after usage spiked beyond projections.
● Massive usage bloat: Leaderboards at major tech firms incentivised workers to generate massive token counts, rewarding raw activity over business impact. Meta employees reportedly consumed tens of trillions of tokens in a single month.
As OpenAI’s head of enterprise Alexander Embiricos recently observed: “Six months ago, conversations were all about ‘What can it do? Is it good enough?’ Today, conversations are never about that. Now they’re about: ‘Hey, we’re spending so much. What visibility do you have? What token controls do you offer?’”
Why are advanced AI systems consuming tokens at an unprecedented rate?
The primary engine behind this inflation is the shift from simple single-prompt chat windows to autonomous agentic workflows (AI programmes designed to independently execute multi-step tasks without human intervention at every step).
Next-generation reasoning models don’t just return immediate text. They think longer, execute multi-day tasks and constantly re-read massive context histories. This structural evolution drives four primary cost drivers:
● Unchecked operational loops:
Development teams frequently build background “agentic loops”, automated routines where an AI constantly queries itself to complete repetitive technical tasks. An operational loop running endless background checks can escalate a single developer’s monthly API bill from US$20 to thousands of dollars overnight.
● Burning tokens to save tokens:
To manage primary AI agents, organisations have started building secondary monitoring agents. This creates a compounding cycle of token consumption, which means, one is spending money on AI compute just to supervise other AI compute.
● Diminishing returns on developer velocity:
Higher token burn does not automatically equal better software output. A study by management platform Jellyfish revealed that developers consuming the highest token volumes were roughly twice as productive as low-usage peers, but consumed 10 times the tokens to achieve that incremental gain.
A two-year study of 20,000 developers by Faros found that while raw code output increased with AI assistance, so did code churn and bugs, and required rewrites. Fast AI code generation without careful supervision often shifts work down the line to expensive debugging.
● Erosion of core software development fundamentals:
Over-reliance on AI can lead engineering teams to bypass basic programming fundamentals. Developers increasingly run expensive AI model queries to solve tasks that standard database updates or traditional software logic could handle at a fraction of a cent.
● The FinOps blind spot:
Traditional cloud computing (AWS, GCP, Azure) spent a decade developing mature FinOps (the business practice of managing and optimising cloud costs). They now use strict billing alarms, resource tagging and centralised spending approvals.
Enterprise AI deployment bypassed much of this financial infrastructure. API keys (the digital passwords that allow software to access and charge AI tools) were decentralised across individual business units and developers. When a massive end-of-month invoice arrives, IT leaders lack the telemetry, that is, the real-time automated tracking and measurement data to identify which specific product, developer team or runaway agent generated the bill.
Exceeding an AI budget is not entirely negative. It signals that teams are actively attempting to adopt digital tools. However, enterprise strategy must pivot from raw usage adoption to strict financial and operational optimisation.
Strategy 1: Establish formal AI FinOps and governance
● AI as a critical budget line item: Treat AI compute spend as a major budget line item on a par with office rents, telco costs or corporate travel.
● Dedicated roles: Medium to large businesses should introduce specialised profiles in the company, such as AI FinOps specialists (experts in managing AI compute costs), AI integration specialists and AI business analysts.
● Centralised procurement: Negotiate enterprise-tier AI provider agreements with centralised key management and built-in spending limits.
Strategy 2: Dynamic AI model selection and tiered access
● Abandon the “all-purpose model” mindset: Top-tier, massive reasoning models should not be used for simple text summaries or basic file formatting. Route simple tasks to smaller, inexpensive language models, reserving flagship frontier models such as OpenAI’s GPT-5.6 Sol, Anthropic’s Claude Fable 5 and Google’s Gemini 3.5 Pro exclusively for high-complexity, business-critical workloads.
● Service tiering: Establish clear access tiers. Using the most expensive AI reasoning models should require formal manager approval, similar to booking last-minute business class travel, while general teams operate within set monthly token budgets.
Strategy 3: Architectural efficiency with inbuilt ‘tokenomics’
● “Create once, run anywhere”: Avoid calling on the AI over and over again to do repetitive daily work. Instead, have the AI write the instructions or automation programme once, export it and run that finished programme on regular, cheap computers without paying for continuous AI usage.
● Granular cost tracking and observability: Implement real-time tracking dashboards to monitor token usage attributed to specific users, business units, product features, developer-level spend or individual agentic loops.
● Standardised AI economics: Adopt emerging frameworks for tracking AI spend. Measure consumption through metrics such as cost-per-intelligence (CPI) — how much cash it costs to accomplish a specific complex outcome — rather than raw token counts.
Strategy 4: Align incentives with business return on investment
● Eliminate usage leaderboards: Stop measuring AI adoption by raw token usage or total hours spent chatting with an AI.
● Tie AI metrics to business outcomes: Evaluate AI success based on actual finished projects, faster customer resolution times or direct profit margins.
● Prioritise quality AI over AI speed: Fast AI-generated code can turn into long-term technical debt. Implement mandatory automated testing. Speed of delivery is meaningless if it creates costly software bugs downstream.
Strategy 5: Partner with outcome-focused AI integration experts
● Engage strategic AI integration partners: Collaborate with AI advisory and integration specialists who design an end-to-end strategy, directly tying model deployment to measurable business outcomes, operational efficiency and financial return on investment.
The first phase of enterprise AI adoption was defined by enthusiasm, experimentation and unconstrained token usage. The next phase will be defined by financial engineering and operational governance.
Organisations that master AI FinOps, matching the right model size to the task and measuring true return on investment, will turn AI from an unpredictable operational expense into a sustainable competitive advantage.
Akhil Gupta is deputy chair of Pikom’s (National Tech Association of Malaysia) AI chapter and group CEO of a conglomerate delivering IT infrastructure, AI and talent solutions globally
Save by subscribing to us for your print and/or digital copy.
P/S: The Edge is also available on Apple's App Store and Android's Google Play.