



AI token prices are falling. AI budgets are still climbing. The answer is a planning discipline finance already knows: driver-based planning.

Here is the strangest line item in enterprise finance right now. Over the past year, the blended market price of AI dropped roughly 67%, from about $18.40 to $6.07 per million tokens. Measured against late 2022, the cost of GPT-4-level output has fallen from around $20 to $0.40 per million tokens. By any normal procurement logic, AI should be getting dramatically cheaper to run.
And yet: 73% of enterprises report that their AI costs exceeded original projections, and more than half admit they don’t understand the full scope of what they’re spending. The unit price fell by two-thirds, and the budgets still broke. Newer data shows the gap widening, not closing: in a February 2026 survey of 500 finance leaders, 79% said their organizations experienced AI cost overruns in the past twelve months.
For finance leaders, that contradiction points to a clear conclusion: the issue is not simply unit price. It is consumption volume, and consumption is a planning problem, not a procurement problem.
Finance has managed this cost pattern before. Payroll is shaped by decisions made across the business, and a single top-line budget number rarely explains the result. Workforce planning works because it models the drivers beneath the expense. That same discipline is the right starting point for getting AI cost under control.
When AI costs rise while token prices fall, the missing variable is consumption. Driver-based planning is built to manage exactly that.
Most organizations initially book AI as software expense, which is understandable: the invoice arrives from a technology vendor. But traditional software licensing is largely contractual: annual terms, known renewal dates, and relatively predictable expense. AI inference is consumption-based. Every prompt, user, product feature, workflow change, and model selection can change the run rate. You are not simply buying software; you are funding computation by the unit.
Finance does not budget payroll by entering one salary number. It models the underlying drivers: role mix, compensation bands, hiring timing, geography, benefits load, and attrition. Payroll is the output; the drivers are the model.
AI token budgeting has the same architecture with different inputs. Monthly inference expense is the output. The drivers are active users, requests per user, average tokens per request, model mix, and price per token, across every product and workflow that uses AI. Budget the output without modeling the drivers, and variance becomes both predictable and difficult to explain.

The planning parallel is direct:
Both are variable, consumption-based, cross-functional costs. Both need driver-based planning, regular reforecasting, and clear ownership to be managed well.
The workforce parallel is the right starting point. But three differences should shape the AI cost model from the outset.
Hiring takes months and attrition is roughly predictable. AI consumption can double in a week: a product launch, a viral internal tool, one engineer restructuring prompts. Agentic workflows compound this: a single autonomous agent workflow can consume 10–50 times the tokens of a simple query, which is why Goldman Sachs projects total token consumption will grow roughly 24-fold by 2030 even as prices fall. An annual AI budget is stale before the second quarter. Rolling forecasts aren’t a best practice here; they’re the only format that makes sense.
Payroll rolls up relatively cleanly: departments own headcount, and headcount rolls to the P&L. AI rarely follows the org chart so neatly. A single model can support customer service, marketing, product, and finance. Without tags and dimensions, the provider invoice cannot be allocated, and costs that cannot be allocated cannot be governed. Finance needs visibility into how engineering instruments API calls, not just the bill that arrives at month-end.
Salary bands typically change on an annual cadence. Model pricing, capabilities, and the right model for a task can change overnight. A budget built on today’s model mix may be materially wrong by the next quarter even if usage is unchanged. AI planning must scenario-model the unit cost as well as the volume, for example by testing the impact of a lower-cost model becoming viable midyear.
NetSuite Planning and Budgeting (NSPB) does not need a dedicated AI cost module to support AI token budgeting. Its driver-based planning capabilities provide the core structure: define the dimensions, model the consumption drivers, forecast scenarios, and compare the plan with actual usage. In practice, the work follows four steps.
Start with the axes that make AI spend explainable: provider (for example, OpenAI, Anthropic, Google, or Azure), model or model tier, consuming team or business unit, product or workflow, and environment (production versus development). These dimensions become both the planning structure and the allocation logic for actuals.
The driver tree mirrors a headcount model with different math. Instead of headcount × compensation rate, the core calculation is:
Active users × requests per user × average tokens per request × model price per token = monthly AI expense
Each driver should be maintained independently. That makes the forecast useful when pricing shifts, usage changes, or a new workflow goes live.
Workforce scenarios test hiring pace and attrition. AI scenarios should test model selection, usage growth, and prompt efficiency: What if usage doubles? What if a high-volume workflow moves to a model at one-third the price? What if prompt optimization reduces tokens per request by 20%? These decisions need engineering inputs, but their financial impact belongs in the planning model before the decision is made, not after.
The model earns trust when actuals flow back into it. Major providers expose usage data through APIs or exports. The practical challenge is tagging API calls with the chosen dimensions and moving provider billing data into the GL structure so NSPB can report budget versus actual at the level of real consumption. The difficult part is rarely the data source; it is agreeing, across finance and engineering, who owns each tag.
If your organization already spends materially on AI inference (or expects to within the next twelve months), these questions reveal whether the planning foundation is ready:
Organizations that cannot answer these questions are managing AI reactively. That may be tolerable while AI remains an innovation-budget line item. It is not tolerable once the bill becomes visible to leadership and the board. The stakes are rising on the value side as well: in the same February 2026 survey, only 15% of finance leaders said they could calculate AI ROI without significant bottlenecks, while 83% expect clear, quantifiable returns within twelve months. A cost model built on explainable drivers is the prerequisite for closing that gap.
Across our NSPB client base, the question has flipped in about eighteen months: from “how can we use AI?” to “how do we govern AI financially?” That mirrors the broader market: 98% of FinOps practitioners now manage AI spend, up from 63% a year earlier. The pattern we see is consistent: a company has been booking AI as a lump sum in a technology expense account, the number has grown large enough to appear in board discussions, and now finance needs a model.
Once the dimensional structure is agreed and instrumentation is available, an NSPB AI cost model can often be built in four to six weeks. The more important work is the finance-and-engineering conversation: what data exists today, what tags are attached to API calls, and what is required to create a reliable actuals feed. That is fundamentally a planning conversation: the same kind of alignment finance and HR establish for workforce planning.
The organizations making the fastest progress are not necessarily the ones spending the most on AI. They are the ones that establish the planning model before the cost becomes material.
Every major technology shift eventually becomes a finance problem. Cloud computing required new infrastructure planning. Subscription models changed revenue forecasting. AI is now reshaping operating-expense planning, but it is not asking finance to invent a new discipline.
AI is not introducing a new planning problem. It is introducing a new consumption driver. Finance already knows how to plan variable, driver-based costs; the opportunity is to apply that discipline before surprises become recurring variances.
Teams that wait for engineering alone to solve AI cost visibility will keep explaining budget variances after the fact. Teams that build the model now, even an imperfect first version, can manage the spend as it changes.
Myers-Holum helps finance teams define the dimensions, build the driver model, and connect provider-billing actuals to NSPB’s budget-versus-actual framework. If AI spending is growing faster than your ability to forecast it, the next step is a diagnostic conversation about the data, drivers, and governance a practical model requires.
This post is part of Myers-Holum’s EPM series for NetSuite Planning and Budgeting leaders: practical planning frameworks you can take back to your team, not product overviews.


