From Cell Phone Minutes to AI Tokens… Have We Seen This Pricing Playbook Before?

I remember when we used to watch the clock when we made a cell phone call. Now we watch the meter when we ask a question. The technology changed, but the pricing conversation sounds remarkably familiar.

I had a realization a few weeks back that made me laugh. I remember when owning a cell phone meant studying your plan before using the phone. How many local minutes did you have? Did long distance count differently? Were nights and weekends free? What happened if you crossed into another calling region? And would one long conversation create an unpleasant surprise on the next bill?

Then came text messages. We bought packages of those, too. Data arrived with its own allowances and overage charges. Every useful thing a phone could do seemed to come with another unit to count.

Today, I listen to the conversation around artificial intelligence and hear an echo: How many tokens? Which model? How much context? How many requests? What is the rate for input versus output? Is the faster tier worth it? What happens when the included usage runs out?

We have moved from counting minutes to counting pieces of a conversation.

The new rate card

AI providers have a legitimate reason to meter usage. Running a model consumes computing resources, and a long, demanding task can cost more to serve than a short, routine one. An API lets organizations pay for what they use, while subscriptions make everyday use easier to budget.

Still, the rate cards are becoming familiar in their complexity. OpenAI lists different rates for model families, input, output, cached input, processing speed, and some tools. Anthropic lists Claude Sonnet 5 at $2 per million input tokens and $10 per million output tokens. Google lists introductory Gemini 3.8 Flash pricing at $0.75 per million input tokens and $3.75 per million output tokens through the end of 2026, with higher standard rates stated for 2027. Those figures are snapshots, not apples-to-apples measures of capability or the total cost of finishing a job. [1][2][3]

The consumer sees a monthly plan. The developer sees a token ledger. The enterprise may see both, plus seats, feature allowances, rate limits, and separate charges for premium workloads. The more AI becomes woven into daily work, the more familiar that combination will feel to anyone who once decoded a wireless bill.

“Good enough” depends on what happens when it is wrong

One sales pitch especially catches my attention… use a cheaper model when you do not need the accuracy of a more capable one. I understand the engineering logic. There is no reason to pay for the most demanding option to sort routine requests or draft a first pass of an internal summary that a person will review.

But imagine a vendor saying, “If you only need 85% accuracy, take the cheaper plan.” Who wakes up hoping for 15% wrong answers?

The trouble is that accuracy is rarely one universal number. Eighty-five percent on one test may say little about your documents, your customers, or the decisions you intend to make. More capable models can also make mistakes. Paying more does not buy a guarantee of truth.

For a marketing brainstorm, an error may cost a few minutes of editing. For a security investigation, it might bury an indicator that matters or turn an assumption into a reported fact. For customer billing, legal review, or a production change, the cost of the wrong answer may dwarf the token savings.

The useful question is not “What percentage of accuracy can we afford?” It is… What kinds of mistakes can this workflow tolerate, how will we catch them, and what will a missed error cost?

That changes the economics. A cheaper model that needs repeated prompts, longer outputs, extra review, and corrections may be more expensive per completed task. A stronger model may cost more per token and less per reliable outcome. Or the inexpensive model may be exactly right for a narrow, well-tested task. Measure the work, not just the rate card.

A chip from another era

This also reminds me of the Intel 486SX and 486DX. The DX offered floating-point capability, while the SX was sold without a functioning floating-point unit. Intel describes the SX as serving market segments that did not need that capability. It is an enduring example of how technology companies package different capabilities for different budgets. [4]

AI models are not literally identical processors with a switch flipped. Smaller models, larger models, reasoning settings, and service tiers can involve real differences in architecture, compute, latency, and operating cost. The resemblance is in the buying experience: what capability do you get at this price, and which capability is reserved for the next tier?

Will AI have its “unlimited” moment?

Wireless pricing eventually changed. Unlimited national calling plans spread, and bundles made texting feel less like a scarce resource. The meters did not vanish entirely; attention shifted to data, speed, and the finer print of the plan. The FCC documented that transition more than a decade ago. [5]

I suspect AI will follow a similar path. Competition and improving efficiency could make broad access to capable everyday models a standard expectation. In fact, the shift has begun… OpenAI describes virtually unlimited messages for eligible base models on ChatGPT Enterprise, while explaining that some enterprise usage can still be billed by tokens. “Unlimited” can describe the experience without promising unlimited access to every model, tool, speed tier, or workload. [6]

My prediction is that the visible meter will move. Basic conversation may feel unlimited. Providers may differentiate on advanced reasoning, autonomous tasks, video, speed, priority capacity, integrations, and the amount of work an agent can perform. Organizations will still want budgets and safeguards, particularly when software can initiate thousands of requests without a human watching each one.

The future AI bill might look less like “How many questions did you ask?” and more like “How much work did you delegate?”

What I believe buyers should ask now

For leaders evaluating AI, I would ask four questions before choosing the cheapest or most capable model:

  1. What is the total cost per completed task? Include retries, tool calls, human review, and correction time.
  2. Which errors matter? Test with your own examples, including the rare failures that carry real consequences.
  3. What happens at the limit? Find out whether work slows, stops, switches models, or adds charges.
  4. Who can see and control usage? A predictable plan needs clear reporting, budgets, and ownership.

I am old enough to remember checking whether a call would use my remaining minutes. I hope we will someday look back at anxiety over every AI token with the same amusement. But if history is any guide, “unlimited” will not end the conversation about price. It will change which parts of the service we count.

And that may be the most familiar part of the story. 🙂


My Article Sources

  1. OpenAI, API pricing (accessed September 23, 2026).
  2. Anthropic, Introducing Claude Sonnet 5 (pricing update August 10, 2026).
  3. Google AI for Developers, What’s new in Gemini 3.8 Flash (accessed September 23, 2026).
  4. Intel, Intel 64 and IA-32 Architectures Software Developer’s Manual, Volume 1, Appendix D.
  5. Federal Communications Commission, Fifteenth Mobile Wireless Competition Report, discussion of unlimited calling plans.
  6. OpenAI, ChatGPT Enterprise and Edu: Models and limits (accessed September 23, 2026).

Leave a Reply

Your email address will not be published. Required fields are marked *