Anthropic now allows paid Claude subscribers to continue working after reaching their included usage limits by purchasing credits billed at standard API rates. Its usage documentation describes five-hour reset periods and shared consumption across Claude and Claude Code. A subscription therefore provides access to a variable amount of work before consumption pricing resumes.

The arrangement reflects a broader transition in generative AI. Early adoption was supported by generous access, rapid capability improvements, and intense competition for users.

As usage becomes operationally important, model providers must reconcile high infrastructure costs and investor expectations with customers that increasingly scrutinize cost, reliability, and control.

The financial expectations are unusually large. OpenAI announced a funding round of roughly $120 billion at a valuation above $850 billion, while Anthropic's 2026 financing reportedly placed its valuation near $1 trillion.

Those valuations require the labs to capture a durable share of the economic value created by AI. This must occur while chip suppliers, cloud platforms, and enterprise integrators compete for the same revenue pool. OpenAI also reported that enterprise customers already represented more than 40 percent of its revenue.

Key Points


  • Token pricing measures computational activity rather than customer value or completed work.
  • Rapid capability diffusion subjects older frontier performance to continuing price competition.
  • Open-weight models provide portability, custody options, and bargaining leverage even when self-hosting carries substantial costs.
  • Outcome-based pricing brings model vendors closer to proprietary customer knowledge and raises new trust concerns.
  • Enterprise deployment requires deterministic controls, auditable evidence, and accountable decision processes around probabilistic models.
  • Pricing surprises and restrictive terms can accelerate investment in architectures designed to keep model suppliers replaceable.

The Contractible Unit


A token is a small unit of text processed or generated by a language model. It is also the cleanest contractible unit available to a vendor operating outside the customer's business. The vendor can count computation consumed without determining whether the resulting work reduced costs, improved a decision, or produced an unusable answer.

That distance creates an incentive problem. A concise workflow can use fewer tokens than an agent that repeatedly reloads context, follows unproductive paths, or retries failed operations.

Consumption growth can therefore represent productive adoption, computational waste, or a mixture of both. Revenue based on token volume does not distinguish among those outcomes.

Customers ultimately care about cost per accepted result. An expensive frontier model may complete a difficult task with fewer retries and less supervision, making it cheaper than a lower-priced alternative. A nominally inexpensive model can impose greater review and correction costs.

Context management, workflow design, and human oversight determine much of the economic result after the published token rate is applied.

The contract can use other units, including seats, reserved capacity, or completed tasks. Each remains measurable at the vendor boundary.

A model supplier seeking a share of a drug discovery, legal outcome, or engineering breakthrough must move closer to the customer's internal definition of value. That move requires access to proprietary context and a negotiated right to participate in the outcome.

Frontier AI valuations depend on this boundary moving. Token revenue can support a large infrastructure business, particularly as global consumption expands. Valuations approaching $1 trillion imply broader platform power and sustained participation in enterprise value.

The central commercial problem concerns how a model vendor can obtain that participation without weakening the customer's control over its data, workflows, and institutional advantage.

More Technology Articles

A Moving Frontier


Model capability is diffusing quickly. The Stanford AI Index found that the inference cost of reaching approximately GPT-3.5 performance fell more than 280-fold between late 2022 and late 2024.

Performance that once required a leading commercial system became available at a small fraction of its earlier cost.

The current frontier can still command a premium where additional capability materially improves completion rates. Stanford's 2026 technical report found that the measured gap between leading open- and closed-weight models had widened after narrowing sharply, while remaining within single digits.

The market contains a moving premium tier above an increasingly capable and competitive middle.

The durability of the model layer depends on the relationship between those two movements. Frontier labs must renew commercially important capability faster than earlier performance diffuses into cheaper alternatives.

Enterprise demand may continue requiring the strongest available models for difficult work. Many routine workloads may settle on models that meet an acceptable threshold at lower cost.

Open-weight models strengthen that competition by allowing multiple operators to host, customize, and price access to the same underlying capability. Their strategic value includes portability, bargaining leverage, and greater control over inference custody.

Self-hosting introduces hardware, security, and staffing costs, while managed open-weight services retain some of the operational dependencies associated with any external provider.

Model origin remains relevant when open weights come from a geopolitical competitor. Internal or domestic hosting can keep prompts and outputs outside a foreign provider's custody.

Security weaknesses, licensing restrictions, and model behavior remain part of the deployment decision. A NIST evaluation of DeepSeek models identified performance, security, and censorship concerns alongside their growing adoption.

Custody, Outcomes, and Control


OpenAI and Anthropic state that commercial inputs and outputs are excluded from model training by default. That policy addresses one form of exposure.

Sensitive information can still pass through stored context, tool calls, and third-party connectors, with retention and access conditions varying across services. Enterprise custody analysis must cover the complete workflow rather than the model-training policy alone.

The distinction matters because prompts increasingly contain institutional knowledge rather than generic requests. An agent working across internal systems may encounter product plans, customer relationships, or operational performance.

These materials can constitute competitive advantage even when they fall outside conventional classifications of protected intellectual property.

Sam Altman's reported discussion of subsidizing AI-assisted drug discovery in exchange for royalties made the value-capture issue explicit. According to Bloomberg Law, the proposal contemplated covering model costs in exchange for a royalty from successful discoveries.

Outcome participation could align price with customer success more closely than token billing. It would also bring the supplier closer to the risks created by the outcome.

A lab taking a royalty from a successful drug is no longer merely selling computation at arm's length. It has acquired an economic interest in the development program. General counsel would therefore examine not only confidentiality, attribution, and intellectual property, but also how the agreement allocates indemnity, regulatory responsibility, and potential liability for downstream harm.

The token meter limits the vendor's value capture partly because it also keeps the vendor distant from the consequences. Moving beyond that meter means accepting or contractually allocating some portion of the risk and governance burden surrounding the upside.

This is the vendor's bind: the downstream value it wants to capture is also the value most expensive to approach. The greater its access to proprietary workflows and successful outcomes, the greater the customer's concerns about custody, lock-in, conflicting economic interests, and responsibility when those outcomes fail.

Institutional trust determines how far a supplier can move beyond metered computation into the systems where the customer's most valuable work occurs.

Enterprise Accountability


Probabilistic models require deterministic controls when their output enters consequential enterprise systems. Access policies must govern what a model can see and do.

Auditable methods must record which evidence, model, and rule set produced an action. Accountable people must remain able to review decisions whose consequences the organization cannot transfer to a software supplier.

The NIST Generative AI Profile identifies governance, pre-deployment testing, content provenance, and incident disclosure as core areas for managing generative AI risk. These controls remain necessary as models improve. Greater capability expands the range and consequence of actions that require controlled access and review.

Operational responsibility remains concentrated at the deployment layer because that is where generated output becomes institutional action. OpenAI's business agreement makes customers responsible for evaluating the accuracy and appropriateness of output for their use cases.

Integrators create value by limiting what the model can do, checking its work, and preserving a record that approved safeguards were followed.

Palantir CEO Alex Karp's criticism of enterprise "tokenmaxxing" translates this operational requirement into a theory of market structure: Palantir benefits when models remain selectable components and the governed deployment layer retains the customer relationship, but consequential systems require permissions, verification, and accountability regardless of which model is strongest.

A related control failure can be described as compliance theater. A model can acknowledge an instruction, explain a previous error, and promise different behavior without creating a persistent rule that binds later output.

The correction becomes operational only when it is translated into enforceable policy, durable configuration, or repeatable testing. Generated assent provides no evidence that the underlying control has changed.

Designing for Replaceability


Enterprise customers are already organizing technical systems around model choice. Amazon Bedrock markets access to many foundation models through common interfaces and provides tools for comparing performance, cost, and accuracy. The customer can change models without reconstructing the entire application around a new provider.

Open standards reinforce that approach. Anthropic donated the Model Context Protocol to a Linux Foundation project after the protocol gained adoption across several major AI products. OpenAI added remote MCP support to its Responses API.

A common integration layer expands the market for connected AI applications while reducing the cost of substituting one compatible model or service for another.

Frontier vendors are simultaneously building above the raw model. OpenAI's Responses API combines models with hosted tools and execution tracing, while Anthropic offers agent tooling with permissions and context management.

These products recognize that customers purchase completed workflows with operational controls. They also place the labs in direct competition with cloud platforms, integrators, and software companies seeking to own the same customer relationship.

The timing creates an asymmetry. A model provider can change prices, limits, and terms quickly. Enterprises change architecture and counterparty policy over much longer periods. A customer that responds to pricing uncertainty by qualifying alternative models may retain that capability after the original restriction is removed.

Customer-hostile monetization can therefore accelerate the contestability that threatens the frontier premium. Each pricing surprise strengthens the internal case for routing layers, portable data, and self-hosted options.

The resulting architecture reduces dependence on a single supplier and usually persists because dismantling it provides little operational benefit.

Visible complaints provide an incomplete measure of this response. Switching costs can preserve usage after goodwill has deteriorated. Customers may continue using a service while testing replacements and moving sensitive work elsewhere.

Churn becomes visible after the strategic decision has already been made and the replacement path has absorbed much of the migration cost.

Restrictions alone provide insufficient evidence of balance-sheet distress. Capacity constraints, licensing obligations, and abuse controls can also shape product policy. Suno, for example, attributed its 2026 download limits to licensing and abuse concerns.

The strategic consequence remains measurable: separating creation from possession reduces portability and changes the value customers receive from the subscription.

Frontier AI is entering an industrial market in which capability diffuses downward while the frontier continues advancing. Model vendors can readily meter computation, yet their valuations require access to a much larger share of the value produced downstream.

That access depends on customer confidence in pricing, custody, and institutional boundaries.

Durable frontier companies will renew their capability advantage while building products above it. They will also need to preserve customer control over proprietary context and support enforceable governance around probabilistic output.

Vendors that exploit temporary scarcity through unpredictable limits or aggressive claims on downstream value give customers a direct incentive to design the next generation of enterprise AI around their replacement.

Sources


Article Credits