The Shortening Half-Life of Intelligence

Unit Cost Discontinuity · A trilogy

The Shortening Half-Life of Intelligence

How a model event becomes an industrial input, then branches into labour and balance-sheet consequences.

Three essays by Ben Luong · 30 July 2026

Read all three as one page

The Shortening Half-Life of Intelligence: fact, mechanism, consequence

Field update · 30 July 2026

OpenAI has just priced Part II into the market.

Twenty-one days after GPT-5.6 entered broad public availability, OpenAI cut Luna’s API price by 80% and Terra’s by 20%, while leaving the frontier Sol price unchanged.

80%Luna cut$1 / $6 to $0.20 / $1.20
20%Terra cut$2.50 / $15 to $2 / $12
UnchangedSol price$5 / $30 per million tokens
About 90 daysIllustrative half-lifeVendor claim, not a time series

The asymmetry is the trilogy’s mechanism in miniature. The apex can retain a premium while the broad, high-volume layer is repriced. OpenAI does not use the phrase accepted-output cost. Its deployment guidance now tells businesses to define the required outcome and quality standard, test where additional intelligence changes the result, and send well-specified implementation and testing to Luna after Sol has resolved uncertainty and planned the work.

The frontier model keeps the championship work. The cheapest adequate model receives the routine volume.

OpenAI says technical efficiency created the room for the cuts: better routing, serving software and context management produce more useful work from the same compute. It reports that Sol-assisted kernel work reduced end-to-end serving cost by 20%, while related experiments improved token-generation efficiency by more than 15%. Axios separately reports that cheaper Chinese open-weight models have increased pressure on OpenAI and Anthropic to justify their premiums. The announcement does not identify how much of the pass-through came from efficiency and how much from competition.

The outside option therefore need not win the account to matter. OpenAI can retain the customer while repricing the tier that carries routine volume. OpenAI also claims Luna now matches models that were frontier-class a year ago at roughly six cents on the dollar per task and nearly nine times the speed. That is a vendor comparison, not a controlled time series. Taken only as an illustration, six cents represents just over four cost halvings in a year, or an economic half-life of about 90 days.

This is not the credit event described in Part III. Efficiency may protect provider margins, and lower prices may stimulate enough demand to increase total compute use. Nor does the cut show that a lower acquisition basis for distressed capacity would reach customers after a crash. That is a separate pass-through mechanism. The announcement does strengthen the prerequisite: the realised selling price of economically adequate intelligence is compressing quickly, while the expensive frontier is being reserved for the narrow steps where its advantage changes the outcome.

The complete transmission

Every arrow is conditional. The model is not linear.

  1. 01
    Near-frontier open weights Capability becomes reproducible outside the originating laboratory.
  2. 02
    Credible outside option Contestability begins before buyers migrate.

Two routes to cheap cognition

03A
Competitive hosting Many providers expose price differences and compete to pass efficiency through.
03B
Integrated undercutting A laboratory or hyperscaler makes its own volume tier cheap.

From contestability, the argument forks

Branch A

Labour economics

Where either delivery route lowers accepted-output cost

  1. 04A
    Lower all-in cost per accepted output Inference, retries, verification, error and supervision all count.
  2. 05A
    Unit Cost Dominance over labour The machine system clears the acceptance threshold at lower total cost.
Branch B

Credit economics

Only if effective supply outruns paid demand, or rental economics fall below debt service

  1. 04B
    Lower rental rates, utilisation or collateral values The leveraged edge loses the economics it underwrote.
  2. 05B
    Credit stress Refinancing, covenants and impairments make the break visible.
  3. 06B
    Written-down capacity and cheaper access Somebody else has absorbed the original capital loss.

Kimi K3 makes the first step difficult to deny. Delivery can be dispersed or concentrated. The labour branch requires lower accepted-output costs. The credit branch requires a paid-workload, rental or debt-service failure. Only the post-write-down feedback requires competitive pass-through.