The Shortening Half-Life of Intelligence

Unit Cost Discontinuity · A trilogy followed by a coda

The Shortening Half-Life of Intelligence

A three-part trilogy on contestable cognition, followed by a coda on where the surviving rent goes.

Three essays and a coda by Ben Luong · 30 July 2026

Read the trilogy and coda as one page Copy it

The Shortening Half-Life of Intelligence: fact, mechanism, consequence

The first three essays remain a trilogy: fact, transmission mechanism, then financial and political consequence. Rent Is Not Conserved follows as a coda. It corrects the trilogy’s final movement from cognitive capital to physical capital while remaining outside the numbered sequence.

Field update · 30 July 2026

OpenAI has just priced Part II into the market.

Twenty-one days after GPT-5.6 entered broad public availability, OpenAI cut Luna’s API price by 80% and Terra’s by 20%, while leaving the frontier Sol price unchanged.

80%Luna cut$1 / $6 to $0.20 / $1.20
20%Terra cut$2.50 / $15 to $2 / $12
UnchangedSol price$5 / $30 per million tokens
About 90 daysIllustrative half-lifeVendor claim, not a time series

The asymmetry is the trilogy’s mechanism in miniature. The apex can retain a premium while the broad, high-volume layer is repriced. OpenAI does not use the phrase accepted-output cost. Its deployment guidance now tells businesses to define the required outcome and quality standard, test where additional intelligence changes the result, and send well-specified implementation and testing to Luna after Sol has resolved uncertainty and planned the work.

The frontier model keeps the championship work. The cheapest adequate model receives the routine volume.

OpenAI says technical efficiency created the room for the cuts: better routing, serving software and context management produce more useful work from the same compute. It reports that Sol-assisted kernel work reduced end-to-end serving cost by 20%, while related experiments improved token-generation efficiency by more than 15%. Axios separately reports that cheaper Chinese open-weight models have increased pressure on OpenAI and Anthropic to justify their premiums. The announcement does not identify how much of the pass-through came from efficiency and how much from competition.

The outside option therefore need not win the account to matter. OpenAI can retain the customer while repricing the tier that carries routine volume. OpenAI also claims Luna now matches models that were frontier-class a year ago at roughly six cents on the dollar per task and nearly nine times the speed. That is a vendor comparison, not a controlled time series. Taken only as an illustration, six cents represents just over four cost halvings in a year, or an economic half-life of about 90 days.

This is not the credit event described in Part III. Efficiency may protect provider margins, and lower prices may stimulate enough demand to increase total compute use. Nor does the cut show that a lower acquisition basis for distressed capacity would reach customers after a crash. That is a separate pass-through mechanism. The announcement does strengthen the prerequisite: the realised selling price of economically adequate intelligence is compressing quickly, while the expensive frontier is being reserved for the narrow steps where its advantage changes the outcome.

The complete transmission

Every arrow is conditional. The model is not linear.

  1. 01
    Near-frontier open weights Capability becomes reproducible outside the originating laboratory.
  2. 02
    Credible outside option Contestability begins before buyers migrate.

Two routes to cheap cognition

03A
Competitive hosting Many providers expose price differences and compete to pass efficiency through.
03B
Integrated undercutting A laboratory or hyperscaler makes its own volume tier cheap.

From contestability, the argument forks

Branch A

Labour economics

Where either delivery route lowers accepted-output cost

  1. 04A
    Lower all-in cost per accepted output Inference, retries, verification, error and supervision all count.
  2. 05A
    Unit Cost Dominance over labour The machine system clears the acceptance threshold at lower total cost.
Branch B

Credit economics

Only if effective supply outruns paid demand, or rental economics fall below debt service

  1. 04B
    Lower rental rates, utilisation or collateral values The leveraged edge loses the economics it underwrote.
  2. 05B
    Credit stress Refinancing, covenants and impairments make the break visible.
  3. 06B
    Written-down capacity and cheaper access Somebody else has absorbed the original capital loss.

Kimi K3 makes the first step difficult to deny. Delivery can be dispersed or concentrated. The labour branch requires lower accepted-output costs. The credit branch requires a paid-workload, rental or debt-service failure. Only the post-write-down feedback requires competitive pass-through.

The complete reading edition

The trilogy, followed by its coda

Read the three essays in sequence, then continue through the corrected conclusion. The coda is deliberately separated from Part III and is not Part IV.

79 min read · 17,574 words

Copies all three essays and the coda as clean Markdown, including citations and the full proposed methodology.

Part I · Fact

The Second DeepSeek Moment

When the scarcity premium became contestable

Kimi K3 did not prove that open models are free to run, that benchmarks are the economy, or that every proprietary laboratory is finished. It proved something narrower and more dangerous: near-frontier capability now exists as a downloadable outside option. The best model may still command a premium. The old assumption that raw intelligence could sustain one while the surrounding capital stack reprices, refinances and amortises no longer looks safe.

By Ben Luong · 30 July 2026 · 10 min read · 2,195 words
Read Part I on its own →

The shrug

Eighteen months ago, a Chinese laboratory called DeepSeek released a model that helped wipe nearly $600 billion from NVIDIA’s market value in one day, then the largest one-day loss in US stock-market history. Commentators reached for Sputnik. Senators demanded investigations. The Western AI industry held a crisis meeting that has, in some sense, never adjourned.

On 16 July 2026, Moonshot AI released Kimi K3. On Artificial Analysis’s launch snapshot, it scored 57 on the Intelligence Index, behind only Claude Fable 5 and GPT-5.6 Sol. It placed third on GDPval-AA v2 at 1,668 Elo and second on AA-Briefcase at 1,547 Elo. Those are dated launch rankings, not permanent titles. The leaderboard changed again within days.

The economic fact is not the medal but the gap. An open-weight model had arrived within three Intelligence Index points of the proprietary leader. On 27 July, Moonshot released the full 2.8-trillion-parameter weights under the K3 licence.

The reaction was a shrug.

The second DeepSeek moment has arrived, and almost nobody is treating it as one. The shrug is not proof of its economic significance. It marks the change in prior that this essay tests through prices, routing and investment.

The first shock was treated as an anomaly. The second is being treated as weather.

What K3 actually is

Strip away the launch noise and the verified picture is both less magical and more economically consequential.

K3 is a 2.8-trillion-parameter mixture-of-experts model with native image input and a one-million-token context window. Moonshot says it still trails the strongest proprietary systems overall. Independent launch testing put it close enough to them to make the distinction between best and adequate commercially important.

Benchmarks do not measure the whole economy. GDPval-AA and AA-Briefcase measure defined distributions of agentic digital work under particular harnesses. They do not settle reliability in a bank, latency in a consumer product, integration into an old enterprise stack, compliance in a hospital, or the expected cost of a confident error. They establish broad near-frontier capability over the work they test. Commercial substitutability still depends on the cost of failure.

The pricing picture requires the same discipline. Moonshot’s first-party API costs $3 per million uncached input tokens and $15 per million output tokens. That is not the cheapest leading API. Claude Sonnet 5, for example, is on introductory pricing of $2 and $10 through 31 August before moving to the same $3 and $15 standard rate. A closed laboratory matching that rate card is evidence that price compression can occur without a balance-sheet event. The credit claim in Part III requires the separate condition that realised rental economics fall below the obligations written against them.

Artificial Analysis estimated K3 at $0.94 per task across its evaluation suite, close to GPT-5.6 Sol at $1.04 and roughly half Claude Opus 4.8 at $1.80. That is an evaluation-suite estimate, not the universal cost of completing commercial work. K3 also remains unusually verbose and expensive relative to open-weight peers. It used about 132 million output tokens across the Intelligence Index evaluation, materially above the reported median, although fewer than K2.6.

K3 is not revolutionary because Moonshot’s own API is the cheapest. It is not. The significance is that near-frontier capability has become downloadable, contestable and available to a serving market whose future price Moonshot no longer controls.

The outside option

The economically significant threshold was never technical supremacy. It is substitutability: the point at which the quality sacrificed by using the cheaper system is worth less than the money saved.

K3 does not need to beat the frontier. It needs to be close enough that a buyer can credibly threaten to move defined workloads after production evaluation. The launch evidence does not establish that threshold across most commercial work. It establishes enough capability to justify testing an expanding set of bounded tasks. Where those tasks clear the acceptance threshold, the open alternative can constrain the premium without taking every request.

Market power is strongest when the customer lacks a credible alternative. Contestable pricing begins before migration. A buyer can dual-source, route routine work elsewhere, demand a discount, reserve the proprietary model for the difficult tail, or build an internal alternative it may never fully deploy.

Contestability begins before commoditisation and can survive substantial switching costs, because negotiation responds to alternatives before deployment does.

A credible option can constrain the premium before most customers exercise it. It does not cap every price. It caps the premium on the tasks for which buyers could credibly switch.

Rents begin to die not only when customers leave, but when enough of them credibly could.

This is why the launch-week objection that K3 does not definitively beat the frontier misses the economic claim. A competitive market does not require the cheaper supplier to be the best. It requires the gap between best and adequate to fall below the price difference for enough work to discipline the seller.

That is not yet true everywhere. It does not need to be.

Downloadable does not mean free

The weights are available without an acquisition fee. They are not public-domain arithmetic floating above law, capital or engineering.

The K3 licence grants broad rights to use, copy, modify, distribute, fine-tune and deploy the model. It also contains conditions. A model-as-a-service operator whose group revenue exceeds $20 million over any consecutive twelve months must reach a separate agreement with Moonshot before commercial use. Very large consumer products face attribution requirements. Applicable law still applies.

Self-hosting changes the supply chain. It does not remove the deployment from law. The GDPR still follows personal-data processing, and the AI Act still allocates obligations to providers and deployers for covered uses.

Nor is a 2.8-trillion-parameter model cheap to serve merely because only part of it is active for each token. Moonshot recommends supernode deployments with at least 64 accelerators. The full expert pool must still be stored, distributed and routed. K3 is computationally sparse and infrastructurally enormous.

The precise claim is therefore not that release instantly makes global inference cheap, or that any competent host can serve K3 tomorrow at the cost of electricity. It is this:

The weights can be downloaded, modified and served by a broad global ecosystem, subject to Moonshot’s licence and the capital required to deploy them.

The licence is a friction. The deployment cost is a barrier. Neither restores exclusive ownership of the capability to Moonshot or to the Western frontier laboratories.

Why confirmation matters more than surprise

Markets and media price surprises. DeepSeek’s information content was that an open-weight laboratory could approach the frontier. That was unexpected and therefore dramatic.

K3’s information content is different. It says the first result was not safely dismissible as an accident. The open tier can return near the frontier on another release cycle, from another laboratory, while the proprietary leader continues the expensive work of moving the frontier again.

The leaderboard position will move. It already has. That is not an embarrassment to the thesis. It is the thesis. The question is not whether K3 remains third. It is whether an economically adequate open tier remains close enough, often enough, to discipline the closed tier.

A surprise might reverse. A repeated result changes the prior. The shrug is what it looks like when an extraordinary claim becomes an ordinary fact, and that normalisation is itself the discontinuity arriving: not with a crash, but with a calendar.

The hidden theorem is not that AI has become a commodity. Compute remains scarce. Reliable deployment remains difficult. Product moats, proprietary data, regulated distribution and customer relationships can remain valuable.

Nor can an outsider establish that any particular frontier model failed to recover its research, training, serving and organisational cost before its premium narrowed. Those fully allocated economics are not public. A short period of exclusivity can still support a profitable succession of products.

The observed claim is narrower:

The proprietary half-life of economically useful cognition is shortening.

The central financial hypothesis follows from that observation, but is not identical to it:

The quality-adjusted selling price of adequate cognition is falling faster than parts of the surrounding capital stack can reprice, refinance and amortise.

A temporary lead can still be worth a fortune, and frontier laboratories may remain profitable by repeatedly creating the next one. The exposed obligations are those whose repayment depends on the previous premium lasting longer than the market now allows.

The standalone model-layer fork

At the standalone model-API layer, the frontier laboratories now face a fork.

Hold price, and they invite customers to move routine volume, dual-source supply and use the open tier as leverage. The premium market may remain large, but it contracts towards the tasks where the reliability gap is still worth paying for.

Cut price towards cost, and they preserve volume by surrendering margin. Capital raises, cloud bundles and introductory rates can delay the accounting. They cannot turn a contested input back into a scarce one.

There is a third corporate move, but it is an exit from the layer: own the workflow, the customer, the distribution, the proprietary data, the device or the regulated relationship. The frontier companies are already moving into agents, coding environments, enterprise products and hardware. That is exactly the strategic behaviour this thesis predicts, although vertical integration alone does not prove why they are doing it.

Those moves may preserve company value. They do not by themselves preserve the old scarcity rent on raw intelligence.

The distinction matters. This is not an obituary for every frontier laboratory. It is an obituary for the assumption that the model layer itself will remain scarce merely because producing the next frontier model is expensive.

What has, and has not, been shown

K3 has not shown that open weights automatically produce low customer prices. Competitive hosting, adequate supply and efficient routing still have to transmit model contestability into the cost of delivered work. Infrastructure rent may survive even where model-layer rent does not.

It has not shown that all providers are interchangeable. A compatible interface can reduce switching costs while prompts, tool schemas, safety behaviour, latency, evaluation and compliance still require revalidation.

It has not shown that every economically valuable task is substitutable. The hardest, highest-risk work can sustain a premium long after routine work moves.

What it has shown is enough for Part I: near-frontier capability has entered the open-weight layer; the closed frontier now faces a credible outside option; and the standalone scarcity premium is now exposed to a visible and testable shortening of its half-life.

The first DeepSeek moment was an alarm: the frontier is reachable. The second is the sound of everyone sleeping through the confirmation.

The discontinuity was never going to announce itself twice. It arrives the second time as an ordinary Tuesday: a leaderboard update, a rate card, a licence and a download link.

A compact falsifiability note

This essay makes a narrow prediction, not a declaration that every model premium disappears. Freeze a representative basket of agentic digital tasks, the harnesses, the provider set, a minimum acceptance threshold and total cost per accepted output, then test it quarterly through 31 July 2027.

The claim takes serious damage if the best downloadable model remains more than 10 per cent behind the best closed model on accepted outcomes for two consecutive quarters. It also takes serious damage if, despite a downloadable model meeting the acceptance threshold, the leading closed providers sustain a quality-adjusted price premium above two times across the broad task basket for four consecutive quarters. A premium confined to the difficult tail does not refute the claim. It is what the claim predicts.

This test measures the shortening commercial premium. It does not establish whether a named laboratory recovered the fully allocated cost of a particular model generation.


Author’s note: this essay was drafted with the assistance of the model it describes, at commodity prices, during its own launch week. The reader may take that as evidence for the thesis, or as the one part of the argument that needed no essay at all.

Sources

  1. Artificial Analysis, “Kimi K3 achieves #3 in the Artificial Analysis Intelligence Index”, 17 July 2026. Launch-snapshot rankings, Elo scores, task-cost estimates and token use.
  2. Moonshot AI, “Kimi K3: Open Frontier Intelligence”. Architecture, context window, API pricing, scaling claims, limitations and deployment guidance.
  3. Moonshot AI, Kimi K3 model repository. Released weights and deployment materials.
  4. Moonshot AI, Kimi K3 Licence. Rights, model-as-a-service threshold and attribution conditions.
  5. Anthropic, “Introducing Claude Sonnet 5”. Introductory and standard API pricing.
  6. European Commission, “Data protection explained”. The GDPR’s technology-neutral application to personal-data processing.
  7. European Commission, “AI Act regulatory framework”. Provider and deployer obligations under the risk-based regime.
  8. Associated Press, “Nvidia posted another strong quarterly report. What to know, by the numbers”. The January 2025 DeepSeek market reaction.

Part II · Mechanism

The frontier labs are building a product Hetzner will sell like bandwidth

How contestability becomes unit cost dominance

A downloadable model is only an outside option. Standard interfaces, optimisation and routing can turn it into an industrial input through competitive hosting or integrated undercutting. The relevant price is the all-in cost of an accepted output.

By Ben Luong · 30 July 2026 · 18 min read · 3,928 words
Read Part II on its own →

Hetzner is experimenting with an LLM inference API. It offers one open-weight model, no billing, no service-level agreement and no promise of production availability.

The announcement matters because Hetzner can do this at all.

Hetzner is not a frontier AI laboratory. It spent nothing training the model. It did not assemble a research organisation or discover a new architecture. It took an available set of weights, ran them on its own infrastructure and exposed the result through an OpenAI-compatible endpoint.

This is not yet a commodity market. It is the mechanism by which one can form.

Hetzner shows the delivery mechanism. Kimi K3 shows the class of capability entering it. On 16 July 2026, Moonshot released a 2.8-trillion-parameter model that scored 57 on the Artificial Analysis Intelligence Index launch snapshot, within three index points of the proprietary leader. On 27 July, Moonshot released the full weights under the K3 licence.

Put the two together. A laboratory creates the capability. A different company can turn that capability into infrastructure.

The frontier laboratories are spending fortunes inventing a product that companies like Hetzner will eventually sell like bandwidth.

A standard endpoint lowers switching cost without making deployments identical

Hetzner’s experiment serves Qwen3.6-35B-A3B-FP8 through an OpenAI-compatible API. A developer points a standard client at Hetzner’s base URL, supplies a key and names the model.

At protocol level, switching may be little more than changing a base URL, key and model name. At production level, it is not. Prompts, tokenisation, tool schemas, structured outputs, safety behaviour, latency, rate limits, caching, evaluations, legal terms and compliance controls must be revalidated.

Standard interfaces do not make models identical. They do not make providers interchangeable in every production deployment. They reduce the cost of comparing and switching them. That is enough for contestability.

The distinction matters because the economic threshold is not zero switching cost. It is a credible outside option. A buyer that can test several providers against the same task basket, through substantially the same interface, has leverage even when migration still takes work. A routing layer can hold more than one approved endpoint and move eligible requests between them. The incumbent no longer negotiates against the cost of rebuilding the entire application.

That is a poor foundation for durable model-layer margins. Traditional software companies try to own something difficult to replace: a workflow, a proprietary dataset, a network of users, a regulated relationship or the customer itself. A hosted open-weight model owns much less. Competing providers can serve the same weights, related weights or an adequate substitute through a familiar protocol. Competition then moves towards the mundane economics of infrastructure: hardware, electricity, cooling, utilisation, reliability, location and markup.

That is precisely the market Hetzner understands. It does not need to become a glamorous AI company. It can remain what it already is, a ruthless operator of cheap European infrastructure.

The Lidl of machine intelligence.

K3 is downloadable, but it is not cheap or simple to deploy

K3 matters because its capability is close enough to the proprietary frontier to discipline it. It does not matter because Moonshot’s own API is the cheapest. It is not.

On Artificial Analysis’s dated launch evaluation, K3 placed third on the Intelligence Index, third on GDPval-AA v2 and second on AA-Briefcase. The same evaluation estimated an average cost of $0.94 per task, close to GPT-5.6 Sol’s $1.04 and below Claude Opus 4.8’s $1.80. That is an evaluation-suite estimate, not the universal price of commercial work. K3 also used about 132 million output tokens across the suite. It remained unusually verbose and expensive relative to open-weight peers.

The release does not abolish deployment economics. It exposes them.

K3 has 2.8 trillion total parameters and activates 104 billion per token. Its mixture-of-experts architecture selects 16 of 896 experts for each token. Sparse activation weakens the relationship between total parameter count and compute per generated token. It does not abolish the cost of storing, distributing and routing across the full model.

K3 is computationally sparse but infrastructurally enormous.

Moonshot’s published deployment guidance centres on supernode configurations with at least 64 accelerators. The released model already uses quantisation-aware training with MXFP4 weights and MXFP8 activations. It supports low, high and max reasoning effort, with max as the default. Third-party quantisations have already appeared, but further compression necessarily trades deployment cost against quality, latency and behaviour.

The weights are also licensed, not placed in the public domain. The K3 licence grants broad rights to use, modify, deploy and distribute the model. It requires a model-as-a-service operator whose group revenue exceeds $20 million over any consecutive twelve months to reach a separate agreement with Moonshot before commercial use. Very large consumer products face attribution requirements. Internal use is treated more permissively.

The correct claim is therefore narrower than “anyone can serve it”. The weights can be downloaded, modified and served by a broad global ecosystem, subject to Moonshot’s licence and the capital required to deploy them.

The licence is a friction. The hardware is a barrier. Neither restores exclusive control to the laboratory that trained the model.

The important event is not that Hetzner can host K3 tomorrow morning. It is that the intelligence K3 represents has entered a cost-reduction process that Moonshot no longer controls.

Industrialisation attacks every component of the accepted-output cost

K3 has already arrived aggressively quantised. The next reductions do not depend on repeating the same trick.

Serving engineers improve kernels, interconnect utilisation, batching, caching, speculative decoding and request scheduling. Researchers distil task-relevant behaviour into smaller descendants. Hardware improves. Routers reserve the full model for the difficult tail. Each optimisation attacks a different component of the accepted-output cost.

That last phrase is the unit that matters.

Raw token price is not the test. Benchmark rank is not the test. The unit is an accepted output. Unit cost dominance begins when the all-in machine cost of producing that accepted output falls below the fully loaded human cost of producing the same thing.

The machine cost includes inference, integration, orchestration, verification, expected error and rework, compliance and liability, and residual human supervision. The comparison must hold at the required quality, latency and risk. If a human still has to reconstruct the entire output, absorb frequent failures or carry prohibitive liability, unit cost dominance has not occurred for that unit.

This definition includes the strongest objection rather than denying it. Human verification is real. Integration is expensive. Errors matter. Regulation can impose cost. The thesis applies only where the entire deployed system, including those burdens, is cheaper than human-only production.

Once that threshold is crossed, the economic value of the output need not fall. A useful report remains useful. What competition pulls down is the market price of producing it and the rent earned by whoever once supplied it under conditions of scarcity.

The market price of performing the task is pulled towards the all-in unit cost of the cheapest adequate production system.

Routing gives the volume to adequacy

A single model does not need to be excellent at everything. A production system needs to select a model that is adequate for the bounded request in front of it.

In a plausible routed architecture, a cheap model handles the large routine majority of calls while difficult cases escalate to a premium system. The exact share is workload-dependent. The economic point is that the expensive model need not receive the volume.

This changes what progress means. The largest model receives the prestige. The deployment layer receives the requests. A frontier system may remain essential for novel research, difficult coding, exceptional judgement or the highest-risk tail. None of that protects its price across routine work that a cheaper model can complete to the required standard.

K3 itself may be too large for the boring majority of workloads. Its descendants will not need to preserve every capability. A distilled or specialised model can be worse in the abstract and still better for the purchaser if it clears the acceptance threshold at lower all-in cost.

This is why adequacy threatens model-layer rent more than technical defeat does. The closed frontier can remain ahead. Its premium survives only where the gap changes the accepted output.

The model identity then recedes. Applications retain a catalogue of evaluated options. A router assigns work by price, latency, privacy, capacity and measured quality. A provider can be replaced without the customer changing the workflow it sells. Intelligence becomes a workload placed into a competitive infrastructure market.

Field update, 30 July 2026: OpenAI prices the routing mechanism

Twenty-one days after GPT-5.6 entered broad public availability, OpenAI cut Luna’s API price by 80 per cent and Terra’s by 20 per cent, while leaving the frontier Sol price unchanged. Luna moved from $1 per million input tokens and $6 per million output tokens to $0.20 and $1.20. Terra moved from $2.50 and $15 to $2 and $12. Sol remained at $5 and $30.

The asymmetry matters more than the headline. The apex retained its premium. The broad, high-volume tier was radically repriced.

OpenAI’s own guidance now begins with the required outcome and quality standard. It tells businesses to evaluate where additional intelligence materially improves the result and where lower-cost processing preserves the same quality. Its coding example uses Sol to resolve uncertainty and define the plan, then Luna to implement well-specified changes, run tests and evaluate the result.

The expensive model resolves uncertainty. The cheap model implements, tests and evaluates. The account stays with OpenAI; the routine volume leaves Sol.

That is accepted-output economics even though OpenAI does not use the term. The purchaser is told to buy the cheapest system that clears the quality threshold at each stage of the workflow, not the most capable model for every token.

OpenAI says efficiency made the reductions possible. Better routing, serving software and context management produce more useful work from the same compute. It says Sol-assisted kernel work reduced end-to-end serving cost by 20 per cent, while related experiments improved token-generation efficiency by more than 15 per cent. Axios reports that cheaper Chinese open-weight models have also increased pressure on OpenAI and Anthropic to justify their higher costs. The announcement does not identify how much of the pass-through came from efficiency and how much from competition.

Nothing here proves that K3 caused a particular percentage point of the cut. It shows what the credible outside option mechanism can look like. The incumbent keeps the customer while repricing its high-volume tier as if defection were possible.

OpenAI also claims that Luna now delivers performance comparable to models that were frontier-class a year ago at roughly six cents on the dollar per task and nearly nine times the speed. That is a vendor comparison, not an independent or controlled time series. Taken only as an illustration, six cents represents just over four cost halvings in one year, or an economic half-life of about 90 days.

This does not establish the credit event in Part III. Efficiency may protect OpenAI’s margins, and lower prices may stimulate enough demand to increase total compute use. It shows one transmission route directly: OpenAI has repriced its adequate tier and is explicitly positioning Sol’s premium around the stages where additional intelligence changes the accepted result.

Two routes to commodity-priced cognition

The OpenAI cut also exposes a fork inside the delivery mechanism.

The first route is competitive commoditisation. Open weights are served by many infrastructure providers. Standard interfaces and routing layers expose price differences. Hosts compete on electricity, hardware, utilisation, reliability and markup. Efficiency gains and written-down capacity can reach customers through defection between providers.

This is the Hetzner route.

The second route is concentrated commoditisation. A vertically integrated laboratory or hyperscaler uses scale, engineering efficiency, cloud bundling or strategic pricing to undercut independent providers. Routine cognition becomes extremely cheap while the serving market remains concentrated.

This is the Luna route.

Both routes can destroy durable scarcity pricing for routine cognition. Only the first implies a dispersed hosting market. Neither, by itself, proves that a lower acquisition cost for distressed capacity will reach customers after a credit event.

That makes the OpenAI cut double-edged. It strengthens the claim that adequate cognition is becoming cheap. It weakens any assumption that independent hosts must be the organisations making it cheap. It may even intensify pressure on merchant GPU providers that cannot match an integrated provider’s rate, leaving the surviving serving layer more concentrated after the distress.

The title of this essay remains a prediction, not an established industrial structure. A company like Hetzner may eventually sell intelligence like bandwidth. A frontier laboratory may instead preserve the account and sell its own volume tier at bandwidth prices. The labour mechanism can operate through either route. Part III’s fire-sale mechanism requires the additional condition of competitive pass-through after restructuring.

Open weights do not guarantee cheap inference

The transmission remains conditional.

Competition pushes price towards underlying cost only when enough providers can serve adequate models, licensing does not block entry, and effective supply outruns paid workload demand. If infrastructure consolidates into a few hosts, if power and accelerator scarcity dominate the stack, or if demand absorbs every efficiency gain, model-layer rent can die while infrastructure rent survives.

Contestability is the necessary condition. Competitive pass-through is the transmission.

This is also why the argument does not depend on Hetzner personally winning. Hetzner’s public GPU estate may remain too small for frontier-scale serving. Its experiment may end. If Google, Amazon or a specialist inference company owns the cheap layer instead, the model premium can still disappear. The surplus has merely moved to a different landlord.

That outcome would matter. Concentrated infrastructure could keep customer prices well above physical cost. Scarce power, land, accelerators, high-bandwidth interconnect and reliable capacity can support durable rents even while raw intelligence loses its scarcity premium.

The claim is not that every layer becomes perfectly competitive. It is that a model owner cannot rely on permanent scarcity pricing once an adequate alternative can be deployed, evaluated and routed by other firms.

K3 does not prove that pass-through will occur. Hetzner does not prove that supply will outrun demand. Together they make the industrial pathway visible.

Tasks disappear before jobs do

Unit cost dominance operates on outputs. Organisations employ people in jobs. The bridge between them is task decomposition.

Suppose a worker performs twenty recurring tasks. No single model can perform the whole job. One drafts the reports. One reconciles the spreadsheet. One answers routine email. One inspects screenshots. An orchestration layer moves information between them. A smaller number of people supervise exceptions and carry formal responsibility.

The organisation does not need a perfect digital employee. It needs enough accepted output to require fewer producers.

The first quarter looks like assistance. The second looks like productivity. The headcount decision arrives later.

There is no dramatic moment when a machine becomes a perfect employee. There is simply a budget meeting.

The mistake is to wait for one model to perform every element of an occupation. The economically relevant threshold arrives earlier, when a collection of systems produces enough of the output that the organisation needs fewer people. Whole jobs survive on an organisation chart while the volume of human production inside them contracts.

This is how displacement can begin without a mass redundancy announcement. Hiring slows. Junior roles are not reopened. Contractors replace permanent staff. A five-person team supervises work that once required twenty. Headline unemployment can remain calm while the labour market stops absorbing the next cohort.

Sorites prevents a usable boundary

No single step looks like replacement. A tool drafts one report, then checks one spreadsheet, then handles one client queue. Headcount falls at the next budget round.

Assistance and replacement differ by degree, not by a boundary anyone can verify. That Sorites ambiguity makes restraint unenforceable. Every firm can call its own adoption assistance while treating everyone else’s restraint as an opportunity. Nobody can say when defection began, so everybody defects before anybody agrees that replacement has begun.

The distinction remains meaningful at the endpoints. A person writing unaided is producing. A largely automated workflow with one exception handler has replaced most production labour. The problem lies between them. There is no operational line at which assistance becomes replacement, even though the cumulative change is unmistakable.

A rule can require human oversight. It must then define how much attention, at what frequency and with what authority. A firm can satisfy the label through sampling, escalation or formal approval while continuing to reduce the labour content of each accepted output. The human remains visible. Productive necessity recedes underneath.

Sorites is not a claim that regulation is impossible. It is a claim about a particular regulatory defence. Any regime that must identify the precise moment at which assistance becomes replacement is trying to govern a boundary the workflow does not contain.

Sorites is not the net-displacement theorem either. It explains why substitution can proceed without a defensible moment at which assistance becomes replacement. Whether new work absorbs the displaced wage income is a separate empirical question.

New tasks do not necessarily arrive already automated. Digitally expressed tasks increasingly arrive contestable from birth. The same general systems already operate the browser, document, spreadsheet, inbox, dashboard and code editor through which a new workflow is likely to be performed. Human labour may still win many of those contests. It no longer receives an exclusive window by default while specialised machinery is designed around the new task.

The non-absorption claim fails if new human-complementary work restores displaced wage income at comparable scale and speed. It is supported if output and machine-completed work rise while exposed hiring, junior intake, human hours and labour income do not.

The Multiplayer Prisoner’s Dilemma makes adoption compulsory at the system level

The same game repeats at every level.

Laboratories cut prices because rivals might. Hosts optimise because other hosts will. Firms automate because competitors are reducing labour cost. Workers adopt the tools because refusing makes them individually less employable, even though universal adoption reduces the number of workers required.

Each move is locally rational and collectively accelerative.

A laboratory that preserves margin while a near-frontier rival cuts price loses volume. A host that declines to optimise leaves utilisation and customers to another host. A firm that keeps expensive human production for substitutable output while competitors cross unit cost dominance loses on price or margin. A worker who refuses augmentation is compared with one who can supervise more output.

No actor needs to desire the aggregate result. Each needs only to avoid being the actor who pays for restraint while others defect.

Unit cost dominance supplies the payoff. Sorites prevents a usable boundary. The Multiplayer Prisoner’s Dilemma makes adoption compulsory at the system level wherever the cost advantage is material.

That does not mean every individual firm must automate. Some can sustain a premium around human service, trust, regulation, craft or luxury positioning. Those firms become residual niches inside the new cost structure. They do not remove the competitive pressure on the substitutable volume.

This is the missing link between cheaper inference and labour displacement. Capability does not enter the economy because every institution has agreed on its social value. It enters because the cost advantage can be captured privately while the displacement cost is shared.

Regulation loses its single upstream choke point

When the most capable systems are controlled by a few laboratories, regulation has a convenient object. Governments can impose reporting duties, mandate evaluations, restrict particular exports and negotiate with named executives.

Downloadable weights fragment that object across the model developer, host, application provider, router, deployer and regulated institution. The obligations do not vanish. They multiply.

A European company can run K3 on European infrastructure without sending its prompts to Moonshot or a United States cloud. That reduces supplier dependence and some cross-border-transfer exposure. It does not remove the deployment from law.

The GDPR is technology-neutral and still follows the processing of personal data. The AI Act allocates obligations to providers and deployers for specified uses. Sector rules still apply. Moonshot’s licence still governs the weights.

Self-hosting changes the supply chain. It does not make the arithmetic legally invisible.

Regulation can govern employers, hospitals, banks, insurers, public bodies and harmful uses. It can impose liability and documentation duties. What it cannot easily do is restore one upstream choke point once the capability is distributed across models, hosts and jurisdictions.

The delivery layer turns a model event into an economic event

Hetzner may withdraw the endpoint. K3 may remain expensive to self-host. A stronger proprietary model may appear next month. None of those possibilities answers the mechanism.

The mechanism does not require Hetzner to serve K3. It requires near-frontier capability to become an input available to providers that compete on deployment. It does not require providers to be identical. It requires switching and comparison to become cheap enough to discipline price. It does not require zero human involvement. It requires the all-in cost of an accepted output to fall below the fully loaded cost of equivalent human production.

When those conditions hold, the model stops behaving like a premium product and starts behaving like an infrastructure component. The frontier laboratory can preserve company value by owning the workflow, customer, device, proprietary data or regulated relationship. It cannot assume that raw intelligence will continue to carry the old scarcity premium.

For businesses, that is the attraction. Capability becomes cheaper to acquire and easier to embed.

For labour, that is the mechanism.

Workers will keep being told the current model is not quite good enough.

Then, one ordinary quarter, it will be.

What would falsify this mechanism

Fix a basket of production tasks, quality thresholds and providers before measuring it. This account weakens if standard interfaces do not reduce comparison and migration costs, if neither competitive hosting nor integrated undercutting lowers the all-in cost per accepted output, or if verification, error, compliance, liability and supervision keep that cost above equivalent human production. Persistent quality-adjusted model premiums would refute the pricing claim.

The labour transmission weakens if machine-completed accepted outputs rise without human hours, junior intake or the wage share deteriorating in exposed sectors. The stronger non-absorption claim fails if new human-complementary work restores displaced wage income at comparable scale and speed.

Sources

  1. Jonas Scholz, Hetzner Inference: First Look, including the experiment’s single model, OpenAI-compatible endpoint, lack of billing, SLA and production guarantee, and July 23 test.
  2. Artificial Analysis, Kimi K3 achieves number three in the Artificial Analysis Intelligence Index, 17 July 2026 launch snapshot, benchmark positions, measured task cost, token usage and API pricing.
  3. Moonshot AI, Kimi K3 Tech Blog, model architecture, launch positioning, multimodality and context window.
  4. Moonshot AI, Kimi K3 model card, released weights, activated parameters, quantisation format, deployment engines and reasoning-effort controls.
  5. Moonshot AI, Kimi K3 Licence, commercial model-as-a-service threshold, attribution requirements and internal-use exception.
  6. European Commission, Data protection explained, including the GDPR’s technology-neutral application to personal-data processing.
  7. European Commission, AI Act regulatory framework, provider and deployer obligations under the risk-based regime.
  8. OpenAI, Advancing the price-performance frontier with GPT-5.6, 30 July 2026. Price reductions, workflow-routing guidance, cost-per-task comparison and serving-efficiency claims.
  9. OpenAI, GPT-5.6: Frontier intelligence that scales with your ambition, 9 July 2026. Original public-availability date and launch API prices for Sol, Terra and Luna.
  10. Axios, OpenAI cuts GPT-5.6 prices, 30 July 2026. Price-sensitive customer context and reported pressure from cheaper Chinese open-weight models.

Part III · Consequence

The Margin Call

When scarcity debt meets contested pricing

Open weights do not make compute free, and contested model pricing does not automatically crash GPU finance. The narrower claim is that the quality-adjusted selling price of adequate cognition may fall faster than parts of the surrounding capital stack can reprice, refinance and amortise. If effective supply outruns paid workload demand, the loss need not end the automation cycle. With competitive pass-through, it can finance the next phase of it.

By Ben Luong · 30 July 2026 · 27 min read · 6,021 words
Read Part III on its own →

Who lent against the rent?

Part I argued the fact. Near-frontier capability has entered the open-weight layer and given buyers a credible outside option. Part II argued the mechanism. Standard interfaces, delivery optimisation and routing reduce the cost of comparing providers and push the delivered price of intelligence towards contested infrastructure economics. That delivery can be dispersed across hosts or concentrated inside an integrated laboratory.

Providers are not identical. Production deployments still require prompts, tool schemas, safety behaviour, latency, evaluation and compliance to be revalidated. Compute, power and reliable serving remain scarce. The claim is not that every friction has disappeared. It is that the assumption of exclusive control over frontier-like capability has become contestable.

This essay asks the question neither of the first two asked.

Who lent money against the assumption that none of this would happen?

The build-out was priced before the scarcity premium became visibly contestable

The AI capital expenditure programme now runs into the high hundreds of billions. A Reuters dashboard reports more than $800 billion of estimated hyperscaler capital expenditure for 2026. Reuters’ 29 July bond-market analysis reports that by 7 July Amazon, Alphabet, Meta and Oracle had issued about $194 billion of bonds, 79 per cent more than they issued during all of 2025. It also documents widening borrowing spreads and cites a Goldman Sachs estimate that debt issuance will equal about one-third of hyperscaler capital spending in 2026.

That establishes material exposure. It does not establish the crash.

It is not one capital structure, and the precise claim matters. Cash-rich hyperscalers have funded much of the build from operating cash flow and can absorb disappointing returns far longer than anyone downstream, although the scale of current commitments is drawing them further into debt markets too. Utilities and data-centre landlords earn contracted returns. Nvidia can remain the largest winner during the build-out because open weights expand demand for accelerators. The sequence can reverse if algorithmic efficiency, overcapacity, slower frontier training and a secondary market in impaired silicon weaken demand for new high-margin hardware. Nvidia can win the boom and still be repriced by the overbuild it supplied.

The vulnerable layer sits at the leveraged edge: neoclouds, GPU lessors and project-finance vehicles whose debt was sized to specific compute rental rates, specific utilisation assumptions and specific collateral values for hardware whose economic life is measured in years rather than decades. Those numbers were set when frontier-like intelligence looked scarce for long enough to repay the capital built around it.

Not all leveraged-edge exposure is merchant exposure. In March 2026, CoreWeave closed an $8.5 billion non-recourse delayed-draw facility secured by high-performance-computing infrastructure and an associated customer contract, with a March 2032 maturity. CoreWeave said its equity and debt financing commitments secured over the preceding twelve months totalled approximately $28 billion. Its first-quarter filing says borrowing is tied to depreciable equipment cost, projected debt-service coverage and project-level conditions. Contract backing may protect facilities like this from short-term spot-price compression. The exposed cohort is therefore narrower: structures in which tenant credit, take-or-pay durability, utilisation, collateral value or refinancing assumptions fail together.

This is not a claim that every frontier model fails to recover its training cost. Model-specific, fully allocated research, serving and organisational economics are not public, and a laboratory can remain profitable through a succession of short-lived leads. The central financial hypothesis is narrower:

The quality-adjusted selling price of adequate cognition is falling faster than parts of the surrounding capital stack can reprice, refinance and amortise.

Part I argued that the customer now has somewhere else to go. For buyers capable of evaluating and deploying alternatives, procurement now occurs in the shadow of downloadable weights that sat within three Intelligence Index points of the closed leader at K3’s launch, servable by competing hosts with the infrastructure and licensing rights to do so. The outside option constrains the premium whether or not anyone exercises it, and whatever Moonshot’s own list price does this quarter. Contested model pricing can flow downstream into the rental rates, utilisation and collateral values the leveraged edge borrowed against.

That transmission is not automatic, and naming the condition is the honest move. Open weights can destroy model-layer rent while increasing demand for GPU-hours. More hosts, more deployments and more inference can absorb more compute. The credit event requires effective compute supply, installed capacity multiplied by hardware and algorithmic efficiency, to outrun paid workload demand. Alternatively, competition among hosts must compress realised rental economics below debt service.

That is the additional prediction this essay makes beyond Parts I and II. If Jevons demand keeps merchant rental rates and utilisation at their underwriting levels, the model rent can die without the leveraged edge breaking. The efficiency curve, the overbuild and the routing layer make that failure plausible. They do not show that the coin has already landed. This essay dates the prediction rather than presenting it as a result.

So the arithmetic is now visible. Debt sized to scarcity pricing, serviced by contested pricing. The gap between those two numbers is where the bubble lives. Kimi did not create the gap. The release of its full weights on 27 July made the gap visible enough to date, not automatic enough to declare.

Volume is not solvency

The standard defence is demand. As machine cognition becomes cheaper, firms will use more of it. But the economically relevant variable is not token volume. It is paid accepted output. Better systems can complete more useful work with fewer tokens, while weakening household demand can reduce the total quantity of final output the economy is willing to buy.

Jevons pressure expands the uses to which cheap cognition is put. It does not guarantee that paid workload demand will outrun simultaneous gains in hardware efficiency, algorithmic efficiency and installed capacity. Nor does it guarantee the rental rate, utilisation or collateral value assumed by the capital stack.

Exploding demand does not guarantee scarcity margins. Electricity, bandwidth and freight carry enormous volumes at infrastructure returns.

Be precise about what volume can and cannot do, because the distinction is where the margin call actually lives. Volume can service debt at utility margins. Grids, pipelines and telecoms are debt-financed on exactly that basis. What volume at contested margins cannot do is preserve software valuations, venture returns, residual values for silicon ageing faster than its loans amortise, or refinancing terms written when capacity looked permanently scarce.

The first loss therefore lands on equity: valuations are written down to the extent that they priced reproducible cognition as a durable monopoly. It becomes a credit event wherever fixed obligations were sized to the old rental rate, the old utilisation assumption or the old collateral value.

Merchant GPU capacity. Leveraged neoclouds. Project vehicles dependent on loss-making tenants. Hardware-backed lending whose collateral ages faster than the loan amortises.

The hyperscalers can absorb disappointing returns for far longer. The leveraged intermediaries may not survive the refinancing.

The frontier labs can see this, which is why they are racing to become product companies: agents, coding environments, enterprise platforms and devices. Part II called that race consistent evidence for the thesis. They are trying to own the system around the intelligence before the intelligence becomes cheap. Some will succeed as product companies, or as integrated providers able to undercut independent hosts while preserving value elsewhere in the stack. That possibility weakens the case for a dispersed hosting market. It does not restore scarcity pricing to routine cognition.

Demand can fall while substitution accelerates

Cognition is an intermediate input, so its absolute demand ultimately depends on final demand. If wage loss weakens mass consumption, the total volume of paid workloads may grow more slowly than the Jevons defence assumes, or contract outright. That strengthens the credit risk. It does not necessarily rescue labour.

Three quantities must be kept separate: total final output, absolute machine workload and the machine share of production.

Let Y be final output, c the cognitive input required per unit of output, and s the machine share of that cognitive input.

Machine-handled cognition: M = Y × c × s

Human-handled cognition: H = Y × c × (1 – s)

Suppose final output falls from 100 to 90 and cognitive intensity remains one. If the machine share rises from 20 per cent to 60 per cent, machine-handled cognition rises from 20 to 54 while human-handled cognition falls from 80 to 36.

Final output has fallen by 10 per cent. Machine work has risen by 170 per cent. Human work has fallen by 55 per cent.

Even where machine work does not rise in absolute terms, human work can fall much faster than total output. A firm serving a smaller market can still serve it with one supervisor and a machine system rather than five employees. The displacement variable is not total compute consumption. It is the machine share of the cognition used to produce each remaining accepted output.

This creates a feedback loop. Substitution weakens wage income. Weaker wage income weakens mass demand. Weaker demand intensifies the pressure to reduce unit cost. Lower machine prices make further substitution available. The same contraction can therefore break infrastructure debt through insufficient absolute workload while breaking labour through a rising machine share.

The compute owner needs sufficient absolute revenue to service fixed obligations. The worker needs human labour to retain sufficient relative share of production. Those are not the same requirement.

Earlier automation waves offer a precedent, not a forecast. Jaimovich and Siu found that routine employment losses in the United States were heavily concentrated in downturns and that jobless recoveries were largely accounted for by routine occupations that disappeared. Robert Allen’s account of Engels’ Pause describes a long interval in which productivity and output rose much faster than workers’ living standards before the gains were more broadly shared. Neither result proves that the current transition will follow the same path. They show that restructuring can concentrate in contractions and that eventual recovery is not an answer to a generation living through the delay.

The counterforce is real. Weak firms may lack the cash, organisational capacity or confidence to implement new systems. Consolidation may prevent lower infrastructure costs from reaching them. The acceleration claim fails if those frictions suppress adoption more than falling machine costs and unit-cost pressure encourage it.

During the transition, final demand can still come from wages that remain, household savings and credit, government deficits and transfers, capital income, investment spending, exports and lower prices that stretch residual income. None is automatically a permanent replacement for wage-mediated mass demand. Without redistribution or a new claim on automated output, the process can culminate in underused capacity, repeated demand shortfalls and concentrated production rather than universal abundance. The absence of a stable replacement demand engine is not a premise required to start the transition. It is part of the successor-system problem the trilogy leaves open.

The crash can speed up the displacement

Here is the part the public conversation has exactly backwards.

The assumed sequence runs like this. The bubble bursts. The AI story dies. The pressure on jobs eases. Workers watching the capital expenditure numbers with dread are quietly hoping for the crash, on the theory that a crashed industry stops hiring machines.

The labour sequence may already have begun, quietly. Junior hiring can thin, graduate absorption can weaken and incumbent teams can be asked to do more without producing a clean break in headline unemployment. The No-Scream Principle, the claim that compositional damage appears before aggregate measures scream, predicts exactly this sequence.

Credit behaves differently. It has payment dates. The prediction here is not that no worker is harmed before a lender is. It is that financing stress at the leveraged edge becomes the first hard, dated and institutionally undeniable break, before AI displacement appears in headline aggregate unemployment.

Under that sequence, the leveraged edge breaks. Merchant clusters are refinanced, sold or absorbed at steep discounts. The durable sites may hold their value. The acquisition basis of the compute inside them does not. Inference can then clear below the return originally required to build the capacity because somebody else has already eaten the capital loss.

One counterforce deserves more weight, because the acceleration claim depends on beating it. The most plausible buyers of distressed capacity may be hyperscalers that already own adjacent infrastructure. If a few incumbents absorb the written-down assets and preserve the prevailing price umbrella, the lower acquisition basis becomes landlord surplus rather than cheaper intelligence.

Historic cost is sunk. A host with cheap assets will charge the prevailing rate unless idle capacity and competitive defection force it to cut. Standard interfaces and routing make price differences visible and switching easier, but they do not create competitors. The discount reaches customers only if capacity remains contestable after restructuring, through independent hosts, forced disposals, open access or another mechanism that keeps survivors competing to fill the assets.

This creates two separable claims. A credit event can occur without pass-through. The acceleration claim requires both the write-down and competitive pass-through. If consolidation removes the second, the crash does not subsidise diffusion even though lenders may still take the loss.

If competition survives, consultancies, insurers, law firms, software companies and government departments that found frontier API pricing hard to justify get a different calculation. The marginal adopter, the firm that ran the pilot and shelved it on cost grounds, reruns the numbers and may adopt. Substitution can accelerate into the downturn, not out of it, because downturns are precisely when firms cut headcount and hunt for cheaper inputs. Under those conditions, the crash has delivered a cheaper cognitive input.

The liquidation sale reaches individuals too, and here it intensifies a dynamic already running. Freelancers are underbidding agencies today, at today’s prices. Agencies are underbidding in-house teams. In-house seniors are running departments without juniors. Juniors are using the same tools to look employable a little longer. Consultants are selling transformation plans that automate the clients who commissioned them.

Each actor adopts to survive the round, and widespread adoption removes the reason many of them were needed. That is the Multiplayer Prisoner’s Dilemma operating at the level of individual careers, and it does not wait for any crash.

What the crash changes is the floor, but only where competing hosts pass the economics of written-down clusters through an API. That does not require every freelancer to run a 2.8-trillion-parameter model on a second-hand gaming card. It requires temporary access to infrastructure whose original owner has already eaten the capital loss to become a cheaper weapon in the pit.

The game is already being played. A pass-through crash makes each move cheaper.

There is precedent, and the structure is familiar. Railway investors were ruined in the 1840s. The railways remained, and their fire-sale capacity industrialised freight. Telecom investors were ruined in 2001. The fibre remained, and its stranded capacity carried the internet economy for twenty years at prices the original investors never planned.

One honest difference matters. Rail and fibre were durable assets. This generation of accelerators is not. What survives an AI crash is the durable shell around the silicon: the grid connections, substations, buildings, cooling, power contracts and permitted sites. The shell is the reason the capacity survives the crash. It is not necessarily the discounted asset. The subsidy comes from the excess compute, the impaired silicon and the capital claims written off above them, while the permitted shell removes years from the next owner’s deployment timetable.

The crash also has opposite effects on development and diffusion. Capital destruction may thin the trillion-parameter training runs and slow the frontier. It can accelerate diffusion of capability already created because the systems already shipped are sufficient for a large class of commercial substitution and their delivery cost has fallen.

The objection that bankrupt AI firms cannot advance AI is true and beside the point.

The displacement does not need the next model. It needs the last one, cheaper.

The financial parasite dies. The automation organism spreads through its corpse.

A crash meeting these conditions may be reported as the end of the AI era while functioning as a subsidy to the replacement of labour. It becomes a liquidation sale on human obsolescence only where the discounts are passed through to employers able to use them.

The state is standing on four rotting floorboards

The credit event and the fiscal crisis run on different clocks. The first is a late-2020s refinancing prediction. The second compounds into the 2030s as compositional labour losses accumulate. Payroll and consumption receipts can weaken at the margin before headline unemployment breaks, but a state-wide cash-flow crisis is not the same-quarter consequence of the margin call.

Now put the government in the picture, because the fiscal exposure may be worse than the financial one and almost nobody has added it up.

The state taxes wages. AI compresses wages, first through quiet non-absorption at the entry level, then through mid-career restructuring. One of the largest and most automatically collected tax bases in every developed economy erodes at the source: broad, domestic, visible and withheld at payroll.

The state taxes corporate profits. The rents that survive do not necessarily vanish. They migrate out of the contested model layer and into clouds, chips, energy, land and distribution. They accumulate in fewer hands, become more internationally mobile, and sit with owners better placed to bargain with any individual government. A base that was broad and captive becomes narrow and negotiable.

The state taxes consumption. Displaced households consume less. VAT receipts follow wages down.

The state borrows against future growth. Here cheap AI may perversely fulfil the GDP forecast while breaking the fiscal one. Output can rise even as wage income, payroll receipts and mass consumption fall, if the gain accrues to a small and internationally mobile capital base. Productivity is not tax capacity. The state can receive the growth it was promised and still lose the revenue it borrowed against.

These are four gross pressures, not an accounting identity. Cheap AI could also raise taxable profits, lower prices, expand consumption and reduce some public-service costs. The claim is about timing and incidence. Payroll losses arrive automatically and domestically. Replacement rents are narrower, more mobile and slower to capture. Transfer demands arrive before legislation, treaties and enforcement can redirect the surplus.

The state can receive the productivity gain and still suffer the cash-flow crisis.

The state does not lose because output disappears. It loses because the tax base moves faster than the tax code. Payroll taxes itself in real time. The surplus replacing payroll must be chased through treaties, legislation, lobbying and elections, and it can relocate while the bill is in committee. Meanwhile the transfer demands arrive when employment income weakens. The fiscal crisis is a timing mismatch as much as a revenue loss. The state must fund the transition before it has captured anything winning from it.

The surplus may be internationally mobile, but the scarce physical complements are territorially fixed.

That is why the tax system will crawl towards what cannot relocate: land, power generation, grid connections, water rights, data-centre footprint, resource extraction and territorial permission itself. States retain other instruments, including wealth taxes, financial taxes, procurement and sovereign investment. They will use them. The most dependable base is the physical one because it cannot move to another jurisdiction after the finance minister speaks.

The crawl towards it will be presented as strategy when it is elimination.

Note what that implies for bargaining position. A state whose dependable base is land, power and grid access is a state negotiating with the owners of land, power and grid access. It arrives at that table late, stretched, and needing them more than they need it.

The window for permits-for-equity is open and closing

The timing of the constructive move matters more than its design.

I set out the mechanism in an earlier essay, where the constraint analysis identified permits-for-equity as the cleanest pre-commitment instrument under this framework. It is not literally the only instrument. Governments can invest directly, attach royalties, use development banks, retain public land, acquire rescue equity or tax gross flows.

The distinctive advantage is ex ante conversion of public scarcity into a permanent claim before capital is sunk.

The build-out still needs things only states can grant. Land. Planning approval. Grid connections. Water. Transmission. Political protection. Permission is not the state’s only leverage, but it is its cleanest. Before construction, a permanent public claim can attach without a bailout, a retrospective tax fight or a threat of closure. It becomes a transparent condition of access that the developer can price before committing capital.

That leverage is uneven. A jurisdiction demanding equity can lose a sufficiently mobile project to one that does not. Permits-for-equity works unilaterally only where permission attaches to a binding, location-specific bottleneck: a constrained grid connection, publicly financed transmission, public land, sovereign procurement, guaranteed demand or rescue support. Where sites and jurisdictions are genuine substitutes, the claim requires regional coordination or it will be competed away.

The state’s leverage does not disappear after construction. Expansions, power, water, tax treatment and regulation continue. It becomes costlier and more politically adversarial to exercise once capital is sunk. The cleanest leverage exists on the way in and degrades the moment the permit is irreversible.

The exchange rate matters. Trade permission for equity. Trade it for permanent public claims on the infrastructure.

Do not trade permanent access to land, power and grid capacity for headline employment commitments. Construction employment is temporary, and permanent operating payroll is small relative to the capital value and public scarcity being allocated. Data centres still employ people in operations, networking, security, power engineering and their supply chains. The point is not that they employ nobody. It is that automation will place continuing downward pressure on a payroll already small beside the concession.

The state that swaps permanent physical access for temporary payroll promises will spend the next decade taxing the asset it forgot to keep a share of.

The window is narrow because a crash compresses it. Distressed compute changes hands quickly, the surviving owners consolidate, and owners with sunk permits need less from the state. The hopeful reply is that a crash hands the state fresh leverage as workout referee. It does.

Workout leverage can even buy equity, but only if the state supplies rescue capital, guarantees or coercive law at the moment its own balance sheet is weakest. Permit leverage acquires the claim before the crisis. Workout leverage requires the state to purchase it during one.

The bolder reply is that the state should wait and buy everything cheaply at the bottom. It fails on four counts.

The fire sale is in the layer the trilogy says is losing value: model businesses, leveraged shells and depreciating GPUs. The assets worth owning, the best grid connections, land, water and generation, retain or gain scarcity value through the crash because the crash is the value migrating into them.

The buyer is fiscally and politically weakest on the day of the sale, with tax receipts eroding and transfer demands rising. Weaker states may also face higher borrowing costs. Stronger currency issuers can fund the purchase cheaply but can still move too slowly to clear the assets that matter.

Private capital clears a distressed auction in a weekend. Nationalisation needs a statute, a valuation and a court. The consolidation can finish before the enabling act passes.

Even a successful purchase buys melting machinery downstream of a chip supply chain the state still does not control. A few currency issuers with cheap energy will run the bottom-buying play well. That is part of the sorting between states, not a plan available to the rest.

Buying the bottom is the strategy of a state that missed the window, executed with money it no longer has, for assets that stopped mattering.

The leverage that converts permission into permanent claims exists cleanly now, while the concrete is still being poured. Afterwards, the same claim must be bought, taxed, litigated or rescued.

What would falsify the argument?

This essay makes four separable claims, and the most exposed one is pass-through.

Do not treat AI demand as one variable. The empirical record must separate three groups of measures:

Infrastructure solvency: realised price per accepted output, paid workload revenue, cluster utilisation, debt service, refinancing conditions and collateral values.

Technological diffusion: deployed production workflows, machine-completed accepted outputs, machine share of task execution and cost per accepted output.

Labour displacement: human hours per accepted output, junior and graduate hiring, headcount relative to output, wage income in exposed occupations and the reinstatement of equivalent wage income in new occupations.

Paid workloads can disappoint while the machine share rises. That combination is not contradictory. It is the route by which infrastructure investors and workers can lose in the same contraction.

The pricing claim takes serious damage if neither independently hosted open-weight systems nor integrated volume tiers approach the best closed systems on accepted commercial outcomes at a materially lower quality-adjusted price, or if closed providers sustain their broad premium through 31 December 2027.

The credit claim fails if model-layer pricing compresses but realised merchant rental economics remain above debt service after the relevant borrowers have tested the refinancing market through 31 December 2029. The sequencing claim fails if AI-exposed employment produces a clear break in headline aggregate unemployment before a composite credit turn at the leveraged edge.

The acceleration claim fails in either of two cleaner ways. If asset values fall but accepted-output costs do not, demand or consolidation has stopped the write-down from reaching customers. If accepted-output costs fall but production adoption and machine share do not accelerate, cheaper supply has not produced the predicted diffusion. The labour claim weakens if human-complementary work restores displaced hours and wage income at comparable scale and speed.

If neither the headline labour break nor the credit composite appears after the relevant maturity window, the result does not merely reorder the sequence. It falsifies the timetable and weakens the wider claim that contestability is transmitting into the economy with the speed and force argued here.

Growing laboratory revenue would not falsify it. Commodity markets have large sellers. Revenue is not rent.

Nor does OpenAI’s 30 July price cut establish post-crash pass-through. It demonstrates delivered-price compression and workflow routing. The separate acceleration claim begins only when a distressed acquisition basis is transmitted into lower accepted-output cost after a credit event.

Methodological appendix: a proposed empirical test

The design below is a research proposal, not a completed preregistration. The 48-task basket, borrower cohort and enterprise panel have not been published, recruited or frozen. The protocol becomes real only when those materials, the acceptance rubric, baseline model IDs and available data are disclosed. Until then, it is the programme a funded research organisation would need to execute.

Any implementation should publish its baseline date in advance and use trailing ninety-day medians so that one promotional rate, one distressed borrower or one noisy employment release cannot decide the result.

Open the proposed protocol details

Pricing. Publish and freeze a basket of 48 commercial tasks across coding, research, document production, spreadsheet work, customer operations and regulated professional workflows. An independent panel must select the basket, or sample it from a published external task frame, under inclusion and exclusion rules disclosed before evaluation. Freeze prompts, tool access, latency limits and the acceptance rubric. At the first run, record the exact production model IDs for the highest-scoring closed API from Anthropic, Google and OpenAI, plus the two strongest independently hosted open-weight models whose licences permit the test. An output is accepted only when two of three blinded domain assessors judge that it needs no substantive repair. Cost per accepted output includes inference, retries, orchestration and human verification. Human comparison uses task-specific fully loaded labour costs taken from a preregistered external wage-and-overhead source. Results are also reported under common-rate sensitivity scenarios so that no single labour-cost assumption determines the conclusion. Before evaluation, the preregistered rubric must set an absolute acceptance floor for each task category, allowing risk thresholds to differ across coding, customer support and regulated professional work. An open option qualifies only if it clears the preregistered absolute acceptance floor for its task category and reaches at least 90 per cent of the best closed model’s acceptance rate. The open-weight contestability claim fails if no open release qualifies across the next two major open-weight generations. The broader pricing claim also fails if neither an open option nor an integrated volume tier produces material quality-adjusted compression and the broad premium falls by less than 10 per cent from its baseline level in four consecutive quarterly runs through 31 December 2027. A fall of at least 25 per cent from baseline, sustained for two quarters, counts as compression on schedule.

Credit. Publish a cohort of AI-infrastructure borrowers with debt equal to at least 30 per cent of invested capital and a refinancing or maturity date between 1 January 2027 and 31 December 2029. Track paid workload revenue and realised quality-adjusted price per accepted output alongside financing measures. A credit turn requires at least three of the following across two unrelated borrowers within two consecutive quarters: a 25 per cent year-on-year fall in quality-adjusted merchant GPU rental rates, a 15 percentage-point fall in disclosed cluster utilisation, a 250 basis-point widening in refinancing spread over the matched sovereign curve, a 20 percentage-point increase in collateral haircuts, a covenant waiver or take-or-pay renegotiation, or an impairment or distressed sale at least 25 per cent below carrying value. The credit claim fails if pricing compression occurs and no composite turn appears by 31 December 2029, provided at least half of the cohort’s principal has passed a contractual refinancing or maturity date. If less than half has tested the market, the clock runs until it has.

Labour and sequence. Pre-register 20 highly exposed occupations and 20 controls matched on baseline wage, education, sector and prior cyclical sensitivity in both the United Kingdom and United States. Track employment, advertised vacancies, the share of vacancies asking for no more than two years’ experience, median advertised real wage, human hours per accepted output and headcount relative to output. Track whether new occupations restore equivalent wage income, not merely whether new job titles appear. A quiet compositional turn occurs when the exposed basket underperforms its controls by at least 10 per cent in vacancies and five percentage points in entry-level share for two consecutive quarters. That may occur before credit and does not falsify the essay. A headline labour break requires unemployment to rise at least one percentage point year on year in either country while employment in the exposed basket underperforms its controls by at least three per cent. The sequencing claim fails if that headline break precedes the credit composite. Credit first supports the prediction that the margin call becomes the first damage nobody can shrug at.

Acceleration. The first credit composite sets the event date. A funded research partner would need to recruit a vendor-neutral panel of 500 enterprises, stratified by country, sector and firm size. It would record quarterly the share of defined production workflows using AI, machine-completed accepted outputs, machine share of task execution and human hours per accepted output, excluding pilots and individual ad hoc use, with the questions and workflow definitions held fixed. The write-down becomes an adoption subsidy only if quality-adjusted cost per accepted output falls at least 20 per cent beyond its pre-event four-quarter trend within the next four quarters, and the panel’s median production-workflow share grows at least 50 per cent faster than its pre-event four-quarter rate for two consecutive quarters. The joint labour-acceleration claim additionally requires human hours per accepted output, exposed hiring or labour income to deteriorate faster after the event than before it. If asset values fall but accepted-output prices do not pass through, consolidation has interrupted the mechanism. If prices pass through but production adoption and machine share do not accelerate, cheaper supply has not produced the predicted diffusion. Either result falsifies the acceleration claim.

Series close: the redistribution, named

You cannot base the successor system solely on frontier-lab monopoly windfalls that competition is already eroding. A viable fiscal claim must attach to volume, ownership and scarce physical complements, not merely to the accounting profits of whoever currently has the smartest model.

Put the three essays together and the whole event compresses to one movement.

Near-frontier open weights create a credible outside option. Standard interfaces and routing transmit it through one of two industrial structures: competitive hosting or integrated undercutting. Both can make routine cognition cheap. Only the first depends on continued competition among multiple serving providers.

From there the causal model forks. Where either delivery route lowers the all-in cost of an accepted output below equivalent human production, unit cost dominance spreads through firms and careers. Separately, where effective supply outruns paid workload demand or realised rental economics fall below debt service, obligations underwritten against the previous scarcity inherit the loss. If written-down capacity then reaches customers as cheaper access, the credit branch feeds back into the labour branch and accelerates diffusion.

The demand loop can tighten both branches at once. Substitution can weaken wage income and final demand, disappointing the absolute paid workload required to service infrastructure debt while unit-cost pressure raises the machine share of the output that remains. Machine workload and machine share are different quantities. They can move in opposite directions.

Every arrow is conditional. The falsification summary and proposed research design identify where each branch can fail.

Where those conditions hold, labour loses bargaining power because reproducible cognition stops being scarce. Frontier laboratories need not become unprofitable; some can prosper through successive frontier leads, products and integration. The routine-tier scarcity rent can still compress underneath them. Leveraged investors lose where obligations were sized to a premium that reprices faster than the loan. The state loses where its broadest and most automatic tax bases move faster than its tax code.

The owners of whichever physical and legal complements remain binding, including leading-edge fabs, constrained grid connections, scarce power, permitted sites, ports and enforcement, inherit the position because the software layer prices itself towards utility economics and the machines still need atoms.

That is the redistribution actually underway. Cognitive capital to physical capital.

The software aristocracy is discovering it was renting the throne from the owners of atoms.

Part I opened with a shrug. The second DeepSeek moment arrived and almost nobody noticed because the quiet parts of a discontinuity are easy to shrug at. Hiring that never happens makes no sound. Entry-level roles evaporate without a press release.

Credit is different.

Credit has covenants, coupons and dates.

The first part of this sequence that nobody manages to shrug at will be the margin call.


Sources

  1. Reuters, “AI boom sparks rally, frenzy and fear”. Estimated 2026 hyperscaler capital expenditure.
  2. Reuters, “Hyperscaler debt binge pushes yields up as investor demand cools”. Bond issuance through 7 July, the comparison with 2025 issuance, debt-financing share and borrowing-spread context.
  3. Artificial Analysis, “Kimi K3 achieves number three in the Artificial Analysis Intelligence Index”. Launch-snapshot capability, task-cost and token-use evidence.
  4. Moonshot AI, Kimi K3 model repository. Full-weight release, model architecture and deployment materials.
  5. Moonshot AI, Kimi K3 licence. Use, modification, distribution and model-as-a-service conditions.
  6. Sliplane, “Hetzner Inference: First Look”. Evidence for conventional infrastructure providers exposing open-weight inference through a standard interface, with launch-stage limitations.
  7. Ben Luong, “The Margin Call”. Original LinkedIn publication.
  8. Nir Jaimovich and Henry E. Siu, “Job Polarization and Jobless Recoveries”. Evidence that routine employment losses in earlier United States automation waves were concentrated in downturns. This is a historical precedent, not a claim that the AI transition must repeat it.
  9. Robert C. Allen, “Engels’ Pause: Technical Change, Capital Accumulation, and Inequality in the British Industrial Revolution”. Historical evidence on the long lag between productivity growth and broadly shared gains.
  10. OpenAI, “GPT-5.6: Frontier intelligence that scales with your ambition”. Vendor evidence that tool-heavy tasks can complete with fewer tokens and round trips, illustrating why token volume is not the economic unit.
  11. CoreWeave, “CoreWeave Closes Landmark $8.5 Billion Financing Facility”. Facility size, non-recourse structure, infrastructure and customer-contract security, March 2032 maturity and twelve-month financing-commitment total.
  12. CoreWeave, first-quarter 2026 Form 10-Q. Debt-sizing conditions, collateral, limited guarantees and the distinction between total facility capacity and drawn debt.

The Shortening Half-Life of Intelligence · Coda

Rent Is Not Conserved

The Bottleneck Thesis: where the surviving rent goes when the model can generate the mechanism

The trilogy ended with rent migrating from cognitive capital to physical capital. That was directionally right and incomplete. Rent is not a conserved quantity. It forms around binding, excludable complements that cannot be cheaply reproduced or routed around, and where several bind at once, their owners divide it according to bargaining power and control of access. When no complement binds, it is not inherited by a hidden final landlord. Competition can pass it through in lower prices and higher quality, or dissipate it through overinvestment and capital loss. Much of what passes through can arrive as consumer surplus, which no tax system was designed to see.

By Ben Luong · 30 July 2026 · 25 min read · 5,430 words
Read the coda on its own →

The claim, before the rhetoric

This coda does not claim that every software incumbent collapses, that application revenue disappears, or that physical assets earn nothing. It makes a narrower, conditional claim. Models make cognition reproducible, and increasingly make application functionality reproducible too. The rent that survives tends to migrate towards whatever accumulated state, default coordination, enforceable right, loss-bearing capital or constrained physical capacity remains binding and excludable, and it tends to migrate towards the complements whose effective substitutability falls most slowly. Sometimes the destination is atoms. Sometimes it is permission or a balance sheet. And sometimes no complement holds, competition passes the margin through or burns it in the fight, and the rent dissolves. The market does not owe the technology industry a replacement margin.

Here, rent means the return above the competitive level made possible by constrained substitution. It does not mean all revenue, accounting profit or payment received by an owner. A business can survive, grow and earn an ordinary return after the particular rent discussed here has disappeared. Accepted, throughout, means accepted by the party whose recognition makes the output actionable: the customer, the auditor, the regulator or the court, depending on the task.

The escape hatch is temporary

The obvious response to model commoditisation is to move up the stack. If the underlying intelligence becomes cheap, the frontier laboratory becomes an application company. It wraps the model in a workflow, adds customer data, builds an interface, secures distribution and charges for the finished product rather than the cognition underneath it.

The trilogy treated that move as an exit from the contested layer. This coda treats it as a stay of execution.

If a sufficiently capable model can write software, operate interfaces, call tools, inspect its own output, repair errors and maintain a working system, then the application is another bundle of cognitive tasks. The laboratory can move from model to application. The model can follow it there.

A business asks for a lead-management system. The model generates the schema, interface, permissions, reports, email automation, analytics and integrations. It tests the result, deploys it and repairs defects. The cost of producing the functionality falls. The application layer has acquired its own proprietary half-life.

That does not mean every company abandons its existing software next quarter. It means the functionality stops being the scarce part of the product. The invoice can survive long after the feature that originally justified it has become reproducible, because the vendor is still charging for everything required to make that feature usable inside a live institution: access to authoritative data, integration with existing systems, implementation, security approval, support, continuity, contractual responsibility and recourse when the output is wrong.

Feature rent dies before the invoice does.

The application is not destroyed in one event. It is hollowed out from the interface inward. The visible functionality becomes cheap first. The surrounding institution decays more slowly.

Generated mechanisms and accumulated state

The distinction that governs everything below is not software versus atoms. Code is a stock. Data can often be copied. A reputation graph can be exported. The useful line runs between generated mechanisms and accumulated, recognised state.

Models are increasingly able to generate mechanisms: code, interfaces, workflows, controls, tests, integrations, monitoring, transaction logic, compliance documentation, even the software of a marketplace. But an economic system contains facts that are not recreated by reproducing its mechanism. These customers already use it. These sellers are already verified. This transaction history actually occurred. This record is recognised as authoritative. This organisation is licensed to perform the activity. This contract is enforceable. This balance sheet stands behind the promise.

A model can generate a marketplace. It cannot, by generating the code, generate the fact that everyone is already trading there. It can create a customer database. It cannot unilaterally decide which copy the organisation, the auditor and the regulator will treat as the real one. It can draft an insurance policy. It cannot produce the capital that makes the claimant whole.

The model can reproduce the mechanism. It cannot instantly reproduce the state on which the mechanism operates.

That is where rent survives longer. Not permanently, and not by law of nature. Longer.

The system of record becomes a database with contracts

The visible application is a collection of forms, dashboards, workflows and reports. Those are precisely the elements most exposed to generation. A model can build a new interface over the same business data, customise it per employee and replace a rigid menu with a conversational agent.

The incumbent’s durable asset was never the screen. It was the organisation’s decision that a particular body of data would count. When two records disagree, which one is authoritative? Which system controls access? Which audit trail does the regulator accept? Which database triggers the invoice, the shipment, the payroll or the legal notice? A system of record is an institutional declaration about which state is recognised when records conflict.

AI compresses the moat around that declaration. It can map schemas, clean data, rebuild connectors, identify dependencies, generate tests and rewrite downstream workflows, turning a multi-year consultancy migration into a much shorter project. But copying the data does not complete the migration. Duplicated data creates two records, not a new authority. The replacement becomes authoritative only when permissions, contracts, reporting rules, audit procedures and organisational behaviour recognise it as such.

The model can clone the database. It cannot unilaterally declare the cutover complete.

The old application therefore becomes something thinner and stranger: a database with contracts, permissions and historical recognition attached. Its feature rent falls. Its authority rent survives on an institutional clock rather than a code-generation clock. Portability rules, open standards and cheaper migration keep shortening even that clock. They do not reset it to zero.

Marketplace code is not marketplace power

The same distinction applies to marketplaces, and here the decay is visibly layered.

A model can generate the software for a marketplace: listings, search, ranking, matching, payments, messaging and reviews. That does not generate marketplace power. A marketplace earns rent because multiple groups have coordinated around the same place. Buyers go there because sellers are there. The software supports the equilibrium; it is not the equilibrium. The scarce state is the verified participants, the transaction history, the accumulated reputation, the payment relationships, the dispute rules and the shared expectation about where demand will appear.

AI attacks that rent one layer at a time, from the outside in.

Discovery rent comes under pressure first. A buyer’s agent can search several platforms, compare offers and identify the cheapest acceptable supplier without accepting any single platform’s ranking. The first take rate under pressure is the charge for helping buyers find sellers.

Identity and reputation last longer. The agent still needs to know whether the seller exists and whether the history can be trusted, and platforms can keep those records proprietary. A model can analyse the evidence it can access. It cannot infer its way around evidence deliberately withheld. Portability standards and regulation can weaken this moat; better inference alone cannot.

Settlement and recourse tend to last longer still. Who processes the payment, absorbs the fraud, decides the dispute and compensates the buyer when the item never arrives? An agent can automate the work of all four. The enforceable promise still belongs to an organisation with contractual authority and assets behind it.

Discovery is usually the first marketplace rent exposed. Standing behind the transaction is usually among the last.

The counter-evidence should be named now rather than discovered later. Several large marketplaces have so far preserved or increased effective monetisation. eBay’s reported take rate rose from 13.77 per cent in 2024 to 13.94 per cent in 2025, with advertising and shipping contributing to revenue growth. Booking Holdings’ revenue rose from 14.3 to 14.5 per cent of gross bookings, partly through payment facilitation. The evidence is not uniform, and some headline seller fees have fallen under competitive pressure, but marketplace tolls have so far proved capable of migrating rather than simply disappearing. The discovery claim is therefore a dated prediction, not a description of visible data. The test is compositional rather than the total take rate. Compression should appear first as falling discovery-related revenue per agent-mediated transaction, while settlement, identity, fulfilment and dispute-resolution charges prove more durable. A rising total take rate would not by itself falsify the claim if the toll had visibly migrated towards those slower layers. The claim is weakened if discovery monetisation per comparable transaction remains stable or rises after buyer-side agents mediate a material share of demand and can genuinely compare suppliers across platforms.

The agent becomes the new gatekeeper

If a personal or enterprise agent can search every marketplace and route each request to the cheapest acceptable source, the marketplaces become backend liquidity pools. The customer does not care which supplier completed the task. The agent controls demand, and that creates a new concentration point: whoever owns the identity the agent uses, the permission to spend, the connection to corporate systems and the definition of acceptable quality can tax the layers underneath.

This is the cleanest explanation of the current platform land grab. The laboratories and platforms racing to make their assistants the default surface, and to surround them with app directories, permissions and certification, are not merely diversifying away from the model layer. They are racing to own the layer that disintermediates the coordinators. The observable evidence is already public: model companies now operate directories through which third-party apps are published inside their assistants, and the major platforms subject AI apps and agents to marketplace certification requirements. Discovery, distribution and governance around generated functionality are being enclosed while the functionality itself commoditises.

Whether the agent gate holds depends on which kind of state it accumulates. A user’s preferences, permissions and learned definition of acceptable are accumulated state, but they are individual state: exportable in principle, potentially reconstructible far faster than a two-sided network or certification regime, and an obvious target for portability regulation. The directory, the certification regime and the default placement are institutional state, and they sit on the slower clocks this coda describes. The agent gate is therefore two moats of different depths. Per-user memory is a shallow one. Distribution governance is not, and it may prove the most durable position the laboratories have yet acquired. Portability rules, multi-agent use and open protocols can still move the bottleneck past it. The coda cannot claim that they must.

If they do, the bottleneck moves again.

Trust performs four different economic functions

It is tempting to end the recursion with trust: the last moat is that people trust the incumbent. That is too vague to price. Trust contains four economic functions with four different half-lives.

Functional trust asks whether the system performs the task correctly. This is the most exposed form. Better models, evaluations, redundancy and formal constraints make functional confidence cheaper to produce. One model can test another. Functional trust is becoming a technical task.

Evidentiary trust asks whether the organisation can prove what happened: reproduce the decision, show which data was used, demonstrate that controls ran. Logging, monitoring and compliance documentation are highly automatable. But the production of evidence is automatable in a way its authority is not. A generated audit trail still depends on trusted provenance, recognised standards, tamper-resistant records and an institution willing to accept it. Evidentiary trust therefore sits across the boundary: the artefact is a mechanism; its admissibility is accumulated and recognised state.

Coordination trust asks whether the relevant participants recognise the same identities, histories, permissions and records. A newly generated reputation system has no reputation history. A copied identity network is not automatically recognised by banks, governments or trading partners. This is accumulated social state. It can be transferred or standardised. It cannot be produced by writing better code.

Financial and legal trust asks the residual question: if the system fails, who pays? Who owes the customer a duty, answers the regulator, funds the remediation, and possesses assets against which the claimant can recover? The surrounding work can all be automated. A model can perform the tests, draft the contract, prepare the audit and recommend the settlement. It cannot make the injured party whole unless a recognised entity has capital behind the promise.

The model can write the insurance policy. It cannot fund the claim.

Trust is not the last moat. The last institutional moat is credible recourse.

Recourse becomes the product

Organisations do not merely buy software. They buy someone to be responsible for it.

That sentence sounds like an objection to application commoditisation. It is the next stage of it. When the functionality becomes cheap, the commercial product shifts towards continuity, responsibility and loss absorption. A company pays a major vendor because the vendor promises the service will remain available, security will be maintained, regulatory commitments will be honoured, support will exist, and there is an entity to sue, fine or compel. The model can automate the tasks involved in meeting those commitments. It cannot eliminate the commitments themselves.

The application business therefore separates rather than vanishes. Some software companies survive as risk-bearing institutions with generated products inside them, their remaining rent resting on contracts, licences, recognised authority and a sufficiently large balance sheet. Others discover that the customer never needed the company once the functionality could be generated and the risk could be insured elsewhere. The split runs along liability rather than sentiment. Where losses are large, regulated or borne by third parties, the recourse premium holds. Across the long tail of low-stakes software, buyers will take the cheap option and carry the risk themselves, and nobody will be paid to stand behind it.

Functionality becomes cheap. Recourse becomes the product.

The clocks set the drift

Here is the operative part of the thesis, the part that makes it a prediction rather than a description.

Different complements become contestable on different clocks, and the resulting rents are repriced on a second set of clocks. Generatable functionality can become technically reproducible on model-release cycles, measured in months. Its commercial price may not reset until contracts renew, a credible substitute is deployable or migration becomes affordable. Integration and deployment become contestable on implementation cycles. Systems of record move on migration, procurement and audit cycles, measured in years. Networks and marketplaces move when enough participants can coordinate on an alternative. Legal obligations move through legislation, regulation and judicial decision. Loss-bearing capacity moves when insurers, capital markets or governments learn to price a new class of risk, often among the slowest clocks of all.

Capability half-life is therefore not identical to rent half-life. The first determines when a complement can be reproduced. The second determines when its owner can no longer charge as though it cannot. The technical moat can vanish while the supplier keeps collecting through contracts, inertia, bundling or the absence of a deployable substitute. That is what feature rent dying before the invoice does looks like from the inside.

The rents do not decay on the same clock. Within an exposed sector, value tends to migrate towards the complements whose effective substitutability falls most slowly. But the order is conditional. Regulation, vertical integration, liability, capital structure and sector-specific institutions can bundle layers, skip them or reverse the apparent sequence. A medical application can jump almost directly from generated functionality to regulatory permission and liability. A consumer design tool can jump from functionality to distribution. Insurance capital can reprice overnight while a statute stands unchanged for decades.

The prediction is therefore not that every industry follows one ladder. It is that, as faster-reproduced complements lose pricing power, a growing share of the remaining margin will be justified by the slower-reproduced complements that still bind in that particular market.

That migration tendency is the falsifiable core of the coda. It predicts where price justification moves inside any given business, and it explains why the transition looks contradictory from outside. A software incumbent can remain highly profitable after its technical moat has disappeared, because it is living on the slower half-life of installed state, contractual position or customer coordination. A marketplace can lose discovery while preserving its take rate through settlement. A frontier laboratory can lose model-layer scarcity while gaining default-agent distribution. Revenue persistence is not proof that the original moat remains. It may be the harvest of a slower-decaying one.

The same tendency explains why policy is perpetually late. By the time regulators respond to model concentration, value has moved into distribution. By the time they regulate application stores, it has moved into identity and payment. By the time they mandate data portability, it has moved into liability, licensing or sovereign permission.

The policy response is usually aimed at last year’s bottleneck.

The lineage of this argument is worth naming, because it strengthens it. Teece showed in 1986 that when imitation is easy, the profits from an innovation can flow to the owners of complementary assets rather than to the innovator, with the allocation depending on who can control and contract over those assets. The Bottleneck Thesis is that result applied to a technology that can increasingly generate the complements as well, run at a speed Teece’s cases never approached, with one added branch he did not need: sometimes every complement becomes contestable, and the rent has nowhere left to sit.

Atoms are not the automatic winner

The trilogy’s final movement towards physical capital was correct in one sense. No model runs without machinery, electricity, fabrication, materials and logistics. Digital abundance rests on physical systems.

But necessity does not guarantee rent. Electricity is indispensable, and most electricity producers hold no exceptional pricing power. Servers are necessary, and ordinary server manufacturing is not a monopoly. Land is physical, and most of it is economically irrelevant to high-density computation.

Necessity determines what must be purchased. Scarcity and exclusion determine who captures rent.

Physical owners capture exceptional returns only where the complement remains hard to expand, substitute or route around: leading-edge fabrication, advanced packaging, constrained grid connections, scarce low-cost power, permitted high-density sites, critical minerals, restricted routes. Even those scarcities decay. AI improves chip design, grid operation, materials discovery and construction. Governments subsidise capacity. Demand can disappoint. A previously scarce physical input can become an oversupplied commodity, and the bottleneck does not stop moving when it reaches atoms.

In many cases the enduring constraint is not the machinery but permission to use it: the grid connection, the export licence, the environmental permit, the right to serve a national market, the legal authority to process the data. Sovereign permission can outlast the scarcity of the physical asset it governs.

The endpoint is not physical capital. It is the least contestable claim still required to complete the outcome.

The sovereign gate

At the bottom of most private rights sits the state. It recognises property, enforces contracts, defines legal identity, licenses regulated activities, decides which records and signatures count, and provides the courts through which recourse becomes more than a promise.

A model can generate a contract; the state determines whether it is enforceable. A protocol can issue a token; the state determines whether the token grants a recognised claim over a house or a bank account. A marketplace can build a private reputation system; the state can still remove the licence under which its participants operate.

The sovereign can create, allocate or capture rent wherever permission itself is binding and enforceable. Whether the public receives that rent depends on how the right is issued, whether through auctions, taxes, royalties, equity or revenue-linked claims, and on whether jurisdictional competition forces the state to surrender the value to the private operator. Its position can decay more slowly because the sovereign often defines the legal clock, although constitutional limits, supranational rules, political legitimacy and competition among jurisdictions constrain that control.

Even sovereign rent is constrained. States compete for investment and tax base, and a mobile activity can relocate. The state captures durable rent only where it controls something firms cannot cheaply replace: access to a large market, a constrained grid, public land, legal recognition, procurement demand, rescue finance, physical security. The state is powerful where its gate is real.

This revises Part III’s advice without withdrawing it. The case for permits-for-equity never actually depended on atoms holding their scarcity value through the transition. It depends on permission being the slowest-decaying complement the state owns, and on the state designing the capture mechanism rather than assuming that ownership of the gate automatically puts the rent into the treasury. The instruction stands, with sharper wording: the state should convert its permission into durable, contingent claims while the permission remains binding, and it should take those claims in the operating entity and the revenue stream rather than in the depreciating machinery, precisely because the coda cannot guarantee which physical assets stay scarce. Trade the gate for equity, royalties or revenue-linked participation. Do not trade it solely for an interest in machinery the recursion may commoditise.

Sometimes nobody captures the rent

The debate about AI rents usually assumes the surplus must have an owner. If the laboratory loses the margin, the application vendor gains it. If the application commoditises, the marketplace gains it. If agents weaken marketplaces, the platform gains it. If the cloud commoditises, the power company gains it.

That reasoning treats rent as a conserved quantity. It is not.

Suppose models are competitive, applications are easily generated, data is portable, marketplaces interoperate, identity is standardised, risk can be insured competitively and infrastructure is abundant. No layer holds a binding, excludable bottleneck. The previous margin is not inherited by a hidden final landlord. Some of the former producer rent is passed through as lower prices, higher quality or greater output. Some is dissipated through duplicated investment, strategic subsidy and capital write-offs. The essential claim is not that customers inherit every euro of the previous margin. It is that no successor producer is economically guaranteed to inherit it.

The market does not appoint a new rentier merely because the old one disappeared.

This is the most economically attractive outcome. It is also the politically dangerous one.

The fiscal paradox of consumer surplus

Consumer surplus is a real gain. If a service that cost €1,000 can be produced for €10, the customer is better off, and real living standards can rise even as nominal expenditure falls. Standard GDP is not a measure of total economic welfare and does not record consumer surplus directly; Brynjolfsson and colleagues built GDP-B to capture welfare contributions from new and free digital goods that conventional national accounts can miss. The fiscal implication drawn here is mine, not theirs: welfare that does not appear as a wage, profit or taxable transaction is not automatically converted into state revenue.

But consumer surplus is not a payroll. It is not a corporate profit, a dividend or a rent payment. The state cannot send a tax demand to the invisible difference between what a consumer would have paid and what the service now costs.

Some of the gain returns to the tax base indirectly, as released income is spent elsewhere. That transmission is not automatic. If lower prices arrive alongside displaced wages, compressed producer margins, reduced billable hours and falling nominal business expenditure, then real abundance can coexist with fiscal weakness. Society can afford more in material terms while the institutions responsible for funding the transition collect less in nominal claims.

The productivity gain exists. The fiscal claim on it does not arise automatically.

This gives the trilogy’s fiscal crisis a second, harder-to-tax form. In the first, rent migrates to concentrated physical, platform or sovereign bottlenecks, and the political fight is over who owns and taxes them. In the second, producer rent is competed away faster than the wage and tax systems can be rebuilt, and there is no equivalent corporate winner from whom to recover the cost of displaced wages and institutional adaptation. The output becomes cheap. The transition remains expensive.

There is a third form between them, and it may be the most common. Where several weak complements bind at once, the surviving rent fragments into thin, contested margins spread across many holders, none dominant. That can broaden the ordinary tax base and reduce the political power of any single rentier, but it also leaves less concentrated surplus for the state to capture through targeted bargaining. Whether fragmentation is fiscally better or worse depends on mobility, reporting, profit allocation and the tax instruments available. The treasury no longer has one obvious counterparty. It is negotiating with a diaspora.

Cheap does not mean equitably distributed either. A generated service can be nearly free at the point of production while remaining inaccessible to anyone lacking a device, connectivity, recognised identity, payment credentials or the ability to bear residual risk. The productive mechanism becomes abundant; the rights required to use it remain scarce. The model can generate a legal argument, but only a recognised court can enforce it. It can design a treatment plan, but only a licensed institution may deliver it. The scarcity shifts from production to participation, and the important question stops being who owns the model.

Who is permitted to turn the model’s output into an enforceable claim on the world?

That owner may capture more rent than the model provider ever did.

The Bottleneck Thesis

The theory can now be stated plainly.

An accepted economic outcome requires a collection of complements: cognition, software, data, integration, distribution, identity, permission, settlement, recourse and physical capacity. Most can be supplied competitively. Actors controlling complements with low effective substitutability can charge above competitive cost, and that premium is the rent. Where several such complements are required at once, their owners bargain over the resulting surplus rather than one layer automatically inheriting it all.

Models reduce the scarcity of cognition. Code generation reduces the scarcity of application functionality. Agents reduce discovery costs. Interoperability reduces integration lock-in. Portable identity weakens network capture. Competitive insurance weakens liability rent. Infrastructure build-out weakens physical scarcity. Regulation can weaken or create any of these positions.

At each stage, one of four things happens. The previous owner retains the bottleneck. The bottleneck moves to another layer. The rent fragments among several owners who each hold part of the gate. Or no binding bottleneck remains and the rent is competed away or dissipated.

Not all difficulty earns rent; the asset must be binding. Not all scarcity earns private rent; the asset must be excludable. Not all control is durable; the customer must be unable to route around it.

Economic rent forms around binding, excludable complements that cannot be cheaply reproduced or routed around. Where several complements bind, the rent is divided according to control and bargaining power. As models make cognition, software and adjacent mechanisms more reproducible, surviving rent tends to migrate towards the accumulated state, coordination, enforceable rights, loss-bearing capital, physical capacity or sovereign permission that remains less substitutable. Competition can also fragment or destroy the rent without appointing a successor owner.

What would falsify the coda

This is a taxonomy with a migration tendency, so most tests are directional rather than dated. Freeze the predicates before reading the data.

Feature-level differentiation should become harder to monetise as generation quality rises, with vendors increasingly justifying price through integration, authoritative data, compliance and recourse rather than unique functionality. Systems of record should retain pricing power longer than feature vendors and lose it as migration and institutional cutover become cheaper. Marketplace discovery fees should compress before settlement, identity and dispute-resolution fees, and total take rates should fall towards the layer still bearing loss as reputation and payment state become portable. Model providers that fail to convert capability into distribution, state, permission or risk-bearing should compress faster than platforms holding those complements. Physical infrastructure should earn exceptional rent only while capacity or permission remains constrained, reverting to utility economics where supply expands and access is contestable. And where no layer retains a defensible bottleneck, prices should fall without an equivalent producer profit appearing elsewhere.

One near-term observable deserves its own line. Watch not merely whether AI-native insurance products emerge that price generated-software risk, but whether underwriting becomes standardised, coverage broadens, exclusions narrow and risk-adjusted premiums fall. If insurers can repeatedly price that risk without bespoke investigation or exceptional capital charges, the recourse clock, treated above as among the slowest, is visibly accelerating, and the taxonomy is being tested at its most load-bearing joint.

To prevent the framework becoming a retrospective naming exercise, each sectoral test must identify the candidate complements before observing the resulting margins. For each complement, specify the exclusion mechanism, switching cost, portability, time to substitute and predicted source of pricing power. A previously unlisted hidden bottleneck cannot be introduced after the result without independent evidence that it constrained substitution during the test period.

The empirical claim is weakened if durable excess margins remain after the identified exclusion rights, switching costs and portability barriers have demonstrably fallen. The ordering claim is weakened if, across comparable markets and absent a separate regulatory or contractual shock, readily generated and portable layers retain pricing power more persistently than the less portable state, recourse or permission layers predicted to replace them.

The foundational statement, that constrained substitution permits rent, is a definition-level economic proposition. The falsifiable contribution of this coda is the proposed location and relative decay of those constraints as models climb the stack.

The corrected ending

The trilogy’s conclusion was that value migrates from cognitive capital to physical capital.

The complete conclusion is this:

Rent follows the slowest-decaying complement still binding, splits when several bind at once, and vanishes when none do.

Stated in full: economic rent forms around binding, excludable complements that cannot be cheaply reproduced or routed around. Where several complements bind, the rent is divided according to control and bargaining power. As models make cognition, software and adjacent mechanisms more reproducible, surviving rent tends to migrate towards the accumulated state, coordination, enforceable rights, loss-bearing capital, physical capacity or sovereign permission that remains less substitutable. Competition can also fragment or destroy the rent without appointing a successor owner.

Physical capital is one possible destination. It is not the economic terminus.

The model can generate the application. It can generate the marketplace software, much of the compliance machinery, the integration and the maintenance patch. What it cannot instantly generate is the fact that everyone is already there. It cannot generate the history on which a reputation rests, the shared decision that one record counts, the legal right to perform a regulated act, or the balance sheet that pays when the system fails.

That is where rent survives. Not permanently. Not necessarily in atoms. In the last binding claim the model cannot reproduce or route around.

When that claim becomes contestable, the bottleneck moves again.

When no bottleneck remains, the market does not appoint a new rentier.

The rent disappears.

Sources

  1. David J. Teece, “Profiting from technological innovation: Implications for integration, collaboration, licensing and public policy”, Research Policy 15(6), 1986. The complementary-assets result: when imitation is easy, profits from innovation can flow to the owners of complementary assets rather than to the innovator, with the allocation depending on the appropriability regime and control of those assets.
  2. Erik Brynjolfsson, Avinash Collis, W. Erwin Diewert, Felix Eggers and Kevin J. Fox, “GDP-B: Accounting for the Value of New and Free Goods in the Digital Economy”, NBER Working Paper 25695. The measurement framework for welfare gains that fall outside measured GDP. The fiscal-visibility inference drawn in this coda is the author’s extension, not a claim of the paper.
  3. OpenAI, “Developers can now submit apps to ChatGPT”. Evidence of model companies operating discovery and distribution directories around generated functionality.
  4. Microsoft, commercial marketplace certification policies. Marketplace certification requirements applied to AI apps and agents; evidence of platform governance enclosing the layer around generated functionality.
  5. eBay Inc., 2025 Annual Report (Form 10-K) and Q4 2025 results. Reported marketplace take rate, defined as net revenues divided by GMV, of 13.94 per cent in 2025 versus 13.77 per cent in 2024, with first-party advertising penetration and shipping programmes contributing to revenue growth.
  6. Booking Holdings Inc., 2025 results. Revenue equal to approximately 14.5 per cent of gross bookings in 2025 versus 14.3 per cent in 2024, with increased payment-facilitation revenue contributing to the rise.
  7. Ben Luong, “The Second DeepSeek Moment”, Part I of the trilogy. The contestability of the model layer.
  8. Ben Luong, “The frontier labs are building a product Hetzner will sell like bandwidth”, Part II of the trilogy. Accepted-output economics and the two routes to commodity-priced cognition.
  9. Ben Luong, “The Margin Call”, Part III of the trilogy. Scarcity-priced obligations, the permits-for-equity window revised above, and the fiscal floorboards extended here.

End of the trilogy and coda.

Copies all three essays and the coda as clean Markdown, including citations and the full proposed methodology.

Back to the reading edition ↑