The Shortening Half-Life of Intelligence: Complete Trilogy

← Trilogy overview
The Shortening Half-Life of Intelligence · Complete edition

The Shortening Half-Life of Intelligence: Complete Trilogy

All three essays in one continuous reading edition.

One page, one argument: fact, transmission mechanism, then financial and political consequence.

Three essays by Ben Luong · 30 July 2026

Copies clean Markdown, including citations and the full proposed methodology.

Part I · Fact

The Second DeepSeek Moment

When the scarcity premium became contestable

Kimi K3 did not prove that open models are free to run, that benchmarks are the economy, or that every proprietary laboratory is finished. It proved something narrower and more dangerous: near-frontier capability now exists as a downloadable outside option. The best model may still command a premium. The old assumption that raw intelligence could sustain one while the surrounding capital stack reprices, refinances and amortises no longer looks safe.

By Ben Luong · 30 July 2026

The shrug

Eighteen months ago, a Chinese laboratory called DeepSeek released a model that helped wipe nearly $600 billion from NVIDIA’s market value in one day, then the largest one-day loss in US stock-market history. Commentators reached for Sputnik. Senators demanded investigations. The Western AI industry held a crisis meeting that has, in some sense, never adjourned.

On 16 July 2026, Moonshot AI released Kimi K3. On Artificial Analysis’s launch snapshot, it scored 57 on the Intelligence Index, behind only Claude Fable 5 and GPT-5.6 Sol. It placed third on GDPval-AA v2 at 1,668 Elo and second on AA-Briefcase at 1,547 Elo. Those are dated launch rankings, not permanent titles. The leaderboard changed again within days.

The economic fact is not the medal but the gap. An open-weight model had arrived within three Intelligence Index points of the proprietary leader. On 27 July, Moonshot released the full 2.8-trillion-parameter weights under the K3 licence.

The reaction was a shrug.

The second DeepSeek moment has arrived, and almost nobody is treating it as one. The shrug is not proof of its economic significance. It marks the change in prior that this essay tests through prices, routing and investment.

The first shock was treated as an anomaly. The second is being treated as weather.

What K3 actually is

Strip away the launch noise and the verified picture is both less magical and more economically consequential.

K3 is a 2.8-trillion-parameter mixture-of-experts model with native image input and a one-million-token context window. Moonshot says it still trails the strongest proprietary systems overall. Independent launch testing put it close enough to them to make the distinction between best and adequate commercially important.

Benchmarks do not measure the whole economy. GDPval-AA and AA-Briefcase measure defined distributions of agentic digital work under particular harnesses. They do not settle reliability in a bank, latency in a consumer product, integration into an old enterprise stack, compliance in a hospital, or the expected cost of a confident error. They establish broad near-frontier capability over the work they test. Commercial substitutability still depends on the cost of failure.

The pricing picture requires the same discipline. Moonshot’s first-party API costs $3 per million uncached input tokens and $15 per million output tokens. That is not the cheapest leading API. Claude Sonnet 5, for example, is on introductory pricing of $2 and $10 through 31 August before moving to the same $3 and $15 standard rate. A closed laboratory matching that rate card is evidence that price compression can occur without a balance-sheet event. The credit claim in Part III requires the separate condition that realised rental economics fall below the obligations written against them.

Artificial Analysis estimated K3 at $0.94 per task across its evaluation suite, close to GPT-5.6 Sol at $1.04 and roughly half Claude Opus 4.8 at $1.80. That is an evaluation-suite estimate, not the universal cost of completing commercial work. K3 also remains unusually verbose and expensive relative to open-weight peers. It used about 132 million output tokens across the Intelligence Index evaluation, materially above the reported median, although fewer than K2.6.

K3 is not revolutionary because Moonshot’s own API is the cheapest. It is not. The significance is that near-frontier capability has become downloadable, contestable and available to a serving market whose future price Moonshot no longer controls.

The outside option

The economically significant threshold was never technical supremacy. It is substitutability: the point at which the quality sacrificed by using the cheaper system is worth less than the money saved.

K3 does not need to beat the frontier. It needs to be close enough that a buyer can credibly threaten to move defined workloads after production evaluation. The launch evidence does not establish that threshold across most commercial work. It establishes enough capability to justify testing an expanding set of bounded tasks. Where those tasks clear the acceptance threshold, the open alternative can constrain the premium without taking every request.

Market power is strongest when the customer lacks a credible alternative. Contestable pricing begins before migration. A buyer can dual-source, route routine work elsewhere, demand a discount, reserve the proprietary model for the difficult tail, or build an internal alternative it may never fully deploy.

Contestability begins before commoditisation and can survive substantial switching costs, because negotiation responds to alternatives before deployment does.

A credible option can constrain the premium before most customers exercise it. It does not cap every price. It caps the premium on the tasks for which buyers could credibly switch.

Rents begin to die not only when customers leave, but when enough of them credibly could.

This is why the launch-week objection that K3 does not definitively beat the frontier misses the economic claim. A competitive market does not require the cheaper supplier to be the best. It requires the gap between best and adequate to fall below the price difference for enough work to discipline the seller.

That is not yet true everywhere. It does not need to be.

Downloadable does not mean free

The weights are available without an acquisition fee. They are not public-domain arithmetic floating above law, capital or engineering.

The K3 licence grants broad rights to use, copy, modify, distribute, fine-tune and deploy the model. It also contains conditions. A model-as-a-service operator whose group revenue exceeds $20 million over any consecutive twelve months must reach a separate agreement with Moonshot before commercial use. Very large consumer products face attribution requirements. Applicable law still applies.

Self-hosting changes the supply chain. It does not remove the deployment from law. The GDPR still follows personal-data processing, and the AI Act still allocates obligations to providers and deployers for covered uses.

Nor is a 2.8-trillion-parameter model cheap to serve merely because only part of it is active for each token. Moonshot recommends supernode deployments with at least 64 accelerators. The full expert pool must still be stored, distributed and routed. K3 is computationally sparse and infrastructurally enormous.

The precise claim is therefore not that release instantly makes global inference cheap, or that any competent host can serve K3 tomorrow at the cost of electricity. It is this:

The weights can be downloaded, modified and served by a broad global ecosystem, subject to Moonshot’s licence and the capital required to deploy them.

The licence is a friction. The deployment cost is a barrier. Neither restores exclusive ownership of the capability to Moonshot or to the Western frontier laboratories.

Why confirmation matters more than surprise

Markets and media price surprises. DeepSeek’s information content was that an open-weight laboratory could approach the frontier. That was unexpected and therefore dramatic.

K3’s information content is different. It says the first result was not safely dismissible as an accident. The open tier can return near the frontier on another release cycle, from another laboratory, while the proprietary leader continues the expensive work of moving the frontier again.

The leaderboard position will move. It already has. That is not an embarrassment to the thesis. It is the thesis. The question is not whether K3 remains third. It is whether an economically adequate open tier remains close enough, often enough, to discipline the closed tier.

A surprise might reverse. A repeated result changes the prior. The shrug is what it looks like when an extraordinary claim becomes an ordinary fact, and that normalisation is itself the discontinuity arriving: not with a crash, but with a calendar.

The hidden theorem is not that AI has become a commodity. Compute remains scarce. Reliable deployment remains difficult. Product moats, proprietary data, regulated distribution and customer relationships can remain valuable.

Nor can an outsider establish that any particular frontier model failed to recover its research, training, serving and organisational cost before its premium narrowed. Those fully allocated economics are not public. A short period of exclusivity can still support a profitable succession of products.

The observed claim is narrower:

The proprietary half-life of economically useful cognition is shortening.

The central financial hypothesis follows from that observation, but is not identical to it:

The quality-adjusted selling price of adequate cognition is falling faster than parts of the surrounding capital stack can reprice, refinance and amortise.

A temporary lead can still be worth a fortune, and frontier laboratories may remain profitable by repeatedly creating the next one. The exposed obligations are those whose repayment depends on the previous premium lasting longer than the market now allows.

The standalone model-layer fork

At the standalone model-API layer, the frontier laboratories now face a fork.

Hold price, and they invite customers to move routine volume, dual-source supply and use the open tier as leverage. The premium market may remain large, but it contracts towards the tasks where the reliability gap is still worth paying for.

Cut price towards cost, and they preserve volume by surrendering margin. Capital raises, cloud bundles and introductory rates can delay the accounting. They cannot turn a contested input back into a scarce one.

There is a third corporate move, but it is an exit from the layer: own the workflow, the customer, the distribution, the proprietary data, the device or the regulated relationship. The frontier companies are already moving into agents, coding environments, enterprise products and hardware. That is exactly the strategic behaviour this thesis predicts, although vertical integration alone does not prove why they are doing it.

Those moves may preserve company value. They do not by themselves preserve the old scarcity rent on raw intelligence.

The distinction matters. This is not an obituary for every frontier laboratory. It is an obituary for the assumption that the model layer itself will remain scarce merely because producing the next frontier model is expensive.

What has, and has not, been shown

K3 has not shown that open weights automatically produce low customer prices. Competitive hosting, adequate supply and efficient routing still have to transmit model contestability into the cost of delivered work. Infrastructure rent may survive even where model-layer rent does not.

It has not shown that all providers are interchangeable. A compatible interface can reduce switching costs while prompts, tool schemas, safety behaviour, latency, evaluation and compliance still require revalidation.

It has not shown that every economically valuable task is substitutable. The hardest, highest-risk work can sustain a premium long after routine work moves.

What it has shown is enough for Part I: near-frontier capability has entered the open-weight layer; the closed frontier now faces a credible outside option; and the standalone scarcity premium is now exposed to a visible and testable shortening of its half-life.

The first DeepSeek moment was an alarm: the frontier is reachable. The second is the sound of everyone sleeping through the confirmation.

The discontinuity was never going to announce itself twice. It arrives the second time as an ordinary Tuesday: a leaderboard update, a rate card, a licence and a download link.

A compact falsifiability note

This essay makes a narrow prediction, not a declaration that every model premium disappears. Freeze a representative basket of agentic digital tasks, the harnesses, the provider set, a minimum acceptance threshold and total cost per accepted output, then test it quarterly through 31 July 2027.

The claim takes serious damage if the best downloadable model remains more than 10 per cent behind the best closed model on accepted outcomes for two consecutive quarters. It also takes serious damage if, despite a downloadable model meeting the acceptance threshold, the leading closed providers sustain a quality-adjusted price premium above two times across the broad task basket for four consecutive quarters. A premium confined to the difficult tail does not refute the claim. It is what the claim predicts.

This test measures the shortening commercial premium. It does not establish whether a named laboratory recovered the fully allocated cost of a particular model generation.


Author’s note: this essay was drafted with the assistance of the model it describes, at commodity prices, during its own launch week. The reader may take that as evidence for the thesis, or as the one part of the argument that needed no essay at all.

Sources

  1. Artificial Analysis, “Kimi K3 achieves #3 in the Artificial Analysis Intelligence Index”, 17 July 2026. Launch-snapshot rankings, Elo scores, task-cost estimates and token use.
  2. Moonshot AI, “Kimi K3: Open Frontier Intelligence”. Architecture, context window, API pricing, scaling claims, limitations and deployment guidance.
  3. Moonshot AI, Kimi K3 model repository. Released weights and deployment materials.
  4. Moonshot AI, Kimi K3 Licence. Rights, model-as-a-service threshold and attribution conditions.
  5. Anthropic, “Introducing Claude Sonnet 5”. Introductory and standard API pricing.
  6. European Commission, “Data protection explained”. The GDPR’s technology-neutral application to personal-data processing.
  7. European Commission, “AI Act regulatory framework”. Provider and deployer obligations under the risk-based regime.
  8. Associated Press, “Nvidia posted another strong quarterly report. What to know, by the numbers”. The January 2025 DeepSeek market reaction.

Part II · Mechanism

The frontier labs are building a product Hetzner will sell like bandwidth

How contestability becomes unit cost dominance

A downloadable model is only an outside option. Standard interfaces, optimisation and routing can turn it into an industrial input through competitive hosting or integrated undercutting. The relevant price is the all-in cost of an accepted output.

By Ben Luong · 30 July 2026

Hetzner is experimenting with an LLM inference API. It offers one open-weight model, no billing, no service-level agreement and no promise of production availability.

The announcement matters because Hetzner can do this at all.

Hetzner is not a frontier AI laboratory. It spent nothing training the model. It did not assemble a research organisation or discover a new architecture. It took an available set of weights, ran them on its own infrastructure and exposed the result through an OpenAI-compatible endpoint.

This is not yet a commodity market. It is the mechanism by which one can form.

Hetzner shows the delivery mechanism. Kimi K3 shows the class of capability entering it. On 16 July 2026, Moonshot released a 2.8-trillion-parameter model that scored 57 on the Artificial Analysis Intelligence Index launch snapshot, within three index points of the proprietary leader. On 27 July, Moonshot released the full weights under the K3 licence.

Put the two together. A laboratory creates the capability. A different company can turn that capability into infrastructure.

The frontier laboratories are spending fortunes inventing a product that companies like Hetzner will eventually sell like bandwidth.

A standard endpoint lowers switching cost without making deployments identical

Hetzner’s experiment serves Qwen3.6-35B-A3B-FP8 through an OpenAI-compatible API. A developer points a standard client at Hetzner’s base URL, supplies a key and names the model.

At protocol level, switching may be little more than changing a base URL, key and model name. At production level, it is not. Prompts, tokenisation, tool schemas, structured outputs, safety behaviour, latency, rate limits, caching, evaluations, legal terms and compliance controls must be revalidated.

Standard interfaces do not make models identical. They do not make providers interchangeable in every production deployment. They reduce the cost of comparing and switching them. That is enough for contestability.

The distinction matters because the economic threshold is not zero switching cost. It is a credible outside option. A buyer that can test several providers against the same task basket, through substantially the same interface, has leverage even when migration still takes work. A routing layer can hold more than one approved endpoint and move eligible requests between them. The incumbent no longer negotiates against the cost of rebuilding the entire application.

That is a poor foundation for durable model-layer margins. Traditional software companies try to own something difficult to replace: a workflow, a proprietary dataset, a network of users, a regulated relationship or the customer itself. A hosted open-weight model owns much less. Competing providers can serve the same weights, related weights or an adequate substitute through a familiar protocol. Competition then moves towards the mundane economics of infrastructure: hardware, electricity, cooling, utilisation, reliability, location and markup.

That is precisely the market Hetzner understands. It does not need to become a glamorous AI company. It can remain what it already is, a ruthless operator of cheap European infrastructure.

The Lidl of machine intelligence.

K3 is downloadable, but it is not cheap or simple to deploy

K3 matters because its capability is close enough to the proprietary frontier to discipline it. It does not matter because Moonshot’s own API is the cheapest. It is not.

On Artificial Analysis’s dated launch evaluation, K3 placed third on the Intelligence Index, third on GDPval-AA v2 and second on AA-Briefcase. The same evaluation estimated an average cost of $0.94 per task, close to GPT-5.6 Sol’s $1.04 and below Claude Opus 4.8’s $1.80. That is an evaluation-suite estimate, not the universal price of commercial work. K3 also used about 132 million output tokens across the suite. It remained unusually verbose and expensive relative to open-weight peers.

The release does not abolish deployment economics. It exposes them.

K3 has 2.8 trillion total parameters and activates 104 billion per token. Its mixture-of-experts architecture selects 16 of 896 experts for each token. Sparse activation weakens the relationship between total parameter count and compute per generated token. It does not abolish the cost of storing, distributing and routing across the full model.

K3 is computationally sparse but infrastructurally enormous.

Moonshot’s published deployment guidance centres on supernode configurations with at least 64 accelerators. The released model already uses quantisation-aware training with MXFP4 weights and MXFP8 activations. It supports low, high and max reasoning effort, with max as the default. Third-party quantisations have already appeared, but further compression necessarily trades deployment cost against quality, latency and behaviour.

The weights are also licensed, not placed in the public domain. The K3 licence grants broad rights to use, modify, deploy and distribute the model. It requires a model-as-a-service operator whose group revenue exceeds $20 million over any consecutive twelve months to reach a separate agreement with Moonshot before commercial use. Very large consumer products face attribution requirements. Internal use is treated more permissively.

The correct claim is therefore narrower than “anyone can serve it”. The weights can be downloaded, modified and served by a broad global ecosystem, subject to Moonshot’s licence and the capital required to deploy them.

The licence is a friction. The hardware is a barrier. Neither restores exclusive control to the laboratory that trained the model.

The important event is not that Hetzner can host K3 tomorrow morning. It is that the intelligence K3 represents has entered a cost-reduction process that Moonshot no longer controls.

Industrialisation attacks every component of the accepted-output cost

K3 has already arrived aggressively quantised. The next reductions do not depend on repeating the same trick.

Serving engineers improve kernels, interconnect utilisation, batching, caching, speculative decoding and request scheduling. Researchers distil task-relevant behaviour into smaller descendants. Hardware improves. Routers reserve the full model for the difficult tail. Each optimisation attacks a different component of the accepted-output cost.

That last phrase is the unit that matters.

Raw token price is not the test. Benchmark rank is not the test. The unit is an accepted output. Unit cost dominance begins when the all-in machine cost of producing that accepted output falls below the fully loaded human cost of producing the same thing.

The machine cost includes inference, integration, orchestration, verification, expected error and rework, compliance and liability, and residual human supervision. The comparison must hold at the required quality, latency and risk. If a human still has to reconstruct the entire output, absorb frequent failures or carry prohibitive liability, unit cost dominance has not occurred for that unit.

This definition includes the strongest objection rather than denying it. Human verification is real. Integration is expensive. Errors matter. Regulation can impose cost. The thesis applies only where the entire deployed system, including those burdens, is cheaper than human-only production.

Once that threshold is crossed, the economic value of the output need not fall. A useful report remains useful. What competition pulls down is the market price of producing it and the rent earned by whoever once supplied it under conditions of scarcity.

The market price of performing the task is pulled towards the all-in unit cost of the cheapest adequate production system.

Routing gives the volume to adequacy

A single model does not need to be excellent at everything. A production system needs to select a model that is adequate for the bounded request in front of it.

In a plausible routed architecture, a cheap model handles the large routine majority of calls while difficult cases escalate to a premium system. The exact share is workload-dependent. The economic point is that the expensive model need not receive the volume.

This changes what progress means. The largest model receives the prestige. The deployment layer receives the requests. A frontier system may remain essential for novel research, difficult coding, exceptional judgement or the highest-risk tail. None of that protects its price across routine work that a cheaper model can complete to the required standard.

K3 itself may be too large for the boring majority of workloads. Its descendants will not need to preserve every capability. A distilled or specialised model can be worse in the abstract and still better for the purchaser if it clears the acceptance threshold at lower all-in cost.

This is why adequacy threatens model-layer rent more than technical defeat does. The closed frontier can remain ahead. Its premium survives only where the gap changes the accepted output.

The model identity then recedes. Applications retain a catalogue of evaluated options. A router assigns work by price, latency, privacy, capacity and measured quality. A provider can be replaced without the customer changing the workflow it sells. Intelligence becomes a workload placed into a competitive infrastructure market.

Field update, 30 July 2026: OpenAI prices the routing mechanism

Twenty-one days after GPT-5.6 entered broad public availability, OpenAI cut Luna’s API price by 80 per cent and Terra’s by 20 per cent, while leaving the frontier Sol price unchanged. Luna moved from $1 per million input tokens and $6 per million output tokens to $0.20 and $1.20. Terra moved from $2.50 and $15 to $2 and $12. Sol remained at $5 and $30.

The asymmetry matters more than the headline. The apex retained its premium. The broad, high-volume tier was radically repriced.

OpenAI’s own guidance now begins with the required outcome and quality standard. It tells businesses to evaluate where additional intelligence materially improves the result and where lower-cost processing preserves the same quality. Its coding example uses Sol to resolve uncertainty and define the plan, then Luna to implement well-specified changes, run tests and evaluate the result.

The expensive model resolves uncertainty. The cheap model implements, tests and evaluates. The account stays with OpenAI; the routine volume leaves Sol.

That is accepted-output economics even though OpenAI does not use the term. The purchaser is told to buy the cheapest system that clears the quality threshold at each stage of the workflow, not the most capable model for every token.

OpenAI says efficiency made the reductions possible. Better routing, serving software and context management produce more useful work from the same compute. It says Sol-assisted kernel work reduced end-to-end serving cost by 20 per cent, while related experiments improved token-generation efficiency by more than 15 per cent. Axios reports that cheaper Chinese open-weight models have also increased pressure on OpenAI and Anthropic to justify their higher costs. The announcement does not identify how much of the pass-through came from efficiency and how much from competition.

Nothing here proves that K3 caused a particular percentage point of the cut. It shows what the credible outside option mechanism can look like. The incumbent keeps the customer while repricing its high-volume tier as if defection were possible.

OpenAI also claims that Luna now delivers performance comparable to models that were frontier-class a year ago at roughly six cents on the dollar per task and nearly nine times the speed. That is a vendor comparison, not an independent or controlled time series. Taken only as an illustration, six cents represents just over four cost halvings in one year, or an economic half-life of about 90 days.

This does not establish the credit event in Part III. Efficiency may protect OpenAI’s margins, and lower prices may stimulate enough demand to increase total compute use. It shows one transmission route directly: OpenAI has repriced its adequate tier and is explicitly positioning Sol’s premium around the stages where additional intelligence changes the accepted result.

Two routes to commodity-priced cognition

The OpenAI cut also exposes a fork inside the delivery mechanism.

The first route is competitive commoditisation. Open weights are served by many infrastructure providers. Standard interfaces and routing layers expose price differences. Hosts compete on electricity, hardware, utilisation, reliability and markup. Efficiency gains and written-down capacity can reach customers through defection between providers.

This is the Hetzner route.

The second route is concentrated commoditisation. A vertically integrated laboratory or hyperscaler uses scale, engineering efficiency, cloud bundling or strategic pricing to undercut independent providers. Routine cognition becomes extremely cheap while the serving market remains concentrated.

This is the Luna route.

Both routes can destroy durable scarcity pricing for routine cognition. Only the first implies a dispersed hosting market. Neither, by itself, proves that a lower acquisition cost for distressed capacity will reach customers after a credit event.

That makes the OpenAI cut double-edged. It strengthens the claim that adequate cognition is becoming cheap. It weakens any assumption that independent hosts must be the organisations making it cheap. It may even intensify pressure on merchant GPU providers that cannot match an integrated provider’s rate, leaving the surviving serving layer more concentrated after the distress.

The title of this essay remains a prediction, not an established industrial structure. A company like Hetzner may eventually sell intelligence like bandwidth. A frontier laboratory may instead preserve the account and sell its own volume tier at bandwidth prices. The labour mechanism can operate through either route. Part III’s fire-sale mechanism requires the additional condition of competitive pass-through after restructuring.

Open weights do not guarantee cheap inference

The transmission remains conditional.

Competition pushes price towards underlying cost only when enough providers can serve adequate models, licensing does not block entry, and effective supply outruns paid workload demand. If infrastructure consolidates into a few hosts, if power and accelerator scarcity dominate the stack, or if demand absorbs every efficiency gain, model-layer rent can die while infrastructure rent survives.

Contestability is the necessary condition. Competitive pass-through is the transmission.

This is also why the argument does not depend on Hetzner personally winning. Hetzner’s public GPU estate may remain too small for frontier-scale serving. Its experiment may end. If Google, Amazon or a specialist inference company owns the cheap layer instead, the model premium can still disappear. The surplus has merely moved to a different landlord.

That outcome would matter. Concentrated infrastructure could keep customer prices well above physical cost. Scarce power, land, accelerators, high-bandwidth interconnect and reliable capacity can support durable rents even while raw intelligence loses its scarcity premium.

The claim is not that every layer becomes perfectly competitive. It is that a model owner cannot rely on permanent scarcity pricing once an adequate alternative can be deployed, evaluated and routed by other firms.

K3 does not prove that pass-through will occur. Hetzner does not prove that supply will outrun demand. Together they make the industrial pathway visible.

Tasks disappear before jobs do

Unit cost dominance operates on outputs. Organisations employ people in jobs. The bridge between them is task decomposition.

Suppose a worker performs twenty recurring tasks. No single model can perform the whole job. One drafts the reports. One reconciles the spreadsheet. One answers routine email. One inspects screenshots. An orchestration layer moves information between them. A smaller number of people supervise exceptions and carry formal responsibility.

The organisation does not need a perfect digital employee. It needs enough accepted output to require fewer producers.

The first quarter looks like assistance. The second looks like productivity. The headcount decision arrives later.

There is no dramatic moment when a machine becomes a perfect employee. There is simply a budget meeting.

The mistake is to wait for one model to perform every element of an occupation. The economically relevant threshold arrives earlier, when a collection of systems produces enough of the output that the organisation needs fewer people. Whole jobs survive on an organisation chart while the volume of human production inside them contracts.

This is how displacement can begin without a mass redundancy announcement. Hiring slows. Junior roles are not reopened. Contractors replace permanent staff. A five-person team supervises work that once required twenty. Headline unemployment can remain calm while the labour market stops absorbing the next cohort.

Sorites prevents a usable boundary

No single step looks like replacement. A tool drafts one report, then checks one spreadsheet, then handles one client queue. Headcount falls at the next budget round.

Assistance and replacement differ by degree, not by a boundary anyone can verify. That Sorites ambiguity makes restraint unenforceable. Every firm can call its own adoption assistance while treating everyone else’s restraint as an opportunity. Nobody can say when defection began, so everybody defects before anybody agrees that replacement has begun.

The distinction remains meaningful at the endpoints. A person writing unaided is producing. A largely automated workflow with one exception handler has replaced most production labour. The problem lies between them. There is no operational line at which assistance becomes replacement, even though the cumulative change is unmistakable.

A rule can require human oversight. It must then define how much attention, at what frequency and with what authority. A firm can satisfy the label through sampling, escalation or formal approval while continuing to reduce the labour content of each accepted output. The human remains visible. Productive necessity recedes underneath.

Sorites is not a claim that regulation is impossible. It is a claim about a particular regulatory defence. Any regime that must identify the precise moment at which assistance becomes replacement is trying to govern a boundary the workflow does not contain.

Sorites is not the net-displacement theorem either. It explains why substitution can proceed without a defensible moment at which assistance becomes replacement. Whether new work absorbs the displaced wage income is a separate empirical question.

New tasks do not necessarily arrive already automated. Digitally expressed tasks increasingly arrive contestable from birth. The same general systems already operate the browser, document, spreadsheet, inbox, dashboard and code editor through which a new workflow is likely to be performed. Human labour may still win many of those contests. It no longer receives an exclusive window by default while specialised machinery is designed around the new task.

The non-absorption claim fails if new human-complementary work restores displaced wage income at comparable scale and speed. It is supported if output and machine-completed work rise while exposed hiring, junior intake, human hours and labour income do not.

The Multiplayer Prisoner’s Dilemma makes adoption compulsory at the system level

The same game repeats at every level.

Laboratories cut prices because rivals might. Hosts optimise because other hosts will. Firms automate because competitors are reducing labour cost. Workers adopt the tools because refusing makes them individually less employable, even though universal adoption reduces the number of workers required.

Each move is locally rational and collectively accelerative.

A laboratory that preserves margin while a near-frontier rival cuts price loses volume. A host that declines to optimise leaves utilisation and customers to another host. A firm that keeps expensive human production for substitutable output while competitors cross unit cost dominance loses on price or margin. A worker who refuses augmentation is compared with one who can supervise more output.

No actor needs to desire the aggregate result. Each needs only to avoid being the actor who pays for restraint while others defect.

Unit cost dominance supplies the payoff. Sorites prevents a usable boundary. The Multiplayer Prisoner’s Dilemma makes adoption compulsory at the system level wherever the cost advantage is material.

That does not mean every individual firm must automate. Some can sustain a premium around human service, trust, regulation, craft or luxury positioning. Those firms become residual niches inside the new cost structure. They do not remove the competitive pressure on the substitutable volume.

This is the missing link between cheaper inference and labour displacement. Capability does not enter the economy because every institution has agreed on its social value. It enters because the cost advantage can be captured privately while the displacement cost is shared.

Regulation loses its single upstream choke point

When the most capable systems are controlled by a few laboratories, regulation has a convenient object. Governments can impose reporting duties, mandate evaluations, restrict particular exports and negotiate with named executives.

Downloadable weights fragment that object across the model developer, host, application provider, router, deployer and regulated institution. The obligations do not vanish. They multiply.

A European company can run K3 on European infrastructure without sending its prompts to Moonshot or a United States cloud. That reduces supplier dependence and some cross-border-transfer exposure. It does not remove the deployment from law.

The GDPR is technology-neutral and still follows the processing of personal data. The AI Act allocates obligations to providers and deployers for specified uses. Sector rules still apply. Moonshot’s licence still governs the weights.

Self-hosting changes the supply chain. It does not make the arithmetic legally invisible.

Regulation can govern employers, hospitals, banks, insurers, public bodies and harmful uses. It can impose liability and documentation duties. What it cannot easily do is restore one upstream choke point once the capability is distributed across models, hosts and jurisdictions.

The delivery layer turns a model event into an economic event

Hetzner may withdraw the endpoint. K3 may remain expensive to self-host. A stronger proprietary model may appear next month. None of those possibilities answers the mechanism.

The mechanism does not require Hetzner to serve K3. It requires near-frontier capability to become an input available to providers that compete on deployment. It does not require providers to be identical. It requires switching and comparison to become cheap enough to discipline price. It does not require zero human involvement. It requires the all-in cost of an accepted output to fall below the fully loaded cost of equivalent human production.

When those conditions hold, the model stops behaving like a premium product and starts behaving like an infrastructure component. The frontier laboratory can preserve company value by owning the workflow, customer, device, proprietary data or regulated relationship. It cannot assume that raw intelligence will continue to carry the old scarcity premium.

For businesses, that is the attraction. Capability becomes cheaper to acquire and easier to embed.

For labour, that is the mechanism.

Workers will keep being told the current model is not quite good enough.

Then, one ordinary quarter, it will be.

What would falsify this mechanism

Fix a basket of production tasks, quality thresholds and providers before measuring it. This account weakens if standard interfaces do not reduce comparison and migration costs, if neither competitive hosting nor integrated undercutting lowers the all-in cost per accepted output, or if verification, error, compliance, liability and supervision keep that cost above equivalent human production. Persistent quality-adjusted model premiums would refute the pricing claim.

The labour transmission weakens if machine-completed accepted outputs rise without human hours, junior intake or the wage share deteriorating in exposed sectors. The stronger non-absorption claim fails if new human-complementary work restores displaced wage income at comparable scale and speed.

Sources

  1. Jonas Scholz, Hetzner Inference: First Look, including the experiment’s single model, OpenAI-compatible endpoint, lack of billing, SLA and production guarantee, and July 23 test.
  2. Artificial Analysis, Kimi K3 achieves number three in the Artificial Analysis Intelligence Index, 17 July 2026 launch snapshot, benchmark positions, measured task cost, token usage and API pricing.
  3. Moonshot AI, Kimi K3 Tech Blog, model architecture, launch positioning, multimodality and context window.
  4. Moonshot AI, Kimi K3 model card, released weights, activated parameters, quantisation format, deployment engines and reasoning-effort controls.
  5. Moonshot AI, Kimi K3 Licence, commercial model-as-a-service threshold, attribution requirements and internal-use exception.
  6. European Commission, Data protection explained, including the GDPR’s technology-neutral application to personal-data processing.
  7. European Commission, AI Act regulatory framework, provider and deployer obligations under the risk-based regime.
  8. OpenAI, Advancing the price-performance frontier with GPT-5.6, 30 July 2026. Price reductions, workflow-routing guidance, cost-per-task comparison and serving-efficiency claims.
  9. OpenAI, GPT-5.6: Frontier intelligence that scales with your ambition, 9 July 2026. Original public-availability date and launch API prices for Sol, Terra and Luna.
  10. Axios, OpenAI cuts GPT-5.6 prices, 30 July 2026. Price-sensitive customer context and reported pressure from cheaper Chinese open-weight models.

Part III · Consequence

The Margin Call

When scarcity debt meets contested pricing

Open weights do not make compute free, and contested model pricing does not automatically crash GPU finance. The narrower claim is that the quality-adjusted selling price of adequate cognition may fall faster than parts of the surrounding capital stack can reprice, refinance and amortise. If effective supply outruns paid workload demand, the loss need not end the automation cycle. With competitive pass-through, it can finance the next phase of it.

By Ben Luong · 30 July 2026

Who lent against the rent?

Part I argued the fact. Near-frontier capability has entered the open-weight layer and given buyers a credible outside option. Part II argued the mechanism. Standard interfaces, delivery optimisation and routing reduce the cost of comparing providers and push the delivered price of intelligence towards contested infrastructure economics. That delivery can be dispersed across hosts or concentrated inside an integrated laboratory.

Providers are not identical. Production deployments still require prompts, tool schemas, safety behaviour, latency, evaluation and compliance to be revalidated. Compute, power and reliable serving remain scarce. The claim is not that every friction has disappeared. It is that the assumption of exclusive control over frontier-like capability has become contestable.

This essay asks the question neither of the first two asked.

Who lent money against the assumption that none of this would happen?

The build-out was priced before the scarcity premium became visibly contestable

The AI capital expenditure programme now runs into the high hundreds of billions. A Reuters dashboard reports more than $800 billion of estimated hyperscaler capital expenditure for 2026. Reuters’ 29 July bond-market analysis reports that by 7 July Amazon, Alphabet, Meta and Oracle had issued about $194 billion of bonds, 79 per cent more than they issued during all of 2025. It also documents widening borrowing spreads and cites a Goldman Sachs estimate that debt issuance will equal about one-third of hyperscaler capital spending in 2026.

That establishes material exposure. It does not establish the crash.

It is not one capital structure, and the precise claim matters. Cash-rich hyperscalers have funded much of the build from operating cash flow and can absorb disappointing returns far longer than anyone downstream, although the scale of current commitments is drawing them further into debt markets too. Utilities and data-centre landlords earn contracted returns. Nvidia can remain the largest winner during the build-out because open weights expand demand for accelerators. The sequence can reverse if algorithmic efficiency, overcapacity, slower frontier training and a secondary market in impaired silicon weaken demand for new high-margin hardware. Nvidia can win the boom and still be repriced by the overbuild it supplied.

The vulnerable layer sits at the leveraged edge: neoclouds, GPU lessors and project-finance vehicles whose debt was sized to specific compute rental rates, specific utilisation assumptions and specific collateral values for hardware whose economic life is measured in years rather than decades. Those numbers were set when frontier-like intelligence looked scarce for long enough to repay the capital built around it.

Not all leveraged-edge exposure is merchant exposure. In March 2026, CoreWeave closed an $8.5 billion non-recourse delayed-draw facility secured by high-performance-computing infrastructure and an associated customer contract, with a March 2032 maturity. CoreWeave said its equity and debt financing commitments secured over the preceding twelve months totalled approximately $28 billion. Its first-quarter filing says borrowing is tied to depreciable equipment cost, projected debt-service coverage and project-level conditions. Contract backing may protect facilities like this from short-term spot-price compression. The exposed cohort is therefore narrower: structures in which tenant credit, take-or-pay durability, utilisation, collateral value or refinancing assumptions fail together.

This is not a claim that every frontier model fails to recover its training cost. Model-specific, fully allocated research, serving and organisational economics are not public, and a laboratory can remain profitable through a succession of short-lived leads. The central financial hypothesis is narrower:

The quality-adjusted selling price of adequate cognition is falling faster than parts of the surrounding capital stack can reprice, refinance and amortise.

Part I argued that the customer now has somewhere else to go. For buyers capable of evaluating and deploying alternatives, procurement now occurs in the shadow of downloadable weights that sat within three Intelligence Index points of the closed leader at K3’s launch, servable by competing hosts with the infrastructure and licensing rights to do so. The outside option constrains the premium whether or not anyone exercises it, and whatever Moonshot’s own list price does this quarter. Contested model pricing can flow downstream into the rental rates, utilisation and collateral values the leveraged edge borrowed against.

That transmission is not automatic, and naming the condition is the honest move. Open weights can destroy model-layer rent while increasing demand for GPU-hours. More hosts, more deployments and more inference can absorb more compute. The credit event requires effective compute supply, installed capacity multiplied by hardware and algorithmic efficiency, to outrun paid workload demand. Alternatively, competition among hosts must compress realised rental economics below debt service.

That is the additional prediction this essay makes beyond Parts I and II. If Jevons demand keeps merchant rental rates and utilisation at their underwriting levels, the model rent can die without the leveraged edge breaking. The efficiency curve, the overbuild and the routing layer make that failure plausible. They do not show that the coin has already landed. This essay dates the prediction rather than presenting it as a result.

So the arithmetic is now visible. Debt sized to scarcity pricing, serviced by contested pricing. The gap between those two numbers is where the bubble lives. Kimi did not create the gap. The release of its full weights on 27 July made the gap visible enough to date, not automatic enough to declare.

Volume is not solvency

The standard defence is demand. As machine cognition becomes cheaper, firms will use more of it. But the economically relevant variable is not token volume. It is paid accepted output. Better systems can complete more useful work with fewer tokens, while weakening household demand can reduce the total quantity of final output the economy is willing to buy.

Jevons pressure expands the uses to which cheap cognition is put. It does not guarantee that paid workload demand will outrun simultaneous gains in hardware efficiency, algorithmic efficiency and installed capacity. Nor does it guarantee the rental rate, utilisation or collateral value assumed by the capital stack.

Exploding demand does not guarantee scarcity margins. Electricity, bandwidth and freight carry enormous volumes at infrastructure returns.

Be precise about what volume can and cannot do, because the distinction is where the margin call actually lives. Volume can service debt at utility margins. Grids, pipelines and telecoms are debt-financed on exactly that basis. What volume at contested margins cannot do is preserve software valuations, venture returns, residual values for silicon ageing faster than its loans amortise, or refinancing terms written when capacity looked permanently scarce.

The first loss therefore lands on equity: valuations are written down to the extent that they priced reproducible cognition as a durable monopoly. It becomes a credit event wherever fixed obligations were sized to the old rental rate, the old utilisation assumption or the old collateral value.

Merchant GPU capacity. Leveraged neoclouds. Project vehicles dependent on loss-making tenants. Hardware-backed lending whose collateral ages faster than the loan amortises.

The hyperscalers can absorb disappointing returns for far longer. The leveraged intermediaries may not survive the refinancing.

The frontier labs can see this, which is why they are racing to become product companies: agents, coding environments, enterprise platforms and devices. Part II called that race consistent evidence for the thesis. They are trying to own the system around the intelligence before the intelligence becomes cheap. Some will succeed as product companies, or as integrated providers able to undercut independent hosts while preserving value elsewhere in the stack. That possibility weakens the case for a dispersed hosting market. It does not restore scarcity pricing to routine cognition.

Demand can fall while substitution accelerates

Cognition is an intermediate input, so its absolute demand ultimately depends on final demand. If wage loss weakens mass consumption, the total volume of paid workloads may grow more slowly than the Jevons defence assumes, or contract outright. That strengthens the credit risk. It does not necessarily rescue labour.

Three quantities must be kept separate: total final output, absolute machine workload and the machine share of production.

Let Y be final output, c the cognitive input required per unit of output, and s the machine share of that cognitive input.

Machine-handled cognition: M = Y × c × s

Human-handled cognition: H = Y × c × (1 – s)

Suppose final output falls from 100 to 90 and cognitive intensity remains one. If the machine share rises from 20 per cent to 60 per cent, machine-handled cognition rises from 20 to 54 while human-handled cognition falls from 80 to 36.

Final output has fallen by 10 per cent. Machine work has risen by 170 per cent. Human work has fallen by 55 per cent.

Even where machine work does not rise in absolute terms, human work can fall much faster than total output. A firm serving a smaller market can still serve it with one supervisor and a machine system rather than five employees. The displacement variable is not total compute consumption. It is the machine share of the cognition used to produce each remaining accepted output.

This creates a feedback loop. Substitution weakens wage income. Weaker wage income weakens mass demand. Weaker demand intensifies the pressure to reduce unit cost. Lower machine prices make further substitution available. The same contraction can therefore break infrastructure debt through insufficient absolute workload while breaking labour through a rising machine share.

The compute owner needs sufficient absolute revenue to service fixed obligations. The worker needs human labour to retain sufficient relative share of production. Those are not the same requirement.

Earlier automation waves offer a precedent, not a forecast. Jaimovich and Siu found that routine employment losses in the United States were heavily concentrated in downturns and that jobless recoveries were largely accounted for by routine occupations that disappeared. Robert Allen’s account of Engels’ Pause describes a long interval in which productivity and output rose much faster than workers’ living standards before the gains were more broadly shared. Neither result proves that the current transition will follow the same path. They show that restructuring can concentrate in contractions and that eventual recovery is not an answer to a generation living through the delay.

The counterforce is real. Weak firms may lack the cash, organisational capacity or confidence to implement new systems. Consolidation may prevent lower infrastructure costs from reaching them. The acceleration claim fails if those frictions suppress adoption more than falling machine costs and unit-cost pressure encourage it.

During the transition, final demand can still come from wages that remain, household savings and credit, government deficits and transfers, capital income, investment spending, exports and lower prices that stretch residual income. None is automatically a permanent replacement for wage-mediated mass demand. Without redistribution or a new claim on automated output, the process can culminate in underused capacity, repeated demand shortfalls and concentrated production rather than universal abundance. The absence of a stable replacement demand engine is not a premise required to start the transition. It is part of the successor-system problem the trilogy leaves open.

The crash can speed up the displacement

Here is the part the public conversation has exactly backwards.

The assumed sequence runs like this. The bubble bursts. The AI story dies. The pressure on jobs eases. Workers watching the capital expenditure numbers with dread are quietly hoping for the crash, on the theory that a crashed industry stops hiring machines.

The labour sequence may already have begun, quietly. Junior hiring can thin, graduate absorption can weaken and incumbent teams can be asked to do more without producing a clean break in headline unemployment. The No-Scream Principle, the claim that compositional damage appears before aggregate measures scream, predicts exactly this sequence.

Credit behaves differently. It has payment dates. The prediction here is not that no worker is harmed before a lender is. It is that financing stress at the leveraged edge becomes the first hard, dated and institutionally undeniable break, before AI displacement appears in headline aggregate unemployment.

Under that sequence, the leveraged edge breaks. Merchant clusters are refinanced, sold or absorbed at steep discounts. The durable sites may hold their value. The acquisition basis of the compute inside them does not. Inference can then clear below the return originally required to build the capacity because somebody else has already eaten the capital loss.

One counterforce deserves more weight, because the acceleration claim depends on beating it. The most plausible buyers of distressed capacity may be hyperscalers that already own adjacent infrastructure. If a few incumbents absorb the written-down assets and preserve the prevailing price umbrella, the lower acquisition basis becomes landlord surplus rather than cheaper intelligence.

Historic cost is sunk. A host with cheap assets will charge the prevailing rate unless idle capacity and competitive defection force it to cut. Standard interfaces and routing make price differences visible and switching easier, but they do not create competitors. The discount reaches customers only if capacity remains contestable after restructuring, through independent hosts, forced disposals, open access or another mechanism that keeps survivors competing to fill the assets.

This creates two separable claims. A credit event can occur without pass-through. The acceleration claim requires both the write-down and competitive pass-through. If consolidation removes the second, the crash does not subsidise diffusion even though lenders may still take the loss.

If competition survives, consultancies, insurers, law firms, software companies and government departments that found frontier API pricing hard to justify get a different calculation. The marginal adopter, the firm that ran the pilot and shelved it on cost grounds, reruns the numbers and may adopt. Substitution can accelerate into the downturn, not out of it, because downturns are precisely when firms cut headcount and hunt for cheaper inputs. Under those conditions, the crash has delivered a cheaper cognitive input.

The liquidation sale reaches individuals too, and here it intensifies a dynamic already running. Freelancers are underbidding agencies today, at today’s prices. Agencies are underbidding in-house teams. In-house seniors are running departments without juniors. Juniors are using the same tools to look employable a little longer. Consultants are selling transformation plans that automate the clients who commissioned them.

Each actor adopts to survive the round, and widespread adoption removes the reason many of them were needed. That is the Multiplayer Prisoner’s Dilemma operating at the level of individual careers, and it does not wait for any crash.

What the crash changes is the floor, but only where competing hosts pass the economics of written-down clusters through an API. That does not require every freelancer to run a 2.8-trillion-parameter model on a second-hand gaming card. It requires temporary access to infrastructure whose original owner has already eaten the capital loss to become a cheaper weapon in the pit.

The game is already being played. A pass-through crash makes each move cheaper.

There is precedent, and the structure is familiar. Railway investors were ruined in the 1840s. The railways remained, and their fire-sale capacity industrialised freight. Telecom investors were ruined in 2001. The fibre remained, and its stranded capacity carried the internet economy for twenty years at prices the original investors never planned.

One honest difference matters. Rail and fibre were durable assets. This generation of accelerators is not. What survives an AI crash is the durable shell around the silicon: the grid connections, substations, buildings, cooling, power contracts and permitted sites. The shell is the reason the capacity survives the crash. It is not necessarily the discounted asset. The subsidy comes from the excess compute, the impaired silicon and the capital claims written off above them, while the permitted shell removes years from the next owner’s deployment timetable.

The crash also has opposite effects on development and diffusion. Capital destruction may thin the trillion-parameter training runs and slow the frontier. It can accelerate diffusion of capability already created because the systems already shipped are sufficient for a large class of commercial substitution and their delivery cost has fallen.

The objection that bankrupt AI firms cannot advance AI is true and beside the point.

The displacement does not need the next model. It needs the last one, cheaper.

The financial parasite dies. The automation organism spreads through its corpse.

A crash meeting these conditions may be reported as the end of the AI era while functioning as a subsidy to the replacement of labour. It becomes a liquidation sale on human obsolescence only where the discounts are passed through to employers able to use them.

The state is standing on four rotting floorboards

The credit event and the fiscal crisis run on different clocks. The first is a late-2020s refinancing prediction. The second compounds into the 2030s as compositional labour losses accumulate. Payroll and consumption receipts can weaken at the margin before headline unemployment breaks, but a state-wide cash-flow crisis is not the same-quarter consequence of the margin call.

Now put the government in the picture, because the fiscal exposure may be worse than the financial one and almost nobody has added it up.

The state taxes wages. AI compresses wages, first through quiet non-absorption at the entry level, then through mid-career restructuring. One of the largest and most automatically collected tax bases in every developed economy erodes at the source: broad, domestic, visible and withheld at payroll.

The state taxes corporate profits. The rents that survive do not necessarily vanish. They migrate out of the contested model layer and into clouds, chips, energy, land and distribution. They accumulate in fewer hands, become more internationally mobile, and sit with owners better placed to bargain with any individual government. A base that was broad and captive becomes narrow and negotiable.

The state taxes consumption. Displaced households consume less. VAT receipts follow wages down.

The state borrows against future growth. Here cheap AI may perversely fulfil the GDP forecast while breaking the fiscal one. Output can rise even as wage income, payroll receipts and mass consumption fall, if the gain accrues to a small and internationally mobile capital base. Productivity is not tax capacity. The state can receive the growth it was promised and still lose the revenue it borrowed against.

These are four gross pressures, not an accounting identity. Cheap AI could also raise taxable profits, lower prices, expand consumption and reduce some public-service costs. The claim is about timing and incidence. Payroll losses arrive automatically and domestically. Replacement rents are narrower, more mobile and slower to capture. Transfer demands arrive before legislation, treaties and enforcement can redirect the surplus.

The state can receive the productivity gain and still suffer the cash-flow crisis.

The state does not lose because output disappears. It loses because the tax base moves faster than the tax code. Payroll taxes itself in real time. The surplus replacing payroll must be chased through treaties, legislation, lobbying and elections, and it can relocate while the bill is in committee. Meanwhile the transfer demands arrive when employment income weakens. The fiscal crisis is a timing mismatch as much as a revenue loss. The state must fund the transition before it has captured anything winning from it.

The surplus may be internationally mobile, but the scarce physical complements are territorially fixed.

That is why the tax system will crawl towards what cannot relocate: land, power generation, grid connections, water rights, data-centre footprint, resource extraction and territorial permission itself. States retain other instruments, including wealth taxes, financial taxes, procurement and sovereign investment. They will use them. The most dependable base is the physical one because it cannot move to another jurisdiction after the finance minister speaks.

The crawl towards it will be presented as strategy when it is elimination.

Note what that implies for bargaining position. A state whose dependable base is land, power and grid access is a state negotiating with the owners of land, power and grid access. It arrives at that table late, stretched, and needing them more than they need it.

The window for permits-for-equity is open and closing

The timing of the constructive move matters more than its design.

I set out the mechanism in an earlier essay, where the constraint analysis identified permits-for-equity as the cleanest pre-commitment instrument under this framework. It is not literally the only instrument. Governments can invest directly, attach royalties, use development banks, retain public land, acquire rescue equity or tax gross flows.

The distinctive advantage is ex ante conversion of public scarcity into a permanent claim before capital is sunk.

The build-out still needs things only states can grant. Land. Planning approval. Grid connections. Water. Transmission. Political protection. Permission is not the state’s only leverage, but it is its cleanest. Before construction, a permanent public claim can attach without a bailout, a retrospective tax fight or a threat of closure. It becomes a transparent condition of access that the developer can price before committing capital.

That leverage is uneven. A jurisdiction demanding equity can lose a sufficiently mobile project to one that does not. Permits-for-equity works unilaterally only where permission attaches to a binding, location-specific bottleneck: a constrained grid connection, publicly financed transmission, public land, sovereign procurement, guaranteed demand or rescue support. Where sites and jurisdictions are genuine substitutes, the claim requires regional coordination or it will be competed away.

The state’s leverage does not disappear after construction. Expansions, power, water, tax treatment and regulation continue. It becomes costlier and more politically adversarial to exercise once capital is sunk. The cleanest leverage exists on the way in and degrades the moment the permit is irreversible.

The exchange rate matters. Trade permission for equity. Trade it for permanent public claims on the infrastructure.

Do not trade permanent access to land, power and grid capacity for headline employment commitments. Construction employment is temporary, and permanent operating payroll is small relative to the capital value and public scarcity being allocated. Data centres still employ people in operations, networking, security, power engineering and their supply chains. The point is not that they employ nobody. It is that automation will place continuing downward pressure on a payroll already small beside the concession.

The state that swaps permanent physical access for temporary payroll promises will spend the next decade taxing the asset it forgot to keep a share of.

The window is narrow because a crash compresses it. Distressed compute changes hands quickly, the surviving owners consolidate, and owners with sunk permits need less from the state. The hopeful reply is that a crash hands the state fresh leverage as workout referee. It does.

Workout leverage can even buy equity, but only if the state supplies rescue capital, guarantees or coercive law at the moment its own balance sheet is weakest. Permit leverage acquires the claim before the crisis. Workout leverage requires the state to purchase it during one.

The bolder reply is that the state should wait and buy everything cheaply at the bottom. It fails on four counts.

The fire sale is in the layer the trilogy says is losing value: model businesses, leveraged shells and depreciating GPUs. The assets worth owning, the best grid connections, land, water and generation, retain or gain scarcity value through the crash because the crash is the value migrating into them.

The buyer is fiscally and politically weakest on the day of the sale, with tax receipts eroding and transfer demands rising. Weaker states may also face higher borrowing costs. Stronger currency issuers can fund the purchase cheaply but can still move too slowly to clear the assets that matter.

Private capital clears a distressed auction in a weekend. Nationalisation needs a statute, a valuation and a court. The consolidation can finish before the enabling act passes.

Even a successful purchase buys melting machinery downstream of a chip supply chain the state still does not control. A few currency issuers with cheap energy will run the bottom-buying play well. That is part of the sorting between states, not a plan available to the rest.

Buying the bottom is the strategy of a state that missed the window, executed with money it no longer has, for assets that stopped mattering.

The leverage that converts permission into permanent claims exists cleanly now, while the concrete is still being poured. Afterwards, the same claim must be bought, taxed, litigated or rescued.

What would falsify the argument?

This essay makes four separable claims, and the most exposed one is pass-through.

Do not treat AI demand as one variable. The empirical record must separate three groups of measures:

Infrastructure solvency: realised price per accepted output, paid workload revenue, cluster utilisation, debt service, refinancing conditions and collateral values.

Technological diffusion: deployed production workflows, machine-completed accepted outputs, machine share of task execution and cost per accepted output.

Labour displacement: human hours per accepted output, junior and graduate hiring, headcount relative to output, wage income in exposed occupations and the reinstatement of equivalent wage income in new occupations.

Paid workloads can disappoint while the machine share rises. That combination is not contradictory. It is the route by which infrastructure investors and workers can lose in the same contraction.

The pricing claim takes serious damage if neither independently hosted open-weight systems nor integrated volume tiers approach the best closed systems on accepted commercial outcomes at a materially lower quality-adjusted price, or if closed providers sustain their broad premium through 31 December 2027.

The credit claim fails if model-layer pricing compresses but realised merchant rental economics remain above debt service after the relevant borrowers have tested the refinancing market through 31 December 2029. The sequencing claim fails if AI-exposed employment produces a clear break in headline aggregate unemployment before a composite credit turn at the leveraged edge.

The acceleration claim fails in either of two cleaner ways. If asset values fall but accepted-output costs do not, demand or consolidation has stopped the write-down from reaching customers. If accepted-output costs fall but production adoption and machine share do not accelerate, cheaper supply has not produced the predicted diffusion. The labour claim weakens if human-complementary work restores displaced hours and wage income at comparable scale and speed.

If neither the headline labour break nor the credit composite appears after the relevant maturity window, the result does not merely reorder the sequence. It falsifies the timetable and weakens the wider claim that contestability is transmitting into the economy with the speed and force argued here.

Growing laboratory revenue would not falsify it. Commodity markets have large sellers. Revenue is not rent.

Nor does OpenAI’s 30 July price cut establish post-crash pass-through. It demonstrates delivered-price compression and workflow routing. The separate acceleration claim begins only when a distressed acquisition basis is transmitted into lower accepted-output cost after a credit event.

Methodological appendix: a proposed empirical test

The design below is a research proposal, not a completed preregistration. The 48-task basket, borrower cohort and enterprise panel have not been published, recruited or frozen. The protocol becomes real only when those materials, the acceptance rubric, baseline model IDs and available data are disclosed. Until then, it is the programme a funded research organisation would need to execute.

Any implementation should publish its baseline date in advance and use trailing ninety-day medians so that one promotional rate, one distressed borrower or one noisy employment release cannot decide the result.

Open the proposed protocol details

Pricing. Publish and freeze a basket of 48 commercial tasks across coding, research, document production, spreadsheet work, customer operations and regulated professional workflows. An independent panel must select the basket, or sample it from a published external task frame, under inclusion and exclusion rules disclosed before evaluation. Freeze prompts, tool access, latency limits and the acceptance rubric. At the first run, record the exact production model IDs for the highest-scoring closed API from Anthropic, Google and OpenAI, plus the two strongest independently hosted open-weight models whose licences permit the test. An output is accepted only when two of three blinded domain assessors judge that it needs no substantive repair. Cost per accepted output includes inference, retries, orchestration and human verification. Human comparison uses task-specific fully loaded labour costs taken from a preregistered external wage-and-overhead source. Results are also reported under common-rate sensitivity scenarios so that no single labour-cost assumption determines the conclusion. Before evaluation, the preregistered rubric must set an absolute acceptance floor for each task category, allowing risk thresholds to differ across coding, customer support and regulated professional work. An open option qualifies only if it clears the preregistered absolute acceptance floor for its task category and reaches at least 90 per cent of the best closed model’s acceptance rate. The open-weight contestability claim fails if no open release qualifies across the next two major open-weight generations. The broader pricing claim also fails if neither an open option nor an integrated volume tier produces material quality-adjusted compression and the broad premium falls by less than 10 per cent from its baseline level in four consecutive quarterly runs through 31 December 2027. A fall of at least 25 per cent from baseline, sustained for two quarters, counts as compression on schedule.

Credit. Publish a cohort of AI-infrastructure borrowers with debt equal to at least 30 per cent of invested capital and a refinancing or maturity date between 1 January 2027 and 31 December 2029. Track paid workload revenue and realised quality-adjusted price per accepted output alongside financing measures. A credit turn requires at least three of the following across two unrelated borrowers within two consecutive quarters: a 25 per cent year-on-year fall in quality-adjusted merchant GPU rental rates, a 15 percentage-point fall in disclosed cluster utilisation, a 250 basis-point widening in refinancing spread over the matched sovereign curve, a 20 percentage-point increase in collateral haircuts, a covenant waiver or take-or-pay renegotiation, or an impairment or distressed sale at least 25 per cent below carrying value. The credit claim fails if pricing compression occurs and no composite turn appears by 31 December 2029, provided at least half of the cohort’s principal has passed a contractual refinancing or maturity date. If less than half has tested the market, the clock runs until it has.

Labour and sequence. Pre-register 20 highly exposed occupations and 20 controls matched on baseline wage, education, sector and prior cyclical sensitivity in both the United Kingdom and United States. Track employment, advertised vacancies, the share of vacancies asking for no more than two years’ experience, median advertised real wage, human hours per accepted output and headcount relative to output. Track whether new occupations restore equivalent wage income, not merely whether new job titles appear. A quiet compositional turn occurs when the exposed basket underperforms its controls by at least 10 per cent in vacancies and five percentage points in entry-level share for two consecutive quarters. That may occur before credit and does not falsify the essay. A headline labour break requires unemployment to rise at least one percentage point year on year in either country while employment in the exposed basket underperforms its controls by at least three per cent. The sequencing claim fails if that headline break precedes the credit composite. Credit first supports the prediction that the margin call becomes the first damage nobody can shrug at.

Acceleration. The first credit composite sets the event date. A funded research partner would need to recruit a vendor-neutral panel of 500 enterprises, stratified by country, sector and firm size. It would record quarterly the share of defined production workflows using AI, machine-completed accepted outputs, machine share of task execution and human hours per accepted output, excluding pilots and individual ad hoc use, with the questions and workflow definitions held fixed. The write-down becomes an adoption subsidy only if quality-adjusted cost per accepted output falls at least 20 per cent beyond its pre-event four-quarter trend within the next four quarters, and the panel’s median production-workflow share grows at least 50 per cent faster than its pre-event four-quarter rate for two consecutive quarters. The joint labour-acceleration claim additionally requires human hours per accepted output, exposed hiring or labour income to deteriorate faster after the event than before it. If asset values fall but accepted-output prices do not pass through, consolidation has interrupted the mechanism. If prices pass through but production adoption and machine share do not accelerate, cheaper supply has not produced the predicted diffusion. Either result falsifies the acceleration claim.

Series close: the redistribution, named

You cannot base the successor system solely on frontier-lab monopoly windfalls that competition is already eroding. A viable fiscal claim must attach to volume, ownership and scarce physical complements, not merely to the accounting profits of whoever currently has the smartest model.

Put the three essays together and the whole event compresses to one movement.

Near-frontier open weights create a credible outside option. Standard interfaces and routing transmit it through one of two industrial structures: competitive hosting or integrated undercutting. Both can make routine cognition cheap. Only the first depends on continued competition among multiple serving providers.

From there the causal model forks. Where either delivery route lowers the all-in cost of an accepted output below equivalent human production, unit cost dominance spreads through firms and careers. Separately, where effective supply outruns paid workload demand or realised rental economics fall below debt service, obligations underwritten against the previous scarcity inherit the loss. If written-down capacity then reaches customers as cheaper access, the credit branch feeds back into the labour branch and accelerates diffusion.

The demand loop can tighten both branches at once. Substitution can weaken wage income and final demand, disappointing the absolute paid workload required to service infrastructure debt while unit-cost pressure raises the machine share of the output that remains. Machine workload and machine share are different quantities. They can move in opposite directions.

Every arrow is conditional. The falsification summary and proposed research design identify where each branch can fail.

Where those conditions hold, labour loses bargaining power because reproducible cognition stops being scarce. Frontier laboratories need not become unprofitable; some can prosper through successive frontier leads, products and integration. The routine-tier scarcity rent can still compress underneath them. Leveraged investors lose where obligations were sized to a premium that reprices faster than the loan. The state loses where its broadest and most automatic tax bases move faster than its tax code.

The owners of whichever physical and legal complements remain binding, including leading-edge fabs, constrained grid connections, scarce power, permitted sites, ports and enforcement, inherit the position because the software layer prices itself towards utility economics and the machines still need atoms.

That is the redistribution actually underway. Cognitive capital to physical capital.

The software aristocracy is discovering it was renting the throne from the owners of atoms.

Part I opened with a shrug. The second DeepSeek moment arrived and almost nobody noticed because the quiet parts of a discontinuity are easy to shrug at. Hiring that never happens makes no sound. Entry-level roles evaporate without a press release.

Credit is different.

Credit has covenants, coupons and dates.

The first part of this sequence that nobody manages to shrug at will be the margin call.


Sources

  1. Reuters, “AI boom sparks rally, frenzy and fear”. Estimated 2026 hyperscaler capital expenditure.
  2. Reuters, “Hyperscaler debt binge pushes yields up as investor demand cools”. Bond issuance through 7 July, the comparison with 2025 issuance, debt-financing share and borrowing-spread context.
  3. Artificial Analysis, “Kimi K3 achieves number three in the Artificial Analysis Intelligence Index”. Launch-snapshot capability, task-cost and token-use evidence.
  4. Moonshot AI, Kimi K3 model repository. Full-weight release, model architecture and deployment materials.
  5. Moonshot AI, Kimi K3 licence. Use, modification, distribution and model-as-a-service conditions.
  6. Sliplane, “Hetzner Inference: First Look”. Evidence for conventional infrastructure providers exposing open-weight inference through a standard interface, with launch-stage limitations.
  7. Ben Luong, “The Margin Call”. Original LinkedIn publication.
  8. Nir Jaimovich and Henry E. Siu, “Job Polarization and Jobless Recoveries”. Evidence that routine employment losses in earlier United States automation waves were concentrated in downturns. This is a historical precedent, not a claim that the AI transition must repeat it.
  9. Robert C. Allen, “Engels’ Pause: Technical Change, Capital Accumulation, and Inequality in the British Industrial Revolution”. Historical evidence on the long lag between productivity growth and broadly shared gains.
  10. OpenAI, “GPT-5.6: Frontier intelligence that scales with your ambition”. Vendor evidence that tool-heavy tasks can complete with fewer tokens and round trips, illustrating why token volume is not the economic unit.
  11. CoreWeave, “CoreWeave Closes Landmark $8.5 Billion Financing Facility”. Facility size, non-recourse structure, infrastructure and customer-contract security, March 2032 maturity and twelve-month financing-commitment total.
  12. CoreWeave, first-quarter 2026 Form 10-Q. Debt-sizing conditions, collateral, limited guarantees and the distinction between total facility capacity and drawn debt.

End of the complete trilogy.

Copies clean Markdown, including citations and the full proposed methodology.

Return to the trilogy overview →