The frontier labs are building a product Hetzner will sell like bandwidth
How contestability becomes unit cost dominance
A downloadable model is only an outside option. Standard interfaces, optimisation and routing can turn it into an industrial input through competitive hosting or integrated undercutting. The relevant price is the all-in cost of an accepted output.
Hetzner is experimenting with an LLM inference API. It offers one open-weight model, no billing, no service-level agreement and no promise of production availability.
The announcement matters because Hetzner can do this at all.
Hetzner is not a frontier AI laboratory. It spent nothing training the model. It did not assemble a research organisation or discover a new architecture. It took an available set of weights, ran them on its own infrastructure and exposed the result through an OpenAI-compatible endpoint.
This is not yet a commodity market. It is the mechanism by which one can form.
Hetzner shows the delivery mechanism. Kimi K3 shows the class of capability entering it. On 16 July 2026, Moonshot released a 2.8-trillion-parameter model that scored 57 on the Artificial Analysis Intelligence Index launch snapshot, within three index points of the proprietary leader. On 27 July, Moonshot released the full weights under the K3 licence.
Put the two together. A laboratory creates the capability. A different company can turn that capability into infrastructure.
The frontier laboratories are spending fortunes inventing a product that companies like Hetzner will eventually sell like bandwidth.
A standard endpoint lowers switching cost without making deployments identical
Hetzner’s experiment serves Qwen3.6-35B-A3B-FP8 through an OpenAI-compatible API. A developer points a standard client at Hetzner’s base URL, supplies a key and names the model.
At protocol level, switching may be little more than changing a base URL, key and model name. At production level, it is not. Prompts, tokenisation, tool schemas, structured outputs, safety behaviour, latency, rate limits, caching, evaluations, legal terms and compliance controls must be revalidated.
Standard interfaces do not make models identical. They do not make providers interchangeable in every production deployment. They reduce the cost of comparing and switching them. That is enough for contestability.
The distinction matters because the economic threshold is not zero switching cost. It is a credible outside option. A buyer that can test several providers against the same task basket, through substantially the same interface, has leverage even when migration still takes work. A routing layer can hold more than one approved endpoint and move eligible requests between them. The incumbent no longer negotiates against the cost of rebuilding the entire application.
That is a poor foundation for durable model-layer margins. Traditional software companies try to own something difficult to replace: a workflow, a proprietary dataset, a network of users, a regulated relationship or the customer itself. A hosted open-weight model owns much less. Competing providers can serve the same weights, related weights or an adequate substitute through a familiar protocol. Competition then moves towards the mundane economics of infrastructure: hardware, electricity, cooling, utilisation, reliability, location and markup.
That is precisely the market Hetzner understands. It does not need to become a glamorous AI company. It can remain what it already is, a ruthless operator of cheap European infrastructure.
The Lidl of machine intelligence.
K3 is downloadable, but it is not cheap or simple to deploy
K3 matters because its capability is close enough to the proprietary frontier to discipline it. It does not matter because Moonshot’s own API is the cheapest. It is not.
On Artificial Analysis’s dated launch evaluation, K3 placed third on the Intelligence Index, third on GDPval-AA v2 and second on AA-Briefcase. The same evaluation estimated an average cost of $0.94 per task, close to GPT-5.6 Sol’s $1.04 and below Claude Opus 4.8’s $1.80. That is an evaluation-suite estimate, not the universal price of commercial work. K3 also used about 132 million output tokens across the suite. It remained unusually verbose and expensive relative to open-weight peers.
The release does not abolish deployment economics. It exposes them.
K3 has 2.8 trillion total parameters and activates 104 billion per token. Its mixture-of-experts architecture selects 16 of 896 experts for each token. Sparse activation weakens the relationship between total parameter count and compute per generated token. It does not abolish the cost of storing, distributing and routing across the full model.
K3 is computationally sparse but infrastructurally enormous.
Moonshot’s published deployment guidance centres on supernode configurations with at least 64 accelerators. The released model already uses quantisation-aware training with MXFP4 weights and MXFP8 activations. It supports low, high and max reasoning effort, with max as the default. Third-party quantisations have already appeared, but further compression necessarily trades deployment cost against quality, latency and behaviour.
The weights are also licensed, not placed in the public domain. The K3 licence grants broad rights to use, modify, deploy and distribute the model. It requires a model-as-a-service operator whose group revenue exceeds $20 million over any consecutive twelve months to reach a separate agreement with Moonshot before commercial use. Very large consumer products face attribution requirements. Internal use is treated more permissively.
The correct claim is therefore narrower than “anyone can serve it”. The weights can be downloaded, modified and served by a broad global ecosystem, subject to Moonshot’s licence and the capital required to deploy them.
The licence is a friction. The hardware is a barrier. Neither restores exclusive control to the laboratory that trained the model.
The important event is not that Hetzner can host K3 tomorrow morning. It is that the intelligence K3 represents has entered a cost-reduction process that Moonshot no longer controls.
Industrialisation attacks every component of the accepted-output cost
K3 has already arrived aggressively quantised. The next reductions do not depend on repeating the same trick.
Serving engineers improve kernels, interconnect utilisation, batching, caching, speculative decoding and request scheduling. Researchers distil task-relevant behaviour into smaller descendants. Hardware improves. Routers reserve the full model for the difficult tail. Each optimisation attacks a different component of the accepted-output cost.
That last phrase is the unit that matters.
Raw token price is not the test. Benchmark rank is not the test. The unit is an accepted output. Unit cost dominance begins when the all-in machine cost of producing that accepted output falls below the fully loaded human cost of producing the same thing.
The machine cost includes inference, integration, orchestration, verification, expected error and rework, compliance and liability, and residual human supervision. The comparison must hold at the required quality, latency and risk. If a human still has to reconstruct the entire output, absorb frequent failures or carry prohibitive liability, unit cost dominance has not occurred for that unit.
This definition includes the strongest objection rather than denying it. Human verification is real. Integration is expensive. Errors matter. Regulation can impose cost. The thesis applies only where the entire deployed system, including those burdens, is cheaper than human-only production.
Once that threshold is crossed, the economic value of the output need not fall. A useful report remains useful. What competition pulls down is the market price of producing it and the rent earned by whoever once supplied it under conditions of scarcity.
The market price of performing the task is pulled towards the all-in unit cost of the cheapest adequate production system.
Routing gives the volume to adequacy
A single model does not need to be excellent at everything. A production system needs to select a model that is adequate for the bounded request in front of it.
In a plausible routed architecture, a cheap model handles the large routine majority of calls while difficult cases escalate to a premium system. The exact share is workload-dependent. The economic point is that the expensive model need not receive the volume.
This changes what progress means. The largest model receives the prestige. The deployment layer receives the requests. A frontier system may remain essential for novel research, difficult coding, exceptional judgement or the highest-risk tail. None of that protects its price across routine work that a cheaper model can complete to the required standard.
K3 itself may be too large for the boring majority of workloads. Its descendants will not need to preserve every capability. A distilled or specialised model can be worse in the abstract and still better for the purchaser if it clears the acceptance threshold at lower all-in cost.
This is why adequacy threatens model-layer rent more than technical defeat does. The closed frontier can remain ahead. Its premium survives only where the gap changes the accepted output.
The model identity then recedes. Applications retain a catalogue of evaluated options. A router assigns work by price, latency, privacy, capacity and measured quality. A provider can be replaced without the customer changing the workflow it sells. Intelligence becomes a workload placed into a competitive infrastructure market.
Field update, 30 July 2026: OpenAI prices the routing mechanism
Twenty-one days after GPT-5.6 entered broad public availability, OpenAI cut Luna’s API price by 80 per cent and Terra’s by 20 per cent, while leaving the frontier Sol price unchanged. Luna moved from $1 per million input tokens and $6 per million output tokens to $0.20 and $1.20. Terra moved from $2.50 and $15 to $2 and $12. Sol remained at $5 and $30.
The asymmetry matters more than the headline. The apex retained its premium. The broad, high-volume tier was radically repriced.
OpenAI’s own guidance now begins with the required outcome and quality standard. It tells businesses to evaluate where additional intelligence materially improves the result and where lower-cost processing preserves the same quality. Its coding example uses Sol to resolve uncertainty and define the plan, then Luna to implement well-specified changes, run tests and evaluate the result.
The expensive model resolves uncertainty. The cheap model implements, tests and evaluates. The account stays with OpenAI; the routine volume leaves Sol.
That is accepted-output economics even though OpenAI does not use the term. The purchaser is told to buy the cheapest system that clears the quality threshold at each stage of the workflow, not the most capable model for every token.
OpenAI says efficiency made the reductions possible. Better routing, serving software and context management produce more useful work from the same compute. It says Sol-assisted kernel work reduced end-to-end serving cost by 20 per cent, while related experiments improved token-generation efficiency by more than 15 per cent. Axios reports that cheaper Chinese open-weight models have also increased pressure on OpenAI and Anthropic to justify their higher costs. The announcement does not identify how much of the pass-through came from efficiency and how much from competition.
Nothing here proves that K3 caused a particular percentage point of the cut. It shows what the credible outside option mechanism can look like. The incumbent keeps the customer while repricing its high-volume tier as if defection were possible.
OpenAI also claims that Luna now delivers performance comparable to models that were frontier-class a year ago at roughly six cents on the dollar per task and nearly nine times the speed. That is a vendor comparison, not an independent or controlled time series. Taken only as an illustration, six cents represents just over four cost halvings in one year, or an economic half-life of about 90 days.
This does not establish the credit event in Part III. Efficiency may protect OpenAI’s margins, and lower prices may stimulate enough demand to increase total compute use. It shows one transmission route directly: OpenAI has repriced its adequate tier and is explicitly positioning Sol’s premium around the stages where additional intelligence changes the accepted result.
Two routes to commodity-priced cognition
The OpenAI cut also exposes a fork inside the delivery mechanism.
The first route is competitive commoditisation. Open weights are served by many infrastructure providers. Standard interfaces and routing layers expose price differences. Hosts compete on electricity, hardware, utilisation, reliability and markup. Efficiency gains and written-down capacity can reach customers through defection between providers.
This is the Hetzner route.
The second route is concentrated commoditisation. A vertically integrated laboratory or hyperscaler uses scale, engineering efficiency, cloud bundling or strategic pricing to undercut independent providers. Routine cognition becomes extremely cheap while the serving market remains concentrated.
This is the Luna route.
Both routes can destroy durable scarcity pricing for routine cognition. Only the first implies a dispersed hosting market. Neither, by itself, proves that a lower acquisition cost for distressed capacity will reach customers after a credit event.
That makes the OpenAI cut double-edged. It strengthens the claim that adequate cognition is becoming cheap. It weakens any assumption that independent hosts must be the organisations making it cheap. It may even intensify pressure on merchant GPU providers that cannot match an integrated provider’s rate, leaving the surviving serving layer more concentrated after the distress.
The title of this essay remains a prediction, not an established industrial structure. A company like Hetzner may eventually sell intelligence like bandwidth. A frontier laboratory may instead preserve the account and sell its own volume tier at bandwidth prices. The labour mechanism can operate through either route. Part III’s fire-sale mechanism requires the additional condition of competitive pass-through after restructuring.
Open weights do not guarantee cheap inference
The transmission remains conditional.
Competition pushes price towards underlying cost only when enough providers can serve adequate models, licensing does not block entry, and effective supply outruns paid workload demand. If infrastructure consolidates into a few hosts, if power and accelerator scarcity dominate the stack, or if demand absorbs every efficiency gain, model-layer rent can die while infrastructure rent survives.
Contestability is the necessary condition. Competitive pass-through is the transmission.
This is also why the argument does not depend on Hetzner personally winning. Hetzner’s public GPU estate may remain too small for frontier-scale serving. Its experiment may end. If Google, Amazon or a specialist inference company owns the cheap layer instead, the model premium can still disappear. The surplus has merely moved to a different landlord.
That outcome would matter. Concentrated infrastructure could keep customer prices well above physical cost. Scarce power, land, accelerators, high-bandwidth interconnect and reliable capacity can support durable rents even while raw intelligence loses its scarcity premium.
The claim is not that every layer becomes perfectly competitive. It is that a model owner cannot rely on permanent scarcity pricing once an adequate alternative can be deployed, evaluated and routed by other firms.
K3 does not prove that pass-through will occur. Hetzner does not prove that supply will outrun demand. Together they make the industrial pathway visible.
Tasks disappear before jobs do
Unit cost dominance operates on outputs. Organisations employ people in jobs. The bridge between them is task decomposition.
Suppose a worker performs twenty recurring tasks. No single model can perform the whole job. One drafts the reports. One reconciles the spreadsheet. One answers routine email. One inspects screenshots. An orchestration layer moves information between them. A smaller number of people supervise exceptions and carry formal responsibility.
The organisation does not need a perfect digital employee. It needs enough accepted output to require fewer producers.
The first quarter looks like assistance. The second looks like productivity. The headcount decision arrives later.
There is no dramatic moment when a machine becomes a perfect employee. There is simply a budget meeting.
The mistake is to wait for one model to perform every element of an occupation. The economically relevant threshold arrives earlier, when a collection of systems produces enough of the output that the organisation needs fewer people. Whole jobs survive on an organisation chart while the volume of human production inside them contracts.
This is how displacement can begin without a mass redundancy announcement. Hiring slows. Junior roles are not reopened. Contractors replace permanent staff. A five-person team supervises work that once required twenty. Headline unemployment can remain calm while the labour market stops absorbing the next cohort.
Sorites prevents a usable boundary
No single step looks like replacement. A tool drafts one report, then checks one spreadsheet, then handles one client queue. Headcount falls at the next budget round.
Assistance and replacement differ by degree, not by a boundary anyone can verify. That Sorites ambiguity makes restraint unenforceable. Every firm can call its own adoption assistance while treating everyone else’s restraint as an opportunity. Nobody can say when defection began, so everybody defects before anybody agrees that replacement has begun.
The distinction remains meaningful at the endpoints. A person writing unaided is producing. A largely automated workflow with one exception handler has replaced most production labour. The problem lies between them. There is no operational line at which assistance becomes replacement, even though the cumulative change is unmistakable.
A rule can require human oversight. It must then define how much attention, at what frequency and with what authority. A firm can satisfy the label through sampling, escalation or formal approval while continuing to reduce the labour content of each accepted output. The human remains visible. Productive necessity recedes underneath.
Sorites is not a claim that regulation is impossible. It is a claim about a particular regulatory defence. Any regime that must identify the precise moment at which assistance becomes replacement is trying to govern a boundary the workflow does not contain.
Sorites is not the net-displacement theorem either. It explains why substitution can proceed without a defensible moment at which assistance becomes replacement. Whether new work absorbs the displaced wage income is a separate empirical question.
New tasks do not necessarily arrive already automated. Digitally expressed tasks increasingly arrive contestable from birth. The same general systems already operate the browser, document, spreadsheet, inbox, dashboard and code editor through which a new workflow is likely to be performed. Human labour may still win many of those contests. It no longer receives an exclusive window by default while specialised machinery is designed around the new task.
The non-absorption claim fails if new human-complementary work restores displaced wage income at comparable scale and speed. It is supported if output and machine-completed work rise while exposed hiring, junior intake, human hours and labour income do not.
The Multiplayer Prisoner’s Dilemma makes adoption compulsory at the system level
The same game repeats at every level.
Laboratories cut prices because rivals might. Hosts optimise because other hosts will. Firms automate because competitors are reducing labour cost. Workers adopt the tools because refusing makes them individually less employable, even though universal adoption reduces the number of workers required.
Each move is locally rational and collectively accelerative.
A laboratory that preserves margin while a near-frontier rival cuts price loses volume. A host that declines to optimise leaves utilisation and customers to another host. A firm that keeps expensive human production for substitutable output while competitors cross unit cost dominance loses on price or margin. A worker who refuses augmentation is compared with one who can supervise more output.
No actor needs to desire the aggregate result. Each needs only to avoid being the actor who pays for restraint while others defect.
Unit cost dominance supplies the payoff. Sorites prevents a usable boundary. The Multiplayer Prisoner’s Dilemma makes adoption compulsory at the system level wherever the cost advantage is material.
That does not mean every individual firm must automate. Some can sustain a premium around human service, trust, regulation, craft or luxury positioning. Those firms become residual niches inside the new cost structure. They do not remove the competitive pressure on the substitutable volume.
This is the missing link between cheaper inference and labour displacement. Capability does not enter the economy because every institution has agreed on its social value. It enters because the cost advantage can be captured privately while the displacement cost is shared.
Regulation loses its single upstream choke point
When the most capable systems are controlled by a few laboratories, regulation has a convenient object. Governments can impose reporting duties, mandate evaluations, restrict particular exports and negotiate with named executives.
Downloadable weights fragment that object across the model developer, host, application provider, router, deployer and regulated institution. The obligations do not vanish. They multiply.
A European company can run K3 on European infrastructure without sending its prompts to Moonshot or a United States cloud. That reduces supplier dependence and some cross-border-transfer exposure. It does not remove the deployment from law.
The GDPR is technology-neutral and still follows the processing of personal data. The AI Act allocates obligations to providers and deployers for specified uses. Sector rules still apply. Moonshot’s licence still governs the weights.
Self-hosting changes the supply chain. It does not make the arithmetic legally invisible.
Regulation can govern employers, hospitals, banks, insurers, public bodies and harmful uses. It can impose liability and documentation duties. What it cannot easily do is restore one upstream choke point once the capability is distributed across models, hosts and jurisdictions.
The delivery layer turns a model event into an economic event
Hetzner may withdraw the endpoint. K3 may remain expensive to self-host. A stronger proprietary model may appear next month. None of those possibilities answers the mechanism.
The mechanism does not require Hetzner to serve K3. It requires near-frontier capability to become an input available to providers that compete on deployment. It does not require providers to be identical. It requires switching and comparison to become cheap enough to discipline price. It does not require zero human involvement. It requires the all-in cost of an accepted output to fall below the fully loaded cost of equivalent human production.
When those conditions hold, the model stops behaving like a premium product and starts behaving like an infrastructure component. The frontier laboratory can preserve company value by owning the workflow, customer, device, proprietary data or regulated relationship. It cannot assume that raw intelligence will continue to carry the old scarcity premium.
For businesses, that is the attraction. Capability becomes cheaper to acquire and easier to embed.
For labour, that is the mechanism.
Workers will keep being told the current model is not quite good enough.
Then, one ordinary quarter, it will be.
What would falsify this mechanism
Fix a basket of production tasks, quality thresholds and providers before measuring it. This account weakens if standard interfaces do not reduce comparison and migration costs, if neither competitive hosting nor integrated undercutting lowers the all-in cost per accepted output, or if verification, error, compliance, liability and supervision keep that cost above equivalent human production. Persistent quality-adjusted model premiums would refute the pricing claim.
The labour transmission weakens if machine-completed accepted outputs rise without human hours, junior intake or the wage share deteriorating in exposed sectors. The stronger non-absorption claim fails if new human-complementary work restores displaced wage income at comparable scale and speed.
Sources
- Jonas Scholz, Hetzner Inference: First Look, including the experiment’s single model, OpenAI-compatible endpoint, lack of billing, SLA and production guarantee, and July 23 test.
- Artificial Analysis, Kimi K3 achieves number three in the Artificial Analysis Intelligence Index, 17 July 2026 launch snapshot, benchmark positions, measured task cost, token usage and API pricing.
- Moonshot AI, Kimi K3 Tech Blog, model architecture, launch positioning, multimodality and context window.
- Moonshot AI, Kimi K3 model card, released weights, activated parameters, quantisation format, deployment engines and reasoning-effort controls.
- Moonshot AI, Kimi K3 Licence, commercial model-as-a-service threshold, attribution requirements and internal-use exception.
- European Commission, Data protection explained, including the GDPR’s technology-neutral application to personal-data processing.
- European Commission, AI Act regulatory framework, provider and deployer obligations under the risk-based regime.
- OpenAI, Advancing the price-performance frontier with GPT-5.6, 30 July 2026. Price reductions, workflow-routing guidance, cost-per-task comparison and serving-efficiency claims.
- OpenAI, GPT-5.6: Frontier intelligence that scales with your ambition, 9 July 2026. Original public-availability date and launch API prices for Sol, Terra and Luna.
- Axios, OpenAI cuts GPT-5.6 prices, 30 July 2026. Price-sensitive customer context and reported pressure from cheaper Chinese open-weight models.
Next in the trilogy: The Margin Call
The model layer has become contestable and the delivery layer is being industrialised. But the delivery layer was financed when both looked scarce. Who lent against that assumption?
