Who Gets to Use Intelligence?

· 12 min read ai infrastructure theory

Who Gets to Use Intelligence?

You can download one of the best open-weight models in the world for free.

Moonshot released the full Kimi K3 weights on July 27: about 1.56TB across 96 files, already in native MXFP4. Divide that by the RTX 5090’s 32GB of memory and you get forty-nine cards, nearly $98,000. That is only the memory arithmetic and doesn’t account for the kv cache and context window. You still need servers, system memory, networking, cooling, and a plan for coordinating forty-nine consumer GPUs with no NVLink.

A large folder full of weights is technically sovereignty in the same way owning a jet engine is technically transportation.

The example separates three questions usually compressed into the word open:

In practice, open weights still leave most people dependent on someone else’s infrastructure.

I’m part of the TxnLab team which is building ZeroSignal, one attempt at addressing this problem. Discount accordingly.

Abundance on lease

AI access feels abundant. Open a tab and a model answers. Add a card and you can call an API. Download weights and, hardware permitting, run one yourself.

That stack is rented. The company serving the model can change the price, retire the model, rewrite its policies, block an account, or stop serving a country. Governments decide where advanced chips are sold and which systems can be offered in a market. Even a model described as open may require an account, an approval, or a license that falls short of open source.

The internet already ran this experiment with payments. PayPal, Visa, and Mastercard cut off WikiLeaks in 2010. Visa and Mastercard stopped processing Pornhub payments in 2020. A handful of infrastructure companies turned their policy into the market’s policy.

The same concentration is forming around intelligence.

We have made it dramatically easier to use and surprisingly easy to revoke.

A Claude account suspension notice reading "Your access to Claude is on hold," explaining that an organization the user belongs to has been paused because of unusual activity.

One such revocation, posted to X.

Sovereignty is measured in exits

Sovereign AI is usually discussed at the national level: domestic data centers, national models, independence from foreign suppliers. But a country can own an impressive amount of infrastructure while its citizens and businesses remain dependent on a few companies and their policies.

For a person or business, sovereignty becomes concrete under failure. Your provider closes your account on Friday. Can you serve traffic by Monday? Do you lose access to your data? Can you easily switch?

Ownership is one exit. It is not the only one. Interoperable clients, competing suppliers, self-hosted software, and settlement that does not depend on a custodial account can create exit without requiring every user to own their own hardware.

So a practical definition: sovereignty is the number and quality of alternatives available when the default path fails. What matters is not how independent a system looks on an ordinary day, but how much permission is still required on the bad day.

Most people will keep renting compute. Most businesses should. Hardware is expensive, demand is bursty, and a competent provider beats an improvised rack. With substitutes, that is a healthy vendor relationship. A dependency without substitutes is a dangerous gate for an individual and especially for a business.

Open weights and expensive freedom

Open-weight models are the strongest counterweight to a future controlled by a few closed labs. They let researchers inspect systems, developers adapt them, businesses deploy on their own terms, and communities keep useful models alive.

The word open still carries several promises. The Open Source Initiative’s definition requires the freedom to use, study, modify, and share a system, along with the data information. A weights download with a commercial restriction or acceptable-use policy can still be enormously valuable; open-weight is simply the more precise term. Kimi K3 ships under Moonshot’s own license, not a standard open-source one.

Licensing is only the first constraint. K3 has 2.8 trillion parameters, 104 billion active per token. Sparsity cuts the computation per token, but storage still holds all 896 experts. SGLang’s recipes run from eight 288GB B300s on a single node to thirty-two H100s; Moonshot recommends production supernodes of 64 GPUs or more. Within days of release, Unsloth’s dynamic 1-bit GGUF cut the checkpoint from 1.56TB to 594GB. Still not a home machine, and compression trades size for quality.

Smaller local models are genuinely useful and improving. When one fits the task, it offers the strongest privacy and control available. But local inference will not cover every model a person wants, and most people/businesses will not buy hardware for every peak in demand.

Open weights give people rights: to inspect, modify, deploy, preserve, and compete. Freedom with a six-figure cost is still freedom, but with a peculiar distribution. Call it “free”.

Inference is where access becomes real

A model becomes useful only when a request can reach running hardware and a response can return, ideally as fast as possible and without being quantized. Four control points sit on that path:

Effective access = model permission × compute availability × endpoint permission × transaction permission.

Open weights solve the first control point. Local inference solves the rest for users with hardware and expertise. A hosted API bundles all four under one provider. Bundling is convenient. It’s also entirely dependent: the provider becomes model selector, infrastructure operator, payment processor, policy maker, and judge of whether the customer remains welcome.

The industry has spent enormous energy making models capable. Much less has gone into making access resilient.

OpenRouter is the biggest player in inference aggregation. One OpenAI-compatible endpoint, hundreds of models, provider routing, fallbacks. It proved that portability has massive value. But the portability still runs through a centralized intermediary: an OpenRouter key, prepaid OpenRouter credits, requests through openrouter.ai. If OpenRouter changes policy, is acquired, or goes offline, every customer replaces the endpoint and billing path at once. That is an aggregator’s normal architecture. It’s also a useful boundary: a universal interface can make providers substitutable while the interface operator remains hard to substitute.

ZeroSignal is our attempt at the third question from the opening: if you cannot run the model, can you transact with someone who can? It allocates responsibilities differently than OpenRouter. Operator selection happens in a user’s local proxy, not a centralized one. Discovery comes from a public on-chain registry. Payment settles per request from a self-custodial account. Anyone meeting public rules can join as an operator. TxnLab does not approve each request, select each provider, or hold customer balances.

Gatekeepers are useful

The strongest argument for centralized providers is that they do many necessary jobs at once: screen operators, enforce quality, handle abuse, comply with law, absorb fraud, and give users someone to blame. Some gating prevents real harm.

This is why decentralization is difficult and all services stop short. The integrated provider bundles useful coordination with concentrated veto power. Separating them preserves more exit, but it also separates responsibility. An indiscriminate network could become an efficient way to dump the costs of abuse on everyone else. A bad actor refused by nine operators may find a tenth. No design gives ordinary users reliable exit while preserving a universal kill switch.

Strong encryption made a similar tradeoff. Criminals acquired private communication along with everyone else, and society has mostly kept encryption anyway. The analogy has limits, since models can generate actions at scale. But the shared question is the same uncomfortable one: how much power should belong to whoever controls the switch?

Every road in a dark concrete maze curving toward a single glowing sphere at the center.

ZeroSignal should limit the damage it could help create. Per-request payment makes bulk abuse cost money. Operators set rate limits, choose the models they serve, apply their own policies, and refuse requests. Users exclude operators they do not trust. Jurisdictions enforce law against users and operators within their reach.

Network-level moderation has a narrower role: fraudulent receipts, non-delivery, malware, denial-of-service, and attacks on the registry or settlement. Those rules should be public and reviewable. Content policy remains attributable to operators and the jurisdictions that impose it.

Neutral rails can be tested

Crypto has a useful term for the goal: credible neutrality. Rules that are visible and predictable, hard to bend for a preferred participant, and open to newcomers on the same terms.

For inference, neutrality has observable consequences: users choose among models and operators; operators compete on price, performance, location, and privacy; new operators enter without negotiating a private platform agreement; clients move without rewriting their applications; settlement follows known rules; and the market keeps functioning when one operator, jurisdiction, or the founding company leaves.

The last test is the severe one. A marketplace can have hundreds of sellers and remain centralized if one company controls admission, ranking, custody, and settlement. Many “decentralized systems” are not, and users usually only learn this once their exit is impaired.

Neutral rails remain subject to law. Operators choose what they serve, model licenses apply, and governments regulate data centers, chips, payments, and network infrastructure. Decentralization changes how far those constraints reach. A restriction on one operator or jurisdiction doesn’t automatically become a restriction on the entire market.

One ZeroSignal request

The easiest way to understand the architecture is to follow one request. A coding tool sends an OpenAI-compatible call to a proxy on the user’s own machine. The proxy reads the public on-chain registry, filters offers by the user’s preferences on price, performance, capability, and privacy, and selects an operator locally. It selects a relay separately, so the operator never receives the user’s IP address. The request travels encrypted to a short-lived key held only by the chosen node. The response returns with a signed receipt, and a payment reserved in escrow settles for the actual charge. To the coding tool, all of this is an ordinary localhost endpoint. In the browser, a passkey is the account: no email, password, or API key.

The path of one ZeroSignal request A coding tool calls a local proxy on the user's own machine. The proxy reads the on-chain registry and independently selects a relay and an operator. The request passes through the relay, which sees the IP address but not the content, to the operator, which sees the prompt but not the sender. A signed receipt returns and escrow settles the actual charge. coding tool localhost local proxy your machine relay sees IP only operator sees prompt only on-chain registry and escrow discovery, then settlement encrypted to a short-lived key
The proxy picks the relay and the operator independently, so no single party sees both who you are and what you asked.

On the supply side, anyone meeting the public stake and registry rules can run a node in front of vLLM, llama.cpp, or another backend, post models and prices, and start serving. There is no application.

Where trust still sits

The privacy model is a decomposition of trust, not its abolition.

PartySeesNever sees
RelayYour IP addressThe content of the request
OperatorThe plaintext promptWho sent it
ChainAddresses and settlementPrompts or identity
Your local clientAll of itNothing is hidden from it

No single row holds the whole picture. If an operator forwards a request to a proprietary upstream model, that provider’s terms enter the trust model too.

Today, the operator can read the prompt while the model runs. The architecture we are working toward uses trusted execution environments and remote attestation: before releasing a request, a client verifies that the expected software is running inside confidential hardware, eventually including the model itself. Receipts already prove which operator charged what and returned which bytes; attestation can bind that to a measured execution environment.

Why trust a random operator before TEE lands? The correct default is skepticism. The present advantage is that no participant holds the whole relationship graph. A conventional provider can correlate identity, IP address, payment, history, and prompts in one database. ZeroSignal splits those observations among the relay, the operator, the settlement layer, and the local client, and the user can change operators request by request. It provides strong anonymity, not yet complete privacy.

Privacy and access go hand in hand. A service that knows your identity, payment, location, and prompts has many ways to profile and exclude you. The same argument runs from the seller’s side. Anyone with idle or purpose-built compute can register, set prices, and earn by serving inference.

Neutrality has to be earned

It would be easy to put a registry and payments on-chain, add encryption, repeat the word sovereign, and draw the inevitable globe covered in glowing lines. The diagram would look excellent. The control points would remain.

ZeroSignal is early. TxnLab builds the clients, node software, and contracts. The proxy and operator distributions are only partially open-source at the moment. Production TEE and attestation are unfinished. Early supply is concentrated to a few operators. Algorand and USDC are dependencies. And real economic questions remain: will reliable operators show up, how does a new operator establish trust, and what keeps a few large operators from dominating the network?

Our job is to keep removing places where one party, including us, can control the outcome: more independent operators, client-side choice, public rules, portable interfaces, reproducible software, production attestation, and governance that relies less on TxnLab.

By December 31, 2026, the complete local proxy, client, and operator node will be published under a source-available license with reproducible build and deployment instructions.

Access needs somewhere else to go

Open models matter. Local inference matters. National AI infrastructure matters. Each makes the future less dependent on a handful of labs. Together they still leave a gap between the right to use intelligence and the practical ability to do so.

The question is not whether companies may deny access, operators may refuse requests, or governments may enforce law. They will, and sometimes they should. The infrastructure question is whether any one refusal becomes universal for you.

Intelligence is becoming a basic input into work, education, science, and speech. The people who control access to it will acquire power even when they exercise that power responsibly. Durable access does not require a world without rules. It requires a world in which no single rule-maker controls every road.

That is the standard ZeroSignal should be judged against: when one route closes, is there still another route that works?