We are about to see a major shift in how AI infrastructure is valued.

Right now, most companies rent intelligence from a small number of providers. The experience feels cheap because the largest AI companies are willing to spend enormous amounts of money building infrastructure, acquiring users, and winning market share.

That price is not guaranteed forever.

API pricing can change. Rate limits can tighten. Models can disappear. Providers can decide which use cases they support and which ones they do not. If your company depends entirely on someone else’s inference layer, you do not own the economics of your AI workforce.

At the same time, local models are reaching an inflection point.

Qwen3.8-27B is a good example. This is a 27-billion-parameter model with native vision capabilities and a 262,144-token context window. According to Qwen’s published evaluations, it scored 61.7 on SWE-bench Pro, compared with 53.4 for Opus 4.6 Max. It also scored 70.7 on CoWorkBench, 90.3 on LiveCodeBench v6, and 84.3 on OSWorld-Verified. These are vendor-reported benchmarks, so they should be treated accordingly, but the direction is difficult to ignore. A model small enough to deploy yourself is becoming competitive with frontier hosted models on real agentic work.

The weights are free, and the hardware barrier is collapsing. A Mac mini or Mac Studio can sit quietly on a desk and run local inference all day without a server rack, special cooling, or an infrastructure team. You still buy the machine and use electricity, but this is starting to look much more like owning an appliance than operating a data center.

The important difference is that you own the infrastructure.

Once the hardware is installed, the marginal cost of keeping an agent running becomes predictable. You can run models around the clock without counting every token. You can process sensitive data without sending it to an outside provider. You can keep the same model available instead of rebuilding your workflows every time an API changes.

This changes how we should think about GPUs.

Everybody looks at giant companies hoarding GPUs and assumes they are trapped in a reckless infrastructure race. Maybe some of them are. But those GPUs are not only useful for training the next frontier model.

They can power enormous AI workforces.

A warehouse full of GPUs could run persistent agents that monitor operations, write software, investigate problems, process documents, and coordinate work every hour of the day. The hardware does not need to produce artificial general intelligence to become extraordinarily valuable. It only needs to run capable models reliably at scale.

We are basically doing this already. The models are imperfect, but the direction is obvious. In another year or two, today’s expensive GPU clusters may look less like speculative infrastructure and more like factories for digital labor.

That is why I am buying as many GPUs as I reasonably can.

I do not want every important system we build to depend on an API price, a rate limit, or another company’s roadmap. Hosted frontier models will still matter. We will use them when they are the best tool for the job. But they should be one part of the infrastructure, not the infrastructure itself.

I am building a sovereign inference layer.

A private layer where we control the models, the hardware, and the operating economics. It can route difficult work to frontier APIs when necessary, while keeping high-volume and sensitive workloads on infrastructure we own.

The next AI advantage will not come from having access to the same model as everyone else.

It will come from owning the system that keeps working when the economics change.