Thesis

Every device becomes an inference endpoint.

The internet connected documents, then people, then software services. The next 20 years connect intelligence to everything with a sensor, screen, motor, workflow, account, or API. The unit of demand is no longer a page view or a search query. It is an inference event.

An inference endpoint is any device or application that sends context to a model and receives a prediction, plan, action, summary, decision, or control signal back. Phones, PCs, cars, robots, cameras, medical devices, factory machines, wearables, and enterprise apps all become attached to these endpoints.

Value accrual

2026-2046

Scarce inputs

Power, chips, HBM, packaging, optics, land, cooling

Endpoint distribution

Phones, PCs, vehicles, machines, wearables, SaaS workflows

Trusted execution

Identity, routing, memory, policy, security, evaluation, settlement

This is the core thesis: inference becomes a metered utility for cognition, and value accrues wherever the stack controls scarce inputs, endpoint distribution, or trusted execution.

The argument

Chatbots were the browser moment. Agents are the endpoint moment.

A chatbot answers one human prompt. An agent decomposes a goal, reads systems, calls tools, runs code, checks itself, retries failures, asks other models to verify the work, updates memory, and keeps monitoring after the user leaves. One request can become dozens, hundreds, or thousands of inference events. Then that pattern gets embedded into every device and workflow.

Demand compounds

More capable models do not simply replace cheap calls. They unlock harder tasks, longer context, more verification, and more autonomous execution.

All apps become AI apps

CRMs, security tools, finance systems, developer tools, analytics products, and support desks will serve models continuously instead of adding a single chat box.

Control becomes mandatory

Enterprises cannot let agents spend, retry, retrieve data, or take actions without budgets, permissions, evaluation, security, and auditability.

10-20 year view

The architecture moves from cloud AI to hybrid endpoint intelligence.

2026-2030

AI factories become the new industrial base

The first phase is centralized. Frontier labs, hyperscalers, neoclouds, and sovereign buyers race to secure accelerators, HBM, networking, land, substations, cooling, and power. Inference begins as a cloud workload because the best models are too large, too expensive, and too fast-moving to live at the edge.

2030-2036

Every premium device attaches to inference

The second phase is hybrid. Phones, PCs, glasses, cars, cameras, robots, and industrial machines ship with NPUs for local small-model work, but still call larger cloud models for planning, reasoning, verification, memory, and specialized tasks. The winning architecture is not cloud or edge. It is routing between both.

2036-2046

Inference becomes the default interface to the physical world

The third phase is ambient. Devices stop being passive clients and become model-mediated actors. A camera does not merely stream video; it emits structured understanding. A car does not merely navigate; it negotiates with maps, fleets, insurers, chargers, service networks, and regulators. A factory machine does not merely report telemetry; it asks for maintenance, orders parts, and optimizes throughput.

Research anchors

The endpoint attach rate is the number to watch.

The market is already moving from model demos to device distribution. AI phones, AI PCs, connected machines, and cloud AI factories are all different expressions of the same underlying shift: more endpoints attached to inference.

AI phones by 2028

912M

Generative AI smartphones are forecast to become a majority of the phone market, turning the default consumer device into an inference endpoint.

AI PCs by 2028

166M

AI-capable PCs are expected to move from premium feature to mainstream replacement cycle as local NPUs become standard.

IoT connections by 2030

37.4B

Connected sensors, machines, vehicles, cameras, meters, and industrial systems become raw surfaces for semantic inference.

Inference cost decline

>90%

By 2030, frontier-scale inference costs are expected to fall dramatically, which should expand usage rather than cap demand.

Global data center power

945 TWh

Global data center electricity demand is projected to more than double by 2030, with AI as the largest driver.

Cloud capex shock

$830B

Estimated 2026 capex for the top nine cloud service providers sits near this level.

Endpoint surfaces

Inference will attach wherever context meets action.

The endpoint is not just a data center API. It is the place where a device, workflow, or machine converts raw context into a decision. The long-run attach surface is much larger than chat.

Phones and PCs

The personal agent lives here. Local inference handles privacy, latency, summarization, search, and UI control. Cloud inference handles deep reasoning, long memory, and expensive orchestration.

Cars and robots

Physical AI turns inference into a safety-critical workload. The endpoint must perceive, plan, verify, and act under latency and reliability constraints.

Cameras and sensors

Raw telemetry becomes semantic telemetry. Instead of sending everything upstream, endpoints classify, summarize, detect anomalies, and trigger actions.

Enterprise software

Every SaaS workflow becomes an agent workflow. CRM, ERP, security, finance, support, legal, and analytics systems attach to model endpoints for continuous work.

Industrial machines

Factories, energy assets, logistics networks, farms, and utilities use inference to predict, optimize, dispatch, and coordinate real-world operations.

Wearables and glasses

The interface moves from app icons to ambient context. Always-on devices need local inference for privacy and cloud inference for complex reasoning.

The 1000x question

The 1000x company is not another model. It is the inference endpoint network.

Mature public AI infrastructure can still compound, but it is unlikely to produce a venture-style 1000x from today's scale. The 1000x outcome has to start where the market is still structurally unresolved: the runtime layer between billions of endpoints and thousands of models.

That layer decides which model runs, where it runs, what context it receives, what tools it can call, which identity it acts under, what it is allowed to spend, how its work is evaluated, and how the economic value is attributed. In the web era, the default layers were search, payments, identity, cloud delivery, and app distribution. In the inference era, the default layer is the endpoint network that routes, secures, meters, and governs intelligence.

The investable point of view: deploy capital toward companies that can become neutral inference infrastructure for every device and agent, not toward companies that only wrap one model, sell one workflow, or rent one generation of compute.

Endpoint attach

The company must sit at the moment a device, agent, or workflow calls inference. If it is downstream of the call, it is analytics. If it is upstream of the call, it can become infrastructure.

Model independence

The company cannot depend on one foundation model winning. It should benefit when enterprises use many models, many clouds, local NPUs, and specialized inference providers.

Runtime trust

The company must decide what an agent is allowed to know, spend, retrieve, and do before the action happens. Post-hoc dashboards are not enough.

Economic metering

The company must connect inference consumption to useful work: cost per task, outcome, user, device, workflow, and agent. Whoever meters the work can price the value.

Developer distribution

The wedge must spread through SDKs, APIs, protocols, observability hooks, security controls, or device integrations. The category winner needs default distribution, not only a better dashboard.

Data flywheel

The company should improve as it sees more endpoint behavior: routing decisions, failure modes, permissions, latency, cost, quality, and outcome data across production workloads.

Companies positioned to benefit

The AI stack rewards scarcity first.

Tier 1

Compute, memory, and packaging

NVIDIA, AMD, Broadcom, Marvell, TSMC, SK Hynix, Micron, Samsung

The first value pool is still close to the silicon. Accelerators, custom XPUs, networking ASICs, optical links, advanced packaging, and high-bandwidth memory sit on the critical path for every AI factory and high-volume inference endpoint. The next 10 years reward companies that increase tokens per watt, memory bandwidth per package, and useful inference per dollar.

Tier 2

AI factories and cloud capacity

AWS, Microsoft Azure, Google Cloud, Oracle Cloud, CoreWeave, Lambda, Crusoe, Nebius, IREN, Applied Digital

Centralized AI factories train frontier models and serve the hardest inference. Clouds and neoclouds rent scarce accelerated compute into a market where demand can outrun capacity for years. The reward can be enormous, but this is a depreciation-heavy business where utilization, customer concentration, cost of capital, power access, and chip obsolescence decide whether revenue becomes durable equity value.

Tier 3

Networking, power, cooling, and sites

Arista, Cisco, Vertiv, Eaton, Schneider Electric, GE Vernova, Siemens Energy, Equinix, Digital Realty, Corning, Coherent, Lumentum

Inference at scale is not just GPUs. It is switching, optics, transformers, substations, switchgear, backup power, cooling, racks, fiber, land, interconnection, and construction execution. As clusters grow from thousands of accelerators to campus-scale AI factories, the surrounding physical layer can become the bottleneck and sometimes the better risk-adjusted bet.

Tier 4

Device and endpoint ecosystems

Apple, Qualcomm, Arm, MediaTek, Samsung, Microsoft, Google, Tesla, Mobileye, Ambarella, Sony Semiconductor

The second decade shifts more inference toward endpoints. Phones, PCs, glasses, cars, cameras, industrial machines, medical devices, and robots will all need local NPUs plus persistent access to cloud inference. The endpoint owner controls distribution, privacy boundaries, latency, default assistants, and which model calls happen locally versus remotely.

Tier 5

Software control points

Highest-conviction categories: edge networks, enterprise data clouds, systems of record, identity, security, observability, routing, evaluation, and agent governance.

This layer has the best capital intensity but also the highest displacement risk. Generic wrappers, vector databases, dashboards, and narrow agent tools can be rebuilt in-house or absorbed by model providers. The durable winners need a control point that is hard to replace: distribution at the edge, ownership of enterprise data, workflow system-of-record status, identity and permission enforcement, security telemetry, or runtime governance before an agent acts.

Full stack

The inference economy is an energy-to-application stack.

The mistake is to analyze AI as one market. It is a chain of bottlenecks. The capital prize migrates as each constraint is solved: first power and chips, then memory and networking, then endpoint distribution, then data rights, trust, routing, and applications.

Energy and grid

Power becomes the first physical constraint: generation, transmission, substations, transformers, switchgear, backup power, grid queues, and behind-the-meter energy. Without watts, there are no tokens.

Data centers and cooling

Land, permits, water, liquid cooling, high-density racks, construction execution, and utilization determine whether AI factories earn their cost of capital.

Chips and packaging

Accelerators, custom ASICs, CPUs, advanced packaging, interposers, chiplets, and foundry capacity decide how much useful inference can be produced per watt and per dollar.

Memory and storage

HBM, DRAM, NAND, KV cache, vector stores, and retrieval systems become central because agents need long context, persistent memory, and repeated access to state.

Networking and optics

Scale-out inference requires fabrics, switches, optics, fiber, coherent links, and edge delivery. The network is what turns isolated chips into an AI factory.

Data rights and context

The scarce data shifts from public internet text to private enterprise state, real-world sensor history, workflow traces, user memory, and verified outcome data.

Models and inference providers

Foundation models, small models, vertical models, open models, and specialized inference clouds compete to serve tasks. Model quality matters, but model plurality is the base case.

Applications and workflows

The application layer captures user intent and business process ownership. The best apps will not add AI; they will become agentic operating surfaces.

Security and governance

Agent identity, permissions, audit, policy, data loss prevention, model risk, and human approval become mandatory once inference can take action.

Capital deployment

1000x capital requires a software-like monopoly on top of a physical supercycle.

The physical layers are massive and necessary. But the largest upside from capital deployed today is more likely to come from a new default layer that can sit across devices, models, applications, data, and clouds. The goal is not to own every data center. It is to become the route through which inference work is trusted, priced, and executed.

Power, memory, packaging, optics

Best risk-adjusted infrastructure

These layers are already bottlenecks and can compound with the buildout. They are essential, but many winners are large, capital-intensive, cyclical, or already public. They can be excellent investments without being realistic 1000x outcomes from today.

Undifferentiated GPU rental

Most dangerous trap

Buying depreciating compute with borrowed capital is not enough. If the provider lacks power advantage, customer durability, utilization discipline, financing edge, or software differentiation, gross demand can still destroy equity value.

Inference endpoint network

Highest 1000x asymmetry

The most asymmetric capital should target the neutral layer that connects endpoints to models and governs identity, routing, memory, spend, security, evaluation, and settlement. This can have software margins, network effects, and transaction volume across every device and application.

Proprietary action data

Second 1000x candidate

The next scarce dataset is not scraped text. It is verified action data: what agents tried, what failed, what succeeded, which tool calls mattered, which human approvals were needed, and what economic outcome followed.

Investment implication

The long-run winners either own scarce inputs, endpoint distribution, or control layers.

The first wave rewards scarce physical inputs: accelerators, HBM, packaging, optics, power, cooling, land, and cloud capacity. These businesses are easiest to underwrite because demand shows up in revenue, capex, lead times, and power constraints.

The second wave rewards endpoint distribution. If the phone, PC, car, robot, headset, camera, or enterprise app is where inference starts, then default assistants, local NPUs, operating systems, app stores, identity layers, and device clouds become strategic control points.

The third wave rewards governance and routing. As inference becomes cheap enough to use everywhere, demand does not disappear. It explodes. The bottleneck becomes deciding which model runs where, which endpoint is trusted, what data can be used, what actions are allowed, and whether the work created economic value.

Why now

Falling inference costs expand the market.

The common mistake is to assume cheaper inference lowers the opportunity. The opposite is more likely. When tokens become cheaper, software calls more models, devices stay connected for longer, agents retry and verify more work, and physical systems use inference continuously instead of occasionally.

The thesis in one line

The next 20 years attach every meaningful device and workflow to inference; the market rewards the companies that own the compute, memory, power, endpoint distribution, and trust layers that make that possible.