Appearance
Nobody is talking about the actual shift that happened this month. All the headlines were about new model benchmarks, funding rounds and regulatory leaks. No one connected the four stories that landed within 72 hours.
This is not about training anymore. For the last five years every fight in AI was about who could get enough GPUs to train a bigger model. That era ended this month. The next fight is over who runs the models after they are trained.
Nobody builds models. Everybody runs them.
Baseten closed $1.5B funding at $130B valuation last week. 18 months ago this company had effectively zero revenue. Last quarter it hit $6B annualized run rate. 2000% growth in 12 months.
Baseten does not train models. It does not build chips. It does not write foundation models. It solves one problem: once you have a model weight file, how do you run it reliably, cheaply, at scale, for actual users.
That is now the most valuable unsolved problem in AI. Every single fast growing AI application company is a customer. Cursor, Abridge, Clay, Lovable, OpenEvidence. None of these companies run their own inference stack. None of them want to.
For three years everyone parroted the line that moats would be in models. That was wrong. Moats are in running models.
The inference infrastructure market map
Right now this is the most crowded, highest velocity segment in the entire industry. No clear leader has emerged, but the valuations are already staggering.
| Company | Valuation (June 2026) | Core Value Proposition | Reported Gross Margin |
|---|---|---|---|
| Baseten | $130B | Full stack production inference, multi-cloud scheduling | 28% |
| Together AI | $75B | Open model orchestration, fine tuning + inference | 21% |
| Groq | $48B | Custom latency optimized accelerator hardware | 17% |
| Fireworks AI | $32B | Model routing and per task optimization | 24% |
| Modal | $27B | Serverless GPU runtime | 19% |
| Replicate | $19B | Developer focused model hosting | 12% |
This is not a winner take all market. It will not be one company. It will be at least three. All of them will be larger than most model companies.
Why NVIDIA is betting on Baseten
NVIDIA put $150M into Baseten in January. This was never a normal venture investment. This was a defensive move triggered directly by DeepSeek.
When DeepSeek launched at the start of 2025 it broke the core narrative NVIDIA had sold for four years. For the first time the market asked: if good models can be trained for 1/10th the cost, do we really need to buy this many GPUs? NVIDIA's stock dropped 18% in three days.
NVIDIA did not respond by building faster chips. It responded by buying the layer that decides which chips get used.
Cheap models do not reduce total GPU demand. They increase it. When model cost drops 90%, 100x more companies start running models. Total inference consumption grows faster than training ever did.
NVIDIA does not care who trains the models. It cares who runs them. Every model that Baseten schedules runs on NVIDIA silicon. That is the bet.
OpenAI just called the bluff
Three days before Baseten's funding announcement, OpenAI and Broadcom announced the Jalapeño inference chip. No benchmarks were released. No performance numbers. That was the point.
OpenAI did not build this chip to beat NVIDIA. They built it to gain leverage over NVIDIA. Right now OpenAI accounts for 17% of all H100 consumption on the planet. They have the single strongest negotiating position of any customer. They do not intend to ever actually manufacture Jalapeño at scale. They just needed to prove that they could.
This is the quiet standoff that no one reports. Every large model company is now designing their own inference chip. None of them expect to ship more than 10% of their total workload on it. All of them will use the existence of the chip to negotiate 30-40% discounts on NVIDIA silicon.
The brutal economics of inference
This is the hardest business in AI right now. Everyone sees the revenue growth. Almost no one talks about the margins.
Inference infrastructure is not a pure software business. 70-85% of revenue goes straight out the door to pay for GPUs. Gross margins for most players sit between 12% and 22% right now. That is worse than traditional cloud hosting.
Profit does not come from reselling GPUs. Profit comes from utilization. On the same H100, a platform running at 92% utilization makes 3.7x more margin than one running at 25%. That is the entire game.
Baseten runs their fleet at 89% average utilization. AWS runs public GPU instances at 41%. Google runs at 47%. That is the entire moat. There are no shortcuts here. There is no clever pricing trick. You either schedule work better than everyone else, or you lose money on every request.
NVIDIA's quiet pivot to physical AI
While everyone was watching the inference funding rounds, NVIDIA published the most important research roadmap of the year. It has almost nothing to do with LLMs.
Jim Fan's team laid out the full stack for physical AI data. The entire plan can be summed up in one observation: robot teleoperation data will never scale. You will never get one billion hours of robot demonstration data. You will get one billion hours of human demonstration data.
This is the new flywheel:
Every part of this stack runs on NVIDIA hardware. Every part requires orders of magnitude more inference than all LLMs combined today.
This is why NVIDIA is not panicking about OpenAI building custom chips. They already see the next order of magnitude demand. LLMs were the test case. Physical AI is the actual market.
Chips will now report their location
While all this was happening, the US Chip Security Act cleared the House Foreign Affairs Committee 42-0. This bill will mandate that every advanced AI chip manufactured after 2028 includes permanent location verification hardware.
Chips will phone home. They will throttle performance if moved outside an approved jurisdiction. They will refuse to boot if they detect they are inside a reshipping facility.
This is not a hypothetical. This bill will become law this year.
For the last two years chip smuggling operated on a simple loophole: once a chip left the factory, no one knew where it went. $2.5B worth of H100s were diverted to China in one single case this March. That loophole is being closed permanently.
Perfect enforcement will never happen. This law will not stop chips reaching China. It will raise the cost of diverted chips by 60-80%, add 6-12 months of lag, and make large scale diversion visible. That is the actual goal.
The new fault lines
We are now looking at three completely separate axis of competition that no one was talking about 12 months ago:
- Who schedules inference workloads. This layer will capture more margin than model builders, chip designers or cloud providers combined.
- Who converts human action into training data. The bottleneck for physical AI is not compute. It is labelled physical interaction data.
- Who controls where chips are allowed to run. Hardware will cease to be neutral commodity. It will become an instrument of geopolitical control.
All three are hardware problems. All three will define AI for the next decade. None of them have anything to do with who can train the largest model.
What comes next
Over the next 12 months you will see three things happen:
- At least one inference infrastructure company will go public. It will have a higher market cap than every open source model company combined.
- NVIDIA will acquire at least one human motion capture company. They will not announce it as an AI acquisition.
- By the end of 2027 every new H200 and B200 chip will ship with location tracking enabled by default.
Nobody won the training war. It just ended. Everyone already moved on to the next fight. Most people have not even noticed it started.