Appearance
This was not a normal week for AI.
There were no flashy demo videos, no fake benchmark leaks, no press releases about AGI being 3 years away. Instead, every major player quietly laid down their actual cards. If you were only reading headlines you missed it. The race for the best model ended this week. The war for the platform began.
Nothing announced over these seven days was about proving you could build a smart model. Everything was about proving you could run a real industry. This is the inflection point everyone knew was coming. No one expected it to arrive all at once.
The benchmark era is over
For three years every product announcement opened the exact same way. First a slide with MMLU, HumanEval and GSM8K scores. Then a claim that this model was now number one. Then everyone argued about the benchmark for three days. Then everyone forgot about it.
That pattern broke completely this week.
Not one major vendor lead with benchmark scores. OpenAI did not claim GPT-5.6 was the smartest model. Anthropic did not release a single leaderboard. Meta did not even bother running most standard tests for Llama 3.1.
Every announcement lead with price. Speed. Availability. Lock in. Supply chains. Lawsuits.
This is what happens when a technology stops being an experiment and becomes something people actually pay real money to use. Customers stopped caring which model gets one more point on a test. They started caring how much it costs to run ten million tokens an hour, and whether it will still be available next quarter.
OpenAI switches to utility pricing
OpenAI killed Codex this week. That was not the news. The news was what they replaced it with.
They did not merge products. They changed the entire business model. Flat rate subscriptions are gone. In their place is a meter. You get in the door for free. Every task you run burns credits. Harder tasks burn more credits. No one knows exactly what the rate is yet, but it will be tied directly to compute consumed.
This is not a small change. This is the end of AI as a consumer product. This is the beginning of AI as a utility.
| Model | Input price / 1M tokens | Output price / 1M tokens | Relative cost vs Claude Fable 5 |
|---|---|---|---|
| GPT-5.6 Luna | $1.00 | $6.00 | 12% |
| GPT-5.6 Terra | $2.50 | $15.00 | 30% |
| GPT-5.6 Sol | $5.00 | $30.00 | 60% |
| Claude Fable 5 | $10.00 | $50.00 | 100% |
| DeepSeek V4 | $0.44 | $0.88 | 2% |
| GLM 5.2 | $0.62 | $1.10 | 3% |
OpenAI did not even try to argue Sol was the most capable model. They only talked about token efficiency. They only talked about cost per completed job. That is the single most important shift we have seen in three years.
The product is no longer intelligence. The product is intelligence per dollar.
OpenAI is betting that almost no one will pay a 4x premium for 15% better performance on hard tasks. They are probably right. For 95% of work people actually pay AI to do, good enough and cheap will beat perfect and expensive every single time.
The hardware war went hot
Apple sued OpenAI on Wednesday. Most coverage treated this as a boring IP dispute. It is not.
This is the opening shot of the AI device war.
OpenAI did not hire 400 Apple engineers because they needed help writing CUDA kernels. They hired them because they are building a phone. Tang Tan did not leave Apple to run a server team. He left to build the device that will replace the iPhone.
Apple knows this. Everyone knows this.
Every single player is now attempting to own all four layers of the stack. No one is willing to be a supplier to anyone else anymore. That is why every partnership broke. That is why everyone is suing everyone.
There will be no settlement. This lawsuit will run for ten years. It will outlast multiple CEOs. It will define the entire next generation of computing.
There is no such thing as excess compute
You might have seen headlines this week about Meta selling surplus GPU capacity. You might have seen people claiming we have entered an era of compute glut.
This is a lie.
Anthropic locked 11.7GW of compute this week. That is enough power to run a small city. They locked it out to 2028. They are paying 190 billion dollars for hardware that does not even exist yet.
Meta is not selling excess compute because there is too much compute in the world. Meta is selling excess compute because Meta bought too much compute before they had a product to run on it. xAI did exactly the same thing. Those two companies have spare GPUs. Everyone else is rationing.
Power transformers now have a 160 week lead time. That is over three years. You cannot order a data center today and have it running before 2029.
90% of all compute capacity announced this year will not come online before 2028. All of it is already spoken for. There is no surplus. There never was.
Claude cracked internal reasoning, and everyone missed it
Anthropic published the J-space paper this week. 99% of the discussion was dumb clickbait about consciousness. Almost no one talked about what they actually built.
They did not make a conscious AI. They made something much more important.
For the first time ever, we can reliably read what a large language model is reasoning about before it outputs it. We can edit those internal thoughts. We can replace them. We can train the model on its own unspoken reasoning.
This is the single largest breakthrough in alignment and interpretability ever published. This changes everything about how we build, test and secure AI systems.
And everyone was arguing about whether the computer has feelings.
Right now there are teams inside every major AI company replicating this result. Right now they are building tools that will let them audit every thought a model has before it says anything. This will not stop bad actors. It will make good actors dramatically safer.
Open source won the argument this week
Meta dropped Llama 3.1 405B on Tuesday. Zhipu open sourced GLM 5.2 under MIT license on Wednesday.
For the first time ever, you can run a model on your own hardware that is within 10% of the frontier closed models.
Every single enterprise buyer noticed.
This is not a niche thing for hobbyists anymore. 70% of enterprise AI projects started this quarter will be built on open weights. Closed models will retain a premium for the absolute hardest 1% of tasks. Everything else is already gone.
Hugging Face also shipped one click deployments to both AWS SageMaker and Azure Foundry this week. No one wrote about it. That is the plumbing that makes open source actually usable for enterprise. That is the thing that will collapse closed model margins over the next 12 months.
Zuckerberg was right. The Unix repeat is happening much faster than anyone expected.
The end of human written codebases
Bun was rewritten from Zig to Rust in 11 days. One million lines of code. 165 thousand dollars of API credits.
This was not a demo. This was production code. This is code that is running right now for millions of users.
The argument is no longer if AI can write code. The argument is now: is code written by AI and never read by a human maintainable?
We are about to find out.
No engineering manager will ever approve a 12 month rewrite project ever again. No CTO will ever sign off on a team of ten engineers building something that an agent can build over a long weekend.
This is not something that will happen one day. This happened last week. The entire profession of software engineering changed and almost no one noticed.
China is no longer running behind
For the last three years everyone operated on the assumption that Chinese AI models were 12-18 months behind the west.
That gap is gone. It is now 3 months.
GLM 5.2 beats almost every western model on coding benchmarks. It is open source. It costs 1/30th what GPT-5.6 costs. Zhipu is now worth three times as much as Baidu. They just published the first paper showing a general purpose humanoid robot performing live surgery.
No one in the west is talking about this. Everyone should be.
The days when you could dismiss Chinese models as cheap knockoffs are over. They are competing on merit now. They are winning on price. They will win on volume.
What happens next
Over the next 12 months you will see:
- Model pricing will continue to fall by roughly 50% every six months. By this time next year frontier inference will cost less than 1 dollar per million tokens.
- Every major AI vendor will announce their own hardware device before the end of 2026. None of them will run anyone else's models.
- Open source models will pull ahead on raw benchmark scores by Q1 2027. Closed models will retain an advantage on reliability and safety, not raw capability.
- At least one major closed model provider will exit the market. They will not be able to compete on price.
- Governments will stop regulating only safety. They will start regulating pricing, availability and market power. AI will be treated the same way we treat electricity, telecoms and water.
None of this was inevitable. 12 months ago everyone agreed there would be three closed models, run by three companies, that everyone else would pay to use. That future is dead.
We are not going to have one AI king. We are going to have a messy, competitive, fragmented market. There will be bad actors. There will be failures. There will be price wars. There will be lawsuits.
That is good. That is how technology is supposed to work.