← Back

The Tech Report

AI Bubble: 'Frontier models have run their course' | Eli the Computer Guy

The Tech Report26 min
AI bubbleOpenAINvidiainference costsfrontier models

In this episode of The Tech Report, host Isaac sits down with Eli the Computer Guy to interrogate the economics behind the AI boom. The conversation is triggered by a report from The Information that OpenAI has quietly found a way to halve its inference costs, a technical win that Eli argues could paradoxically undermine the company's own trillion dollar valuation. From there the two unpack a much larger thesis: that the AI industry is caught in what Eli calls a technical Ponzi scheme, where trillions in infrastructure spending keep flowing even though the premier AI company still hasn't demonstrated a clear value proposition. They range across Meta selling excess compute, Nvidia's shrinking runway, the rise of local and open source models, and Eli's provocative claim that frontier models have essentially run their course and become like incremental software version bumps rather than revolutionary leaps.

Zašto OpenAI šuti o prepolovljenim troškovima inferencije

The episode opens with a puzzle. The Information reported that OpenAI has found a way to roughly halve its inference costs, a secret so closely guarded that even most of OpenAI's own engineers reportedly don't know how it works. Isaac's instinct is that if this were genuinely true, OpenAI would be shouting it from the rooftops rather than letting it leak through sources familiar with the topic. Eli's answer reframes the whole thing around business incentives rather than technology.

Eli distinguishes between two parallel AI stories: the technology, meaning embeddings, vector databases and neural networks, and the business, meaning the eye watering valuations. Technically, getting more service out of fewer resources is exactly what everyone wants. But the business argument for these companies rests on the idea that AI is so expensive it justifies hundreds of billions in investment, and that cost is the moat. If OpenAI loudly proves you can get the same output with half the resources, it directly cuts against that moat.

He ties this to Sam Altman's stated position that an IPO under one trillion dollars is a non-starter. Announcing a dramatic efficiency gain immediately invites the question: doesn't that mean you're actually worth 500 billion rather than a trillion? So keeping a lid on the breakthrough protects the valuation narrative even if the engineering is genuinely impressive.

"if the argument with AI is that it is so expensive that's why they need hundreds of billions of dollars in investment and that that is basically the moat around their business then literally by saying that you can get as much response out with half as much resources literally cuts against their business argument"

Nelogirani korisnici kao poligon i opasnost od optimizacije memorije

Isaac notes the inference halving was applied specifically to users who weren't logged in, people who visit the site, type something, and leave. He wonders whether that's because they make a good pool of low complaint test users or whether it reveals something about how the efficiency was achieved. Eli leans toward the reputational explanation: these are the lowest level of users, below even free logged-in users, so if something goes wrong there is minimal brand risk.

Eli then explains the likely mechanism, drawing on recent work IBM has discussed. VRAM, the memory on the GPU, holds the loaded LLM, but it must also store the tokens of responses when concurrent users hit the model. IBM has been exploring predictive ways to optimize how those tokens are stored, work that gets complicated very quickly. In that complexity lies danger: if the system starts saving things in unexpected ways, the returned responses could go far beyond ordinary hallucinations.

Because of that risk of tragic mistakes, testing on non-logged-in users makes sense. As Eli puts it, if you really mess up on people who couldn't be bothered to log in, what is there to lose. He suspects the actual innovation is some memory optimization routine that probably can't be protected by patent, which has major competitive implications.

"if it if it if it starts saving things in weird ways that they're not expecting, then returning that response, you could get something that goes much beyond the normal hallucination you're thinking about"

Utrka prema dnu i nemogućnost naplate uštede

The pair turn to what OpenAI could actually do with the savings. Eli argues the real problem is that almost everyone is using fundamentally the same underlying technology, merely optimized for their own companies. If OpenAI could patent its trick and secure intellectual property protection, there would be an interesting question about monetization. But he bets the innovation is a memory optimization routine that can't be protected, which means everyone else will simply replicate it.

This triggers a race to the bottom. Eli points to DeepSeek already operating at roughly a fifth of OpenAI's cost. If DeepSeek figures out what OpenAI did, he expects it could reach a tenth of OpenAI's cost. The consequence is that OpenAI will be forced to drop its own prices rather than capture additional value, so the efficiency gain never translates into fatter margins.

The deeper irony Isaac raises is that even halving inference costs wouldn't balance the compute cost of tokens against the price of a subscription. So the breakthrough helps the technology story while doing nothing to rescue the business story, because commoditization erases any pricing advantage almost immediately.

"OpenAI is going to have to drop the drop the price that they're charging and so they're they're never they're never going to get an additional value out of this simply because it's got to be that that race to the bottom"

Tehnički Ponzijev sustav i politika masovne izgradnje

Isaac raises Eli's recent video about Meta selling its excess compute, something he also discussed with Ed on Friday. The very existence of surplus compute throws the scale of the AI buildout into question: if there's already too much, why keep building. Add inference halving on top and the claim that demand for compute vastly outstrips supply looks even weaker. Eli says it only makes the argument uglier, and he has been warning for years that these trillion dollar buildouts don't make sense.

He underscores that OpenAI has branded 2026 the year of productivity, with Altman promising to finally prove what AI is actually for. Eli finds it damning that trillion dollar companies have poured trillions into infrastructure while the premier AI company still hasn't shown a value proposition. He believes much of the investment has little to do with technical capability and more to do with politics: AI deployments, however poor, are growing the economy by around one percent, so there is pressure to keep investing to keep everything from falling apart.

Isaac compares it to post-crash American infrastructure programs where governments funded dams and bridges almost pointlessly to inject money into the economy, except this time it's a handful of businesses driving it rather than the state. Eli agrees, noting this is the first massive tech buildout in decades that investors feel comfortable with because they believe they understand AI. He rejects the dark fiber comparison, pointing out that we're told there are no dark GPUs, yet SpaceX is renting out something like 340,000 GPUs and Meta, a social media company, has enough spare capacity to rent GPUs at meaningful scale.

"I think we're basically in kind of like a a massive technical Ponzi scheme of we have to keep investing for everything for everything not to fall apart"

Preživljavanje bankrotom konkurencije i olovna cijev u serverskoj sali

Isaac asks whether the hyperscalers are simply locked into spending, since pulling back would end badly while staying the course at least postpones the inevitable and leaves open some unknown future rescue. Eli agrees the best bet is endurance: one company outlasting all the others. In his view, OpenAI's only real play is to hope Anthropic goes bankrupt, because with the current field of players not all of them can succeed.

The danger is a down round. If any of these valuations fall, it triggers broad chaos, so the strategy becomes pushing valuations ever higher and waiting for a competitor to trip and fail so the survivor can consolidate. Isaac offers the familiar image of not needing to outrun the tiger, just the other person. Eli caps it with a darkly comic line about why there's a lead pipe in the server room.

This section captures the zero-sum brutality Eli sees beneath the optimistic public narrative. The industry is not so much building toward collective success as engineering mutual attrition, where staying solvent longer than rivals is the actual business model.

"if you ever wonder why there's a lead pipe in the server room, it's to uh crack crack the kneecap of your competitor when you're running"

Nvidijin kraj tržišta i nacionalna trka za podatkovnim centrima

Turning to hardware, Isaac asks where a real inference breakthrough leaves Nvidia, memory manufacturers and the tens of gigawatts of data centers under construction. Eli says Nvidia is already in a weirdly tough spot as competition mounts, citing Google's TPUs and the Meta and Broadcom partnership for inference processing. He flags an unusual sign of stress: Nvidia is now offering GPUs to startups in exchange for a cut of future revenue rather than selling them outright.

Eli believes Jensen Huang knows the writing is on the wall for the total addressable market for GPUs, which is coming to an end one way or another. That is why Nvidia is diversifying into Vera CPUs and new data center architecture standards for liquid cooling. The inference breakthrough may accelerate this problem, but the problem already exists and the industry is already seeing its effects.

Memory manufacturers, by contrast, he sees as sitting pretty. Because every country now treats AI as more dangerous than a nuclear weapon, France, Germany and India all want their own data centers, driven by national industrial capacity concerns as much as ordinary profit motives. Eli references the Trump administration pulling Mythos 5 and Fable 5 in 90 minutes as the kind of event that spooks other nations into building sovereign capacity. Since only a few manufacturers can produce memory at scale, they will make money for a long time, while Nvidia remains a question mark.

"Jensen Wong knows the writings on the wall where the total addressable market for GPUs they're coming to the end of it one way or the other"

Frontier modeli kao verzije Pythona i lokalni tok inteligencije

The closing argument is Eli's boldest: frontier models have run their course. He was surprised when Google's Gemini 3 arrived, because until then Altman had been pushing integrations, building perhaps ten different integrations on top of a strong model, only to suddenly pivot back toward building a better frontier model. Eli sees no inherent value in a model per se; the value lives in the products and integrations built on top of it.

He compares future frontier model releases to Mac OS updates or Python version numbers. Just as you can skip from Python 3.8 through 3.13 and barely notice, he expects OpenAI 3.13 versus 3.15 to feel similarly incremental, with better security and quality in the background that most users never register. Meanwhile the big industry push is routing requests to the smallest possible model, and Eli predicts small models will solve 99 percent of problems while the frontier model is needed only 1 percent of the time. That again undercuts the trillion dollar valuation and the question of what customers are even paying for.

Eli sees the future in local and on-device intelligence, and in AI becoming an ambient feature rather than a destination product, much like Windows on a computer. He argues Google holds the front runner advantage through a layered intelligence flow: a lightweight model in Chrome, a model on Android, a larger model on a laptop with up to 16GB of RAM, and finally Gemini in the cloud. Companies like Cisco, rolling out agents to 90,000 employees on an in-house stack, and eventually Salesforce, will host their own internal AI rather than call OpenAI or Anthropic. His final worry for OpenAI is business development: unlike Google, Apple and Microsoft, OpenAI is just a thing with no established product suite or sales pipeline to step users through the models, a gap he compares to a promising startup friend who built brilliant technology but ran out of runway because he never learned to sell it.

"I just think for a lot of this that Frontier models have kind of run their course"

Key takeaways

  • OpenAI may have halved inference costs but keeps it quiet because proving AI is cheaper undermines the expensive-moat narrative propping up its trillion dollar valuation.
  • Any efficiency gain is likely an unpatentable memory optimization that rivals like DeepSeek will copy, forcing a race to the bottom on prices rather than fatter margins.
  • Eli frames the entire buildout as a technical Ponzi scheme, sustained by political and economic pressure to keep growing rather than by a proven value proposition.
  • Evidence of surplus compute, from Meta renting out capacity to SpaceX renting 340,000 GPUs, contradicts claims that demand vastly outstrips supply.
  • Nvidia faces the end of its total addressable GPU market and mounting competition, while memory makers thrive on a global national race to build sovereign AI data centers.
  • Frontier models are becoming incremental version bumps; small and local models will handle 99 percent of tasks, pushing AI toward ambient, on-device functionality where Google holds the advantage.

Resources mentioned

  • The Information
  • IBM research on VRAM and token storage optimization
  • DeepSeek
  • Google Gemini 3
  • Nvidia Vera CPUs
  • Eli the Computer Guy