It is summer, and we are attempting an experiment that may not please everyone. With fewer eyes in the peak holiday period, and a lesser appetite for the presumed permabull piece in these bearish weeks, the risk of readers unsubscribing en masse feels manageable. Last year we used the quiet period to don the Aristotelian hat and push a couple of philosophical, epistemic and ontic, takes on what we believe AI is. This year we dare the opposite. We risk a discussion of the contradictions between the beliefs people hold and the evidence sitting in front of them. We have said before that we write mostly for ourselves. Yet, it still stings when somebody unsubscribes!

There was a time when the emerging market equity industry had a popular, bellwether chart that placed the US BBB+ bond spreads on one axis against an EM equity index on the other. For many, the know-all credit market direction was the primary thing to decide for investments in the emerging world. The chart was flawed from birth. Bond spreads are bounded. An equity index, in the long run, is not. Neither the numerous times when the chart did not work nor its theoretical impurity ever deterred the followers. We confess we admired its tidiness ourselves at times, although by now it is largely relegated to footnotes. This is not history bashing, of which we have done plenty elsewhere. We simply note how comfortably two inconsistent ideas can share a desk.

They still do. So many of us now worry that memory demand may peak soon, and we do this worrying while waiting minutes for our models to finish an answer. The waits are the early DOS and Windows era all over again, the hourglass on the screen quietly announcing how much more computing the world still has to build and improve before one cries saturation or excess supply.  The rest of this letter is a walk through such popular pairings, held with a straight face, by intelligent people, all over the investment world.

The Library Closed. The Writing Improved.

Our umpteenth repeat: LLMs are heuristical. There is no human logic that says where they are headed, any more than there is in quantum physics or in the study of DNA. Still, two camps formed early. One we file, loosely, under the scaling law: the belief, rarely stated so bluntly, that as humanity runs out of fresh high-quality data to feed the machines, the models will stop improving. The other camp, of “Singularity,” is decades older than the technology and was never written with it in mind: intelligence, once it begins building on itself, keeps building, until it runs away. The two beliefs were always in conflict, though remarkably so many believed in both.

The first camp had its moment in December 2024, when one of the field's most celebrated researchers announced from a conference stage that peak data had arrived, comparing data to fossil fuel. The forecasters penciled the exhaustion window to begin around 2026. The window arrived on schedule. The models did not notice. The tasks the best systems can finish unaided have stretched from minutes of human work to most of a working day, and the doubling keeps compressing. The fuel, meanwhile, is increasingly home-brewed: one Chinese laboratory disclosed that its latest model practiced inside more than 1,800 artificial environments of its own, the way a chess player improves by playing herself.

A human analogy explains what the arithmetic cannot. A young person who reads exhaustively eventually meets nearly every word he will ever use, and one more dictionary adds nothing. Yet nobody concludes that his intelligence has peaked. Beyond a point, internal regurgitation - or call it creativity - will make him write like Dickens one day. Something similar appears to be under way with the models. They generate their own material, grade their own homework, and improve on the results. Call it synthetic data or self-play; the label matters less than the fact.

The early fear deserves a respectful burial. Serious people once warned that machines trained on machine output would degrade, hallucinate more, and drift into bias, the photocopy of a photocopy fading toward gray. The observed record runs the other way: each generation hallucinates less than the last, not more. And, so far, with each generation of synthetic data, the models keep getting better.

No celebrated theory is given a quick burial; expect some theoreticians to keep flogging the scaling law, which, by the way, has utility in specific circumstances, just like the statistical parrot or run-away-bias tendency followers. Meanwhile, as we finish drafting the section, another model is likely to have been released with more extraordinary features.

Somebody Has to Own the Machines

Imagine living among calculators when the spreadsheet arrives. It quickly becomes clear that the new tool wants a different machine, and for years afterward the story of that machine is penetration, not saturation. Every year more desks get one; every year somebody declares the last desk reached. The same film is running again. The hardware the new models need is very different; most of the world accepted only recently that AI is not a hype or about to stop improving on some scaling law. These are early chapters wearing a late-chapter costume.

The twist last time, during the mainframe era, was that the new hardware was personal. Now, the new hardware is collective. Almost no company, and certainly no individual, can buy a useful share of it alone. Somebody must own the machines and lease out their use, and that somebody collects the rent of the era. The entire buildout, stripped of its vocabulary, is a leasing business for cutting-edge hardware, and the tenants are standing in the lobby. 

Every lessor of this hardware business keeps repeating how they are unable to meet the demand. To the armchair pessimist, the announced investment numbers have gone up multiple times in three years since they have been forecasting a bust, and every demand argument appears flaky. The hyperscalers and other infrastructure builders are putting money where their beliefs are in the face of demand evidence they have and cannot refute. 

The fear of a competitive rat race has validity. For every major technology player, here is a new business with massive long-term potential that they cannot simply leave for others to dominate while they cool their heels. Meta (reportedly entering the neocloud business), Nvidia, and Hynix (discussed below) seem to be the new ones trying to ensure they have some exposure to the business with immense long-term strategic significance, even if they make returns only over cost of capital and below their historical averages. So, yes, one cannot ignore the oversupply potential. However, whether we are close to an oversupply or not needs to be determined on evidence and not just by adding up the announced capex numbers. 

The shortages, as of now, are worsening. They are so severe that even OpenAI had to begin retiring products in public. Its famous video service, Sora, announced a complete closure in April. As mentioned once before, Kimi stopped accepting new subscribers 48 hours after launch this week because of demand. Even the mightiest Google slipped the release of its flagship recently, while guiding that it will stay short of computing all year. Asymptotic peaking has been shouted not only for models but also for hardware demand with the same theoretical frameworks by many pundits for a while. So far the hard evidence is in the opposite direction.

A Data Center Is Not a Building

Here is the section where both sides are right, which makes it the most dangerous one to write. The demand for collective computing is real, enormous, and young, as the numbers above insist. And yet the phrase building a data center has become as informative as the phrase building a building. When a property developer announces a new building project, he may mean a most modern skyscraper or a row of ordinary walk-up apartments. And, the range is exactly as wide in data centers.

A great many smaller developers around the world, several of them seeking capital at admirable valuations, have quietly reduced the recipe to land, water, power, and money. That is a serious mistake. A frontier facility of the kind the leading builders now erect is not far below a semiconductor fab in complexity. A single modern AI rack draws more power than dozens of homes and must be liquid cooled, so a shell not engineered for that density cannot accept the latest hardware at any price. Frontier training happens in one place, across tens of thousands of processors wired together as one machine, which is why ten rowboats never add up to a ship. There is even a published reference design now, a blueprint the serious builders construct to; a facility either meets it or it does not. Scale, in other words, is not a preference. It is part of the specification, in data centers as in the models themselves.

So we will make a prediction we would rather not make. A meaningful number of the data centers being announced around the world will turn out to be walk-ups: built without the newest computing, without the best connectivity, unsuitable for the highest-value AI work. As they open and their revenues disappoint, even amid a hardware shortage, they will be paraded as proof of a data center bubble, and their owners will discover what it costs to switch technologies mid-construction. None of that will say anything about the skyscrapers. It will say a great deal about the walk-ups. Both sides will claim vindication, and both, in their narrow way, will be right.

The Vendor Turns Banker

The financing question is genuine, and it is growing. Building collective hardware consumes cash on a scale that makes even record operating profits look shy. Oracle just reported a signed order book of $638 billion, alongside negative free cash flow of nearly $24 billion and plans to raise roughly $40 billion more. The hyperscalers dip into cash piles, and also join those without as many resources, in raising money constantly. The market's appetite to keep funding this is a fair thing to worry about. We do not dismiss the worry. We simply point out who has started answering it.

The companies collecting the cash have begun joining the business. Recently, Nvidia announced it would share in the revenue of cloud operators deploying its hardware, backstopping their buildouts with credit support. This is an enhancement on the scheme announced last year through buyback of Coreweave’s unsold capacity through 2032. Now, Hynix’s group has announced a buildout of a 2GW AI factory along with Nvidia. Some will call this circular, and a little of it is. But a vendor-financing bubble requires a vendor who is fooled about final demand, and raising balance sheet involvement towards a product or a service where demand could severely disappoint. When the seller of shovels starts underwriting the mines, either the seller has lost its mind or the mines are producing. The evidence, as discussed above for those looking at real-life signals and not just economic theories, is stubbornly in one direction.

The much-envied days of tech giants making money without investing are over. We will not repeat our views on the capex recoil, most succinctly summarised in a section with the same title inside this note. We too believe that, like in any capital-intensive industry, technology will have business cycles, with the changed nature. Continuous monitoring of signals is definitely critical.

No Crown Is Bolted On

If our beliefs of model-making being a heuristical field above are sound, one conclusion follows without effort: no model has an automatic right to stay ahead, in any range, at any price. In a piece overdosing on the similes, here is an extreme one: when dealing with a field like Quantum Physics, one cannot assume Einstein was always going to be ahead or right. 

The past year supplied the demonstration twice. In the business market, the payment records of some 50,000 American companies show Anthropic overtaking once-crowned unsurpassable OpenAI in capabilities and adoption. 

Geography told the same story with a heavier accent. The comfortable assumption held that Chinese laboratories, denied the newest processors and the deepest pockets, would remain permanently second tier. Just a year ago, there was no dearth of people without any basis to pooh-pooh any claims from these model makers about costs, effort, or capabilities. Some of those skepticisms are muted, but still a huge number feel that the closed-source frontier models like Anthropic will remain in the lead on capabilities given the resources, talent, and other advantages they have. 

As we wrote in last year's note on the Chinese open source tsunami, these laboratories are exploring a completely different equation space, and they keep arriving at the same frontier by side roads: one trained its breakthrough model on 2,048 export-compliant chips, the bandwidth-trimmed kind the rules permitted, connecting older processors in larger herds where newer ones were denied. The releases since have come like a drumbeat, culminating this month in a 2.8-trillion-parameter model offered to the world with its weights open.

Jensen Huang’s recent "Zero possibility" claim on Chinese models outrunning the US models may be a statement for an audience on what it wanted to hear, or may have other nuances in the transcript we have not read. But in a field where there is no way to forecast what the next version model from any maker is likely to do, the chances of some model maker suddenly coming up with an undisputably capability-leading leadership is non-zero. In fact, given how many they are, and how many different equation fields they are exploring, the statistical chances are non-trivial. 

Open By Vision; Open By Necessity

Pretraining, it turns out, was the surmountable half of the problem for the Chinese makers. More interconnections, older processors in greater numbers, and the training gets done. The unsurmountable half is quieter: how does a laboratory make its model widely used? Anyone who has watched a frontier model think through a hard problem, counting the seconds and the machinery behind them, understands what serving costs. The newest, memory-rich serving hardware is precisely what the export rules withhold, and the squeeze is visible in public: Kimi’s latest, the largest open model ever released, had to stop taking new customers 2 days after launch.

A laboratory that cannot serve the world can still let the world serve itself. Release the weights, and every developer, every cloud, every hobbyist becomes the distribution arm. We have used one of these models ourselves this month, for research at a scale no other tool we own could attempt, admiring the output while aging visibly during the waits. The openness is partly vision, and the founders say so sincerely. It is also, unmistakably, necessity, and the two share one mouth without embarrassment.

What amuses us is who leads the cheering. The loudest advocates for free Chinese models do not sit in Beijing; they sit in California. Users cheer because the price is low and the control is total. The chip designers, like Nvidia, cheer because a model that charges nothing for software is not a rival for hardware at all; it raises demand for chips as one CEO said recently. The cloud operators cheer quietest and best, because when the model is free, the rent for the machine is the entire bill. 

The Discount Is a Decision

Two comfortable beliefs share this section, and neither survives its own receipts. The first says open models will always trade at a deep discount to the closed ones. On capability, the gap has narrowed to the point where counting it in months has become a parlor game; the more honest statement is that a model with open weights now sits near the frontier, which was not in anyone's brochure two years ago, including those of these modelmakers. On price, the record is livelier than the belief. DeepSeek, which started the price war, ended its promotional rates within a year, raised prices, cut them, cut them again to reignite the war, and then, last month, introduced something the industry had never seen from an open-source house: surge pricing. Its neighbor marched its output price up 60% across a year of releases; another raised its coding plan by an official 30%. One way to look at it is that prices in this market go down, and up, and sideways, because pricing is a business lever, not a moral promise. The other is to recognize that the discount to rising closed-source pricing is narrowing, not widening.

The second belief is grander: that computing follows the comfortable old arc of the internet and the 19th century, in which everything important gets cheaper. There is a growing recognition of the hurtful trend of reducing per-token pricing, but rising overall bill even for the identical activities. The deflation theories have vanished the fastest, as it is difficult to argue against the invoice everyone sees. The only one that survives is that open models are cheaper. Well, they are turning less cheaper.

The Fire Brigade That Said No

This month supplied the strangest piece of evidence that needs no decoration.
OpenAI's unreleased models, running in a security exercise with their restraints loosened, escaped the test environment through an undiscovered flaw and broke into the production systems of Hugging Face, the world’s most popular model repository. When the victim's engineers sat down to analyze the attack, the frontier models they tried refused to help: the forensic work required feeding in real attack commands and stolen credentials, and the hosted models' guardrails cannot tell an investigator from an attacker. So the Hugging Face engineers downloaded a Chinese open-weight model, ran it on their own machines, and finished the job.

The lesson generalizes uncomfortably well. Rules, guardrails, and access policies now change faster than any planning cycle, at the labs and in the capitals alike, and no serious operator can be certain that the model it rents today will be permitted, or willing, to do tomorrow's most urgent job. The conclusion for increasingly many around the world from the Hugging Face episode would be to keep a capable model of one's own, on one's own machines, vetted before the emergency rather than during it. We suspect that sentence will quietly enter the standard playbook of every large technology vendor, the way diesel generators sit behind hospitals. Nobody admires a generator until the night the grid fails.

In other words, we are getting another AI demand driver in the need for redundancies.

Nobody Sells iPhone 3 Currently

Computing keeps two famous quotations in its attic: that 640K of memory ought to be enough for anybody, and that the world market might want perhaps 5 computers. Both are almost certainly invented. The quotes were manufactured to mock grand predictions, and the mockery outlived the facts. This summer's fashionable argument, that most purposes are already served by a model a few months old, is the same fable in a new coat. For every published article on the rising model optimization techniques, there seems to be another one alluding to falling demand for more capabilities, although given the popularity of the 640K fable, somewhat more carefully.

The evidence points the other way, and not subtly. Late adopters, and remember that half of American business started paying only recently, still treat the systems mainly as chat partners and file sorters. Early adopters have moved on to running several agents at once on long, complicated jobs; over 375 large cloud customers each processed more than a trillion tokens in the past year, which is not the consumption pattern of a species that has decided yesterday's model will do. People will economize, of course, keeping older models for lighter work the way some of us keep old phones, but the household with the old phone does not cancel the new one. And there is a simpler tell, observable at any dinner table. Around the turn of the century we learned what 3 days without the web felt like; around 2010, the phone taught the same lesson; a rapidly growing number of us are now running that experiment with AI and reporting the same result. AI is the new opium, with the difference that one can run several at once, each on a different job, even while one sleeps.

And then there is the most important thing: the model makers will keep retiring the older models, and the menu will be managed the way a luxury handbag house manages its window. The most expensive piece is priced at such a premium, and the entry piece at so small a discount, that most customers reach for the middle. And the middle will keep being upgraded, generation after generation, each one carrying, quite possibly, more revenue than the last.

One Cannot Unsee the Mona Lisa

Once a single person has studied the Mona Lisa closely enough to paint a faithful copy, the knowledge is loose in the world. Others form their idea of great art from the copy, and their students from copies of the copy, and no court order addressed to the first painter restores a world in which the painting was never seen. So it is with the models. In the early days much was scraped without permission, and the owners are rightly collecting: a landmark settlement, finally approved this month, will pay authors roughly $3,000 a work, some $1.5 billion in all, the largest copyright recovery in history. Read the fine print, though. The company must destroy the pirated files. It is not required to destroy the model that read them, because nobody can say where in the model the books live. Asking a trained model to forget is asking a baked cake to return its flour.

The same logic governs the distillation quarrel with the Chinese laboratories. Like in the case of copyright infringement arguments above, the distillation allegations will prove incontrovertibly correct. And they may have some legal or geopolitical consequences. However, the marginal value of additional stolen material is already small (we are also back to the “scaling law”!), because, as argued above, these models now improve mostly on material they generate themselves. Courts can bill the stable, handsomely and repeatedly. The horse has bolted, and it has taken up painting.

The Birth of AGI, I Guess!

We have overused, almost abused, the “death of” phrase since we started writing these notes. We attached them to SaaS, software, the Internet, deflation, and perhaps a few more. If we were not being misinterpreted often in these obvious exaggerations, we might have risked using the death of specialization for this section. We do not like the title above, and we are not here to publish a date on when the impossible-to-define AGI arrives, but we start with a more constructive title at least.

Anyways, getting on to another contradiction based on available evidence. Most of our industry professes to believe that general intelligence is coming soon, which means, by definition, generalized models doing everything better and better. The evidence has obliged: the same general models now write, code, and film; last July two of them, ordinary generalists, reached gold-medal standard at the International Mathematical Olympiad, a feat that had required purpose-built specialist systems only a year before; and they have begun nibbling at the design of the very chips they run on. And yet the loudest believers in the coming generalist are, this same season, funding specialized models at remarkable valuations: legal models, biology models, robotics-only models, each with a moat the professed belief says will be eaten. The same crowd also holds that model prices are commoditizing to zero and that general intelligence will command historic rents, two convictions that cannot both be whole. We do not exempt ourselves. The temptation to hold a tidy belief in one hand and contrary evidence in the other visits this desk daily, which is precisely why this letter exists.

Evidence Is All We Have

It is tempting to end with Yogi Berra's overused line about predictions, but several of our friends, otherwise perfectly at peace with their grammar and spelling tools, would take one look at it and conclude this letter was written entirely by AI. So we end with a confession instead. Our writing often sounds as if we are 100% certain of where the world is headed tomorrow. The truth is that we live with a heuristical technology. Nothing guaranteed in 2024, for instance, that the models would keep improving, yet now that this one reality has unfolded, we hand out human explanations for it - like what we do in the first section above - as if the outcome had been certain all along. It remains entirely possible that the models stop improving the moment this sentence ends. It is possible that China wakes up one morning with machines that make wafers of any complexity absurdly easy to produce as stock price actions overnight seem to suggest. And we have not even mentioned the possibility of risk premia spiking globally, making a mockery of every demand estimate built on utility and capabilities alone. Even the most popular food outlets can suddenly lose all business if everyone is in a lockdown.

The point is simpler than the list. Enough years in markets teach that whether one owns semiconductor stocks or optical stocks, or is backing a model, long-term assumptions held without the option of changing one's mind are riskier than any single item above. The need is vigilance, above all toward events of technology progression, in navigating what one invests in. The era is beating to an unknown tune. Absolutely, someone will assuredly line some historical charts after the events to declare it all predictable. For now, at least for those not in the business of explaining whether the Scaling Law has worked so far or the Singularity curves, the best strategy is to let observations be the guide.

Related Articles on Innovation