A Software Company Designed Medicines

This week Anthropic published the results of a protein design campaign. Fifteen targets, a written brief, and an instruction to design small proteins that latch onto them. The model succeeded against fourteen. Between 22.6% and 35.1% of its designs bound when two independent laboratories built and physically tested them, compared with a field norm of nearer 10% to 15%. For one target, RBX1, its best design beat the winning entry in a public competition that drew 245 submissions.

Anthropic sells a chatbot.

Worth restating slowly, because the summary hides the part that matters. Nobody there set out to enter drug discovery. The model was given a brief, the same public design tools that any laboratory can download, 12,500 GPU hours, and sessions of 24 to 48 hours. After that, it received no scientific guidance at all. It chose where to aim on each protein, assembled its own pipeline, ran the optimization cycles, and discarded everything it judged unlikely to work. 1,320 designs went out. 354 came back as binders.

In a market already anxious about AI's utility and the return on several hundred billion dollars of capital, this looks like the sort of obscure data point an enthusiast waves at people who already agree with her. It is not the only one from the past few weeks, and together they point to something considerably larger than better chemistry.

Our working title was "Specialization Is Dead As We Know It." We have overused that construction before, and it invites readers to build a straw man out of the extreme and then knock it down by correctly observing that field specialists of every kind will remain relevant in one form or another. Software development spent three years on that argument and learned very little from it. The question is not what these models cannot do today, or may never do, but what they can suddenly do and where that points.

For readers weighing this in financial terms, it runs in two directions. Utility and the returns on today's capital could surprise on the upside. The disruption could arrive well beyond the level of any single listed or private vehicle, reaching into how entire service sectors, and eventually societies, are organized.

The So What Crowd

Every technology gets its skeptics. AI has got a chorus, and it has been admirably consistent. Its position is two words long.

When the first chatbots appeared, the verdict was that they were statistics in a costume. Pattern matching over a very large pile of text, dressed up as conversation.

Then they got better, and the verdict became an enhanced search. Convenient, certainly. A better index of things other people had already written.

Then came images and video. The verdict was entertainment, with caveats. The machines could not reliably render a hand, reproduced whatever bias sat in the material they had absorbed, and invented facts with complete composure. Anything that could not count to five on a good day was not going to trouble a profession.

Then came drafted emails, meeting notes, and the first agents. Autocorrect with ambitions. And, they were still the models that could not count the number of “r”s in a strawberry. Books were written on hallucinations that were never going away and the biases that were doomed to become worse.

Then programming, where the initial verdict was error correction with better manners. Essays were written on why enterprise programming was untouchable. As the code became real, the verdict adjusted rather than moved. Perhaps some utility, though modest, set against the sums being spent on it.

Agents now run unattended on ordinary machines, doing errands for people working from home. The verdict has settled on pennies. Whatever these things are doing, the value of any single instance of it can be counted in small change.

Notice what that sequence has in common. Almost every judgment in it was defensible when it was made, and several were plainly correct at the time. All of them were about what the models had just done. None was about what they would do next. The objection has had to relocate roughly every eight months, and each relocation has been received as a fresh insight rather than as a retreat.

AI capabilities have been widening incessantly ever since the onset in this round three years ago. The generative models moved beyond text modals years ago, prompting us to ask the question whether they are killer apps or app killers two years ago. Breadth across multiple fields was visible then, and breadth is not the news.

What has changed sits underneath the breadth. These systems have become genuinely deep inside individual fields, deep enough in a few of them to hold their own against people, along with their domain-specific models, who have spent careers there. And because one system now holds many fields at once, it can do something no arrangement of human specialists has ever done cheaply. From Aristotle to Khayyam, Da Vinci, and Feynman, we have gawked at the multi-domain expertise of a few geniuses over centuries and the outsized effects they have had on things they touched. Our models are reaching expert levels in dozens of domains simultaneously, and their output is reaching new heights because of their ability to traverse them in parallel. And that is not just for innovators and innovations. They are also undercutting the value of single-domain specialists.

That appears to be a mundane-sounding exaggeration. Let’s step back to the question that predates every model in this note. Why did we divide knowledge into fields at all?

Nobody decided. It was an accommodation.

An individual’s memory has a capacity constraint. Once the volume of what could be known outgrew, the only way forward was to cut the territory into pieces small enough for a single mind to hold. Specialization has been a necessary civilizational condition. The first time humans settled from their hunter-gatherer days, some took to baking, some to farming, and some to ruling. 

You can date the squeeze by watching the polymaths thin out. Eratosthenes, running the library at Alexandria, was nicknamed Beta, the second letter, on the grounds that he came second in every field he entered. The phrase "the last man who knew everything" is the tell. It has been used in earnest about Leibniz, about Young, about Humboldt, and about Fermi. Four men across three centuries, each time by people convinced the door had just closed behind them. It kept closing because the material kept growing and the skull did not.

The sub-divisions of the recent century have been relentless. Medicine now has 40 specialties and over 130 subspecialties. The divisions may not have as complicated or even well-identified terms elsewhere, but some anthropologist or economist may have measured the progress through the growth in specialization. 

Our technology world inherited the instinct the moment software became an industry in the eighties. Single-purpose programs that do one thing, whether improving photographs or booking taxis, have always been trusted more than applications that span too many verticals. GenAI arrived with the same expectations. From the first months, there were proposals for domain-specific models in medicine, mathematics, and biology, then in languages, then in legal opinion, each on the reasoning that a narrow model trained on the right corpus must beat a broad one.

The Bill Nobody Itemized

Domain specificity carries an assumption that is never said out loud. It assumes the problem arrives already labeled.

Almost none do. A patient does not present with a cardiology problem. A company does not have a tax question. It has a decision, and the decision turns out to touch tax, along with four other things nobody listed at the start. Sorting the problem into fields is the first piece of work, and it is the piece no specialist can perform, because a specialist begins from the assumption that somebody else has already done it.

Inside an organization, this is visible every week. A pricing change touches finance, legal, product, competition law, and the tax treatment of whatever gets bundled. No single person owns it. So it gets divided, and that's where the money leaks.

Three separate losses occur at that junction, and they multiply rather than add.

Specifically, in any situation involving teams of experts, each specialist decides what to say, which is never everything they know. The listeners may absorb a fraction of what was said, and these things happen in loops and steps, with biases and errors included every time, for some to be corrected and others to multiply. What is eventually settled on as a solution is often suboptimal and always time-consuming.

Then there are losses that no one is aware of. No parties ask the question that should originate from their fields but is shaped by the details of other specialists’ domain details, because you cannot enquire about a field whose existence you are unaware of. Your accountant is not withholding the tax consequences of the indemnity clause. It has never occurred to him that the clause has one.

Medicine is the one profession that has bothered to measure this properly, because in medicine, the leak shows up as casualties. The Joint Commission has found handoff failures contributing to adverse outcomes in over 70% of the sentinel events it reviews. A study of paramedics handing patients to hospital trauma teams found that 79% of the information passed over was eventually documented, and that 9% of it was never recorded by anyone at all, including the paramedics who said it. Discharge summaries, the formal record of what happened to a patient, contain omissions or errors in close to half of all cases.

The most instructive result is what happens when somebody fixes the transfer rather than the medicine. A study across nine hospitals covering 10,740 admissions introduced nothing except a standard format for handing a patient from one doctor to the next. Errors fell from 24.5 to 18.8 per 100 admissions. Preventable harm fell by roughly a third. No new drug, no new machine, no new knowledge. A cleaner handover between two competent people, and a fifth of the damage disappeared. 

Software has been running the same experiment for decades without meaning to. Every project has two rooms. One holds people who understand what the business needs. The other holds people who understand what can be built. Between them sits a specification that satisfies neither, and requirements defects, the ones born in that gap, account for something like half of everything that later goes wrong. We have also seen the progress of organizations with individuals who are well-versed in understanding business needs and technological capabilities.

The point is simple. Many problems can be solved more effectively when specialists work together rather than digging deeper into their own fields.

The Work of Jack of All Trades

Take the four results of the past few weeks one at a time, because the summaries flatten the part that matters. The sections are somewhat technical, and many readers may want to skip the details here to move to the next section without losing continuity. 

a. Chemistry: In an experiment reported by Anthropic, a contract laboratory supplied raw data from two instruments used to check a compound. Claude Opus 5 received only the files and two short instructions. No vendor software or chemist guided it.

The LC-MS file was in a proprietary binary format meant for the instrument manufacturer’s software. Claude found no available parser, recognized the file’s outer structure, searched its 281 internal data streams, and worked out how the mass-spectrometry and ultraviolet data were encoded. Before interpreting the results, it checked its decoding by reconstructing the instrument’s own recorded totals exactly across all 2,664 scans. It also confirmed that thousands of ultraviolet-data blocks and their measurement axes had been read consistently.

Claude then produced the separation trace, mass and ultraviolet spectra, a peak table, a purity estimate, and reusable code for opening similar files. It identified the main component at 4.34 minutes and a molecular mass of 504 daltons. Its purity estimate was 96.4%; the laboratory reported 96.33%. The figures were calculated slightly differently, so the closeness should not be treated as exact replication, but the principal component, mass, and timing all agreed. Claude completed the work in 19 minutes. The laboratory’s final report arrived four days after the measurement, largely due to standard queues and reporting processes. Anthropic provides the details in its technical report.

The NMR file required a different kind of work. Claude converted the instrument’s fading radio signal into a readable spectrum using a Fourier transform, cleaned and calibrated it, located 18 signals, and estimated how many hydrogen atoms each represented. Where directly comparable, its hydrogen counts were within 0.08 hydrogen of the laboratory’s.

It also detected four broad signals that might correspond to hydrogens attached to nitrogen or oxygen. Claude proposed the standard follow-up: add heavy water and see which signals disappear. The laboratory had independently performed exactly that experiment three days later. Upon receiving the additional data, Claude initially concluded that all four signals had vanished. Its internal audit caught a contradiction in its own calculations, rechecked the evidence, and corrected the answer to two, matching the laboratory.

Consider how many different kinds of work were crossed in one short task. Decoding an undocumented file was a software engineering task. Fourier transforms and signal cleaning were mathematics and signal processing. Reading the peaks was chemistry. Proposing the heavy-water experiment was a scientific judgment. Checking the reconstructed totals was a quality control step.

The important point is not that Claude replaced the laboratory. The physical measurements had already been made. It removed much of the delay between measurement and understanding. Conventional scientific software usually stops when it encounters a file or workflow its developers did not anticipate. Here, a general AI built the missing software, validated it, processed the signals, interpreted the chemistry, and proposed the next experiment as one continuous piece of work.

b. Proteins: The success rates discussed at the start of this article tell us that the proteins worked. Two less obvious results tell us something more interesting about how Claude designed them.

Consider TNFα, the biological target of some of the largest-selling medicines ever developed. Claude produced binders that attached not only to the human protein, but also to its equivalents in cynomolgus monkeys and mice. This matters because a promising molecule typically must be tested in animals before it can enter human trials. If it binds only to human TNFα, researchers may need to create a separate version merely to conduct those studies. Claude preserved the relevant molecular features across three species, allowing the same binder to progress further through development. Cross-species activity was mentioned only as a secondary objective, yet Claude allowed a consideration that would come much later in drug development to shape the molecule at the design stage.

The second result concerned the shape of the binders themselves. Most computationally designed binders resemble bundles of α-helices, relatively forgiving structures that behave like stable molecular springs. β-sheets are more like folded ribbons: separate strands must align precisely, and errors can cause the protein to misfold or clump together. Claude nevertheless produced 15 confirmed binders with substantial β-sheet content across six different targets. It did not remain inside the safest and most familiar corner of the design space. It repeatedly reached into a structurally harder one and returned with proteins that worked.

c. Mathematics: In May 2026, a general-purpose model inside OpenAI overturned an almost 80-year-old belief associated with Paul Erdős. The conjecture was simple: place (n) dots on a page and ask how many pairs can be exactly one unit apart. Erdős believed that the answer could grow only slightly faster than the number of dots, and that arrangements resembling square grids were essentially as good as one could do. The model disproved this by leaving elementary geometry and importing ideas from algebraic number theory, using high-degree number fields and class-field towers to construct arrangements containing far more unit distances. It had not been specially trained in mathematics or given a system designed to search for proof strategies. The proof was generated in one shot, then simplified and verified by multiple mathematicians.

Anthropic’s Riemann result followed a different route. The Riemann hypothesis predicts that all the non-trivial zeros of the zeta function, which encodes deep information about the distribution of prime numbers, lie on a single special line. Proving the entire conjecture remains out of reach, but mathematicians have gradually proved that at least some proportion must lie there. That lower bound moved from one-third in 1974 to 40% in 1989 and then to 41.6% in 2020. An unreleased version of Claude raised it to 67.2%. The central contribution was recombination. Claude joined Bombieri’s work from 2000 with later results from Aryan and from Baluyot, Goldston, Suriajaya, and Turnage-Butterbaugh in a way that had not previously been assembled. 

How Claude got there is almost as revealing as the theorem. Across two Claude Code sessions, it generated 31 million output tokens. Its first 650 ideas failed. Asked to try again, it spent a day and a half coordinating about 60 subagents, running 2,400 shell commands, writing hundreds of Python scripts, downloading 54 research papers, and performing thousands of numerical checks against known zeta zeros. Some agents developed ideas, others searched for counterexamples, and still others acted as referees. The final result was independently reconstructed, written as a paper, and formalized in Lean, where it passed the standard proof validator. The person supervising the original search was not a mathematician. Once he had set the problem, his principal contribution was repeatedly telling Claude to keep going. 

d. Silicon: In July, Moonshot AI reported that Kimi K3 had spent 48 hours autonomously designing an accelerator for a miniature version of its own model. It produced RTL, the code-level blueprint for the circuitry, together with a bit-exact numerical model, simulation software, synthesized gates, and static timing analysis. Moonshot reported a 4 mm² design running at 100 MHz and decoding more than 8,700 tokens per second in simulation.

This was not a fabricated chip, and the distinction matters. The released repository explicitly states that its area and timing figures are pre-layout estimates, not results from placement, routing, or foundry sign-off. The Nangate 45 nm library is a generic research kit that cannot be fabricated itself, first released publicly in 2008. Read the demonstration for the chain, not the chip. A general model moved through computer architecture, fixed-point arithmetic, Verilog, functional simulation, synthesis, timing analysis, and the software needed to hold the process together. It crossed several engineering boundaries without handing the work to another person.

Other examples: In cryptography, Google Quantum AI published proof that it had sharply reduced the quantum resources needed to attack the elliptic-curve encryption protecting Bitcoin and much of the internet, but concealed the actual circuit behind a zero-knowledge proof, which verifies a result without revealing how it was obtained. At Eigen Labs, a 22-year-old engineer built a machine-checkable evaluator, then opened the problem to swarms of AI agents that read the literature, designed quantum circuits, and exchanged results. They matched Google within 8 hours and surpassed it within 72 hours; by the end of June, their circuit was 47.5% more efficient. The decisive move was not superior knowledge of quantum physics. It was converting the scientific problem into an executable score that thousands of ideas could be tested against. In China, Alibaba reported a similar collapse of boundaries with Qwen3.8-Max. Given only a research paper and access to GPUs, the general model wrote 7,600 lines of code, reproduced the paper’s six principal findings, ran 33 training experiments, invented and tested 18 further ideas, and produced a method that scored 2.71 points better than the original. The shape is becoming familiar: reading, programming, experimental design, measurement, and revision, once divided among different people, are compressed into one continuous feedback loop.

What This Looks Like on a Tuesday

The examples above are exotic because exotic results are easy to see. A theorem is either proved or it is not. A protein binds, or it does not. A circuit passes its tests or fails them. They give us unusually clean demonstrations of what happens when one system can move across several disciplines. But they can also leave the wrong impression: that this matters mainly to mathematicians, chemists, and people designing quantum circuits. The theorem is the spectacle. The spreadsheet is the economy.

At GenInnov, we already have dozens of examples of work that would previously have taken far longer, cost far more, or simply never have been attempted. A single-company investigation can now begin with twenty years of filings in several languages, pass through patents, product specifications, supplier comments, and conference transcripts, and end with a reconstructed financial history, a scenario model, and a list of contradictions requiring further investigation. A regulatory question can travel from the original rulebook to a clause-by-clause document comparison, to a spreadsheet of consequences, and to a draft set of questions for counsel. An idea about portfolio risk can become code, a simulation, a visual explanation, and a formal paper. These are not merely faster versions of the old work. In many cases, the quality is higher because more evidence can be examined, more competing explanations tested, and more awkward questions pursued.

Previously, such work crossed the desks of an analyst, an associate, a data specialist, a programmer, a translator, a lawyer, and sometimes an outside consultant. Each person might be excellent, but every handover costs time and discards some of the original question. Now, a single continuous process can carry the question through all those forms while humans judge the assumptions and the final decision. This is happening, in less dramatic ways, to almost everyone making a serious attempt to use these systems: accountants reconciling messy subsidiaries, engineers tracing failures across hardware and software, doctors assembling fragmented patient histories, lawyers reading entire transaction rooms, and managers turning meeting notes into operating plans and working prototypes.

The largest gain may therefore come not from completing existing work more cheaply, but from making previously uneconomic questions worth asking. When an investigation requires five disciplines and three weeks of coordination, most organizations never begin it. When it can be attempted in an afternoon and abandoned quickly if the evidence fails, the number of worthwhile experiments rises sharply. That is where the scientific stories connect to ordinary work, and where combinatorial explosion stops being a mathematical metaphor and starts becoming an operating model.

The Eigen Labs episode from the previous section is worth a second look because it is the template rather than the exception. The decisive act there was not quantum expertise. It was about building something that could automatically score an answer, after which thousands of attempts became affordable. The obvious objection is that a machine was turned loose on thousands of alternatives, most of which were doomed to fail. What generative systems do that matters is discard. The protein hit rate came from what the model threw away, not from what it produced. The chemistry result came with the model's own caveats attached and a self-correction that it was not asked to make. The mathematics came with a machine-checkable proof. Combination and judgment arrived at the same place: the arrangement no committee has ever managed.

That generalizes uncomfortably well. In any workflow where you can state precisely what a good outcome looks like, and an orchestrator guides the alternative-producing mechanisms steadily towards that outcome, the work becomes different. 

A Conversation Nobody Is Having Yet

Two conclusions follow from all of this. One is obvious. The other is barely being discussed.

The obvious one first. Token consumption has grown 158-fold in two years, according to a recent AMD release. The same period produced the $600bn question, a run of surveys reporting that the overwhelming majority of corporate pilots delivered nothing measurable, and a settled view in parts of the market that the utility is not there. The evidence in this note points the other way and points toward continuation rather than a plateau. It will convert nobody. The two sides of this argument - the one seeing a waste and another an exponential curve - have been talking past each other for three years, and each new result gets filed under whichever heading the reader already keeps.

The second conclusion is far more consequential. Capability that compounds across domains does not improve the way a single-purpose tool improves, and the honest position is that we cannot picture where it lands. Meanwhile, capital keeps veering toward specialist models and awarding them extraordinary valuations. Putting it plainly, just as began happening three years ago with software and services, an increasing number are likely to notice the fast-rising pressures on specialist work.

Let’s have one more piece of evidence before the final words. In a preregistered experiment published in June, 791 Procter & Gamble professionals were randomly assigned to real product-development problems. One person working with a general model matched the output of a two-person team working without one. More tellingly, the technical staff began producing commercially minded solutions, and the commercial staff began producing technical ones. The model reproduced a good part of the reason cross-functional teams exist at all. 

General models also have two advantages that are rarely priced in. Scale lets them approach specialist performance within individual domains at lower cost because the same investment is recovered across all domains simultaneously. And they can manufacture their own material to learn from. A specialist model is bounded by what its field has already produced. A system that can run experiments, check its own work, and keep what survives is not.

None of this is only about models. Organizations are built around large numbers of people performing narrow tasks, then passing partial answers across departments. Expertise will remain essential for judgment, validation, and responsibility, but much of the distance between one expertise and the next may not. We divided knowledge because the skull required it. The machine inherited the knowledge, not the limitation.

Related Articles on Innovation