Buying Bankrupt Biotech: Why Medical AI Needs More Data About Biology

OpenAI is funding wet labs to generate proprietary wet data. But the real shortcut to medical intelligence might be hiding inside the bankruptcy filings of failed biotech startups....

Feed
September 15, 2026
Buying Bankrupt Biotech: Why Medical AI Needs More Data About Biology


Everyone wants an omniscient digital doctor. Silicon Valley keeps throwing massive transformer models at generic internet text, expecting clinical miracles to magically emerge from Wikipedia articles and public health blogs. It is a delusional approach. Real medicine doesn't live in textbooks; it lives in messy, high-dimensional biological reality that current algorithms simply cannot parse because the underlying data doesn't exist in a clean, scrapeable format.

A while back, policy analyst Ruxandra Teslo floated a brilliant, slightly cynical shortcut for supercharging medical AI systems: scavenge the corporate corpses of dead biotech companies. Think about it. When a promising drug discovery outfit goes belly up, its bankruptcy auction holds a goldmine of regulatory filings, toxicological readouts, manufacturing strategies, and failed clinical trial logs. These trade secrets represent millions of dollars of physical trial-and-error. They are the exact ground-truth signals machine learning models starve for.

Now OpenAI is opening its checkbook to generate biological data directly through wet-lab partnerships. They're paying scientists to run experiments and feed the raw outputs straight into neural networks. It is a necessary pivot. Scaling laws hit a hard wall when you manage out of human sentences to ingest, forcing tech giants to finally grapple with the physical world. If you want a model to understand human physiology, you have to stop reading about it and start measuring it.

Buying Bankrupt Biotech: Why Medical AI Needs More Data About Biology

Yet I remain deeply skeptical about how this will play out in practice. Gathering proprietary lab readouts or auctioning off bankrupt clinical data might give well-funded labs an insurmountable moat, but it does little for open science or independent researchers who actually build things. We are watching biology get sucked into a proprietary corporate funnel. Real engineering progress happens when builders can tinker with transparent inputs, not when a handful of tech monopolies hoard the blueprints to human health behind expensive API walls.

Building better tools requires respecting the domain, not just throwing compute at a problem until it submits. We need a massive, open reckoning with how biological information is gathered, shared, and valued if artificial intelligence is going to genuinely assist in drug discovery and clinical diagnostics. Synthetic benchmarks and scraped lab notes are a start. True breakthroughs will demand better measurements, harder engineering, and a lot less hype.