When AI Can Discover a Million Cures, Which 100 Do We Test?

Dario Amodei may be right about AI and still be wrong about medicine. His argument is that intelligence is about to become abundant. Millions of AI scientists could read the literature, generate hypotheses, design molecules, plan experiments and work continuously at a speed no human research organisation can match. I think we should take that possibility seriously.

So assume he is right. Assume AI scientists become better than most human scientists. Assume automated laboratories improve. Assume virtual cells become useful. Assume AI compresses years of early drug discovery into months. Even then, curing most human disease in five to ten years does not necessarily follow, because intelligence and evidence do not scale in the same way.

AI can manufacture hypotheses. It cannot manufacture ground truth.

For most of scientific history, good ideas have been scarce. There are only so many scientists, only so many hours in a day and far more biological possibilities than humans can investigate. AI could reverse that scarcity. We may eventually have far more credible hypotheses than the scientific system can physically test.

DeepMind has already described the beginnings of this as a validation bottleneck. As AI becomes better at generating scientific conjectures, our ability to verify them does not automatically improve at the same rate. Medicine makes this problem especially difficult because validation is not purely computational. It consumes laboratories, biological samples, manufacturing capacity, clinicians, capital and eventually patients.

Imagine an AI system generates one million genuinely plausible cancer treatments. Not hallucinations. One million interventions supported well enough by existing biology that each deserves investigation. We will not test one million of them. We may not even test one thousand.

This changes the economics of science. The question is no longer only, what can we discover? It becomes, what deserves one of our scarce opportunities to find out whether it is true?

The obvious answer is that AI will rank the candidates too. It will use genomics, mechanistic models, animal data, biomarkers, simulations and previous clinical outcomes. That will make the selection process much better. But all of those remain representations of reality. The reason experiments exist is because our representations are incomplete.

At some point prediction ends and measurement begins.

This matters because biomedical development already loses most candidates before approval. Even if AI dramatically improves those odds, it does not make biological uncertainty disappear. A system that moves success from one in twenty serious candidates to one in five would be extraordinary. It would still mean that four of every five serious bets consume time, capital and experimental capacity without becoming medicines.

Now scale candidate generation by 100 or 1,000 times.

The constraint moves.

An AI can generate another molecule in seconds. It cannot generate another patient with a rare disease. It cannot manufacture another person with precisely the mutation required for a clinical study. It cannot turn five years of survival data into five minutes. In rare diseases and narrow biomarker populations, eligible patients are already a limiting resource.

This suggests the next important scaling variable in science is not intelligence.

It is reality throughput.

Reality throughput is the rate at which scientific ideas can encounter enough high-quality evidence to determine whether they are actually true. If AI drives the marginal cost of generating hypotheses toward zero while experiments remain expensive, physical and slow, the value of each opportunity to test reality rises.

That is the real consequence of scientific abundance.

We may soon have too many good ideas, not too few.

This also means we may be benchmarking AI scientists incorrectly. Instead of asking a system to generate 1,000 novel hypotheses, give it 100,000 credible hypotheses and a budget for only 100 experiments. Ask it what to test first. Reveal the result. Then ask again.

The metric should not be how many ideas the AI produces. It should be how much truth it extracts from a fixed amount of reality.

That is a very different kind of intelligence.

The best AI scientist may not be the system that generates the most discoveries. It may be the one that wastes the fewest experiments.

Now return to Amodei’s idea of a country of geniuses in a datacenter. Imagine one million superhuman cancer researchers working continuously. Then add another million. Then ten million more.

Eventually the datacenter contains more intelligence than the physical world can answer.

The scientists form a queue at reality.

This is not an argument against Amodei’s optimism. AI may become the most important technology medicine has ever acquired. But every technological revolution removes one scarcity and exposes another. The internet made information abundant and attention scarce. Generative AI is making cognition abundant and judgment more valuable.

If AI makes scientific hypotheses abundant, ground truth becomes scarce.

And if that happens, the next frontier in AI for science will not be building systems that can think faster.

It will be building systems that know which pieces of reality are worth spending next.

The Defection Graph: The Best Pre-Seed Signal Is a LinkedIn Departure Date From a Lab You Don't Cover

By the time a frontier AI researcher’s departure reaches TechCrunch, the best-connected investors may already be months into the relationship.

That is the problem with treating talent movement as news.

News tells you what happened. Sourcing infrastructure should tell you what is beginning to happen.

Frontier AI labs have become some of the most productive founder factories in venture history. Anthropic came out of OpenAI. Safe Superintelligence was founded by former OpenAI chief scientist Ilya Sutskever and reached a valuation above $30 billion before launching a product. Thinking Machines Lab, founded by former OpenAI CTO Mira Murati, raised $2 billion at a reported $12 billion valuation in July 2025.

The capital was not underwriting revenue. It was underwriting the people who left.

THE DEPARTURE IS NOT THE SIGNAL

A single researcher leaving a lab is noise.

Three researchers leaving the same twelve-person team within six weeks, after years of working and publishing together, is different.

It might be a reorganisation. It might be coordinated recruiting by a rival. Or it might be a founding team assembling before incorporation.

The data already exists to distinguish between these possibilities.

Researchers are nodes. Co-authorship, shared projects and team membership are edges. Departure timing, seniority, technical ownership, new domains, corporate registrations and open-source activity are features.

The objective is not to label someone a “future founder.” It is to estimate how the probability of company formation changes as new evidence appears.

One unexplained departure creates a small increase in probability. A second departure from the same team increases it further. A new domain registration, company filing or renewed collaboration among former colleagues raises it again.

This is not a static founder score. It is a time-to-event problem.

THE RAW LIST WILL NOT BE THE EDGE

The data layer is already being productised.

Evertrace monitors company registries, GitHub, patents, grants, domains and social activity. Specter tracks new founders, stealth hires and job changes. Dealroom has built a talent graph covering millions of companies and profiles.

That means the departure list will become available to everyone.

Buying access to the same list as every other fund is not alpha. It is software procurement.

The real advantage moves to three things: scoring, timing and access.

First, how likely is this person or cluster to form a company?

Second, how likely is that company to attract institutional capital?

Third, at the valuation the market is likely to assign it, will the investment generate venture returns?

These are different questions.

Prestigious lab experience may predict fundraising success while also producing an entry price that absorbs most of the upside. A model that only predicts who will raise money may consistently find the most expensive deals, not the best ones.

THE MODEL DOES NOT GET THE MEETING

Even a good model only creates a timing advantage.

Knowing that three researchers quietly left the same team is not particularly useful if your first interaction is a cold LinkedIn message after the system flags them.

The funds that win these deals already have relationships with the researchers, their former colleagues, professors, founders and angel investors.

The system tells them where activity is forming.

The network gets the meeting.

Reputation converts the meeting into an allocation.

THE BOTTOM LINE

Build the defection graph as infrastructure, not as a Slack alert.

Continuously ingest career transitions, collaboration histories, company registrations, domains, patents and technical activity. Score clusters with a time-dependent model. Test separately for company formation, fundraising probability and investment returns.

Then build relationships with the people the model suggests are most likely to move, before they have decided to raise.

The list will become available to everyone.

The enduring edge belongs to the firms that interpret it correctly, act before the signal becomes obvious and already have the phone number.

The Last Hours

It's 1 AM. You're still awake. Not because of a deadline. The Slack stopped hours ago. You're awake because this is the only hour of the day that belongs to you.

That's not a sleep problem. That's a resource allocation problem.

THE PERFORMANCE OF THE DAY

From the moment you wake up, you are someone's version of you.

You're an answer to an email. A face on a call. A founder projecting confidence. A co-founder being fair. A boss being decisive. A son who should call more. A partner who has been distracted. A friend who keeps canceling.

Every hour of the day is shaped by what other people need from you, what role you're playing in that particular scene. By 10 PM, you've been someone else's version of you for fourteen hours straight.

That is not work. That is a sustained performance of self.

And the body knows it needs somewhere to take the costume off.

MIDNIGHT IS WHEN THE COSTUME COMES OFF

At midnight, no one needs anything from you. No metrics to tend. No posture to hold. No gap to close.

You are not, for now, a founder or a boss or a son or a friend.

You're whoever you actually are. Which, if you're honest, you've been too busy to check in with lately.

That's why you stay up. Not insomnia. Not procrastination. The specific relief of being unobserved.

THE REAL BURNOUT MODEL

Founder burnout is not from working too hard. It's from never being off.

From living eighteen hours a day in a state of mild emergency, always slightly on, always slightly accountable. The exhaustion is not the hours. It is the performance of the hours.

The night is the pressure valve. The only place you're free.

The question is not how to go to sleep earlier. The question is why the night is the only place you've found this and what it costs to keep stealing it at 1 AM instead of building it into the day.

WHAT TO DO ABOUT IT

The night will not always be enough. You already know this. The drag behind your eyes by Thursday. The afternoon you can't think clearly.

You don't need a better sleep schedule. You need more pockets of unobserved time inside the daylight.

A Sunday morning with no agenda. An actual lunch with the door closed. Being honest with the people around you that you need hours each week where nothing is expected.

Or sometimes it just means staying up until 2 AM because you've earned it.

Just don't let the night be the only place you can breathe.

THE BOTTOM LINE

If you're reading this at midnight, I'm not going to tell you to sleep.

I know what you're doing. I know why you're here.

But the version of you underneath all the roles, try to give it a little daylight too.

The night has been generous. It doesn't need to do all the work.

Stage-Agnostic Sourcing Is Systematically Miscalibrated


The standard pitch for ML-driven VC sourcing is correct in the aggregate and wrong at the margin where it matters. Hiring velocity, web traffic, GitHub commits, app store ratings, these signals outperform gut intuition across the full universe of startups. The problem is that the full universe is not your portfolio.

Stage-agnostic scoring models are systematically miscalibrated because the signals that predict success at seed have a different shelf life than the signals that predict success at Series A. Using a stale signal isn't neutral. It's noise dressed as insight.


THE DECAY CURVE IS REAL

Hiring velocity is the clearest example. At seed, a team that is aggressively hiring before revenue is a leading indicator of founder conviction, capital efficiency expectations, and organisational ambition. The signal is fresh and predictive, founders who build fast teams early tend to build fast companies.

By Series A, hiring velocity has largely decayed as a signal. The companies that are hiring fast at Series A are often the ones that raised a lot, not the ones that are executing well. Capital causes hiring, not the reverse. The signal that was predictive at seed is now confounded by the round itself.

The same logic applies in the other direction. Revenue retention like net dollar retention, expansion rate, churn cohorts is meaningless at seed because there is no revenue to retain. It's a powerful signal at Series A and a critical signal at growth. Stage-agnostic models that include revenue retention in the seed-stage scoring layer are not ignoring this fact, they are giving it zero weight implicitly, which is different from having thought about it explicitly.


WHAT STAGE-STRATIFIED STACKS LOOK LIKE

The funds winning the sourcing game are not running one scoring model. They're running signal stacks calibrated to stage, with deliberate choices about which inputs are predictive at which moments.

Seed signal stack: Team composition and background data (founders who have built before in the domain), early hiring decisions (which functions they hired first, how fast they hired after funding), product shipping cadence (GitHub activity, product launches, beta user counts), and network graph positioning (who is on the cap table, who is advising, who they have access to).

Series A signal stack: Revenue trajectory and retention shape (not just the number, but whether the curve is accelerating or flattening), sales cycle characteristics (ACV relative to sector, time-to-close relative to cohort), product engagement depth (not DAU, but whether the product is in the critical path of the customer's workflow), and competitive positioning relative to named alternatives.

The transition between these stacks matters as much as the stacks themselves. There is a window between seed and Series A, typically between $500K and $2M ARR where the seed signals are losing predictive power before the Series A signals have enough data to be reliable. This is the hardest scoring environment. The funds that perform here supplement quantitative signals with qualitative inputs they've structured: reference calls that follow a consistent schema, founder conversation recordings analyzed for linguistic patterns, board update language that gets embedded and monitored.


THE PRACTICAL IMPLICATION

If your sourcing model was trained on a corpus that includes seed, A, and B outcomes without stage stratification, it has learned signal weightings that are incoherent. It is applying the predictive logic of seed-stage signals to growth-stage companies and missing the predictive logic of growth-stage signals at seed. The model is not wrong — it is systematically miscalibrated in ways that are hard to see in aggregate and expensive in individual decisions.

The fix is not to discard the model. It is to stratify it. Train separate models for each stage, or apply stage as a conditioning variable that reshapes signal weights rather than treating it as one feature among many.

The funds that have done this will tell you it is a materially different output. The companies that score well on a seed-stage model look different from the companies that score well on a stage-stratified model at the same funding level. That difference is where the alpha is.


THE BOTTOM LINE

Alt-data sourcing works. The evidence is solid enough that this is no longer the argument. The argument is whether you are running a model that understands signal decay or one that treats all inputs as equally predictive regardless of when in the company lifecycle they're measured.

The signal decay curve, the half-life function of each alt-data input as a function of company stage is not a theoretical refinement. It is the difference between a sourcing model that finds the companies that already look like winners and one that finds the companies that will become winners. At seed stage, those are very different lists.

Your signal is fresh at seed. It's stale at Series A. Price them accordingly.