Happy Friday! I’ve created a little game called ‘The clinical trial bottlenecks game’. Naturally, it’s bottle-themed, with bubbles and literal bottlenecks. And there are leaderboards as well. You can simply go and play the game instead of reading this post. But if you want to stick around, I’ll describe the motivation behind it, some of my takeaways, and what it tells us about whether AI will cure all diseases in ten years. (For real.)
‘Bottlenecks discourse’ needs more data and modeling
Last year, Demis Hassabis said that progress in AI would cure all diseases within a decade. Dario Amodei made a similar prediction: initially saying it would cure ‘most human diseases’ and more recently, ‘most major diseases’.
Many people I know are fairly skeptical. It’s not that they think progress in AI will be underwhelming, but that they think the bottlenecks to progress in medicine are not necessarily in drug discovery, or making cures in the lab. Rather, they’re about how it plays out in the real world, including clinical trials, regulation, manufacturing and financing. Once a drug or vaccine enters human testing, it takes an average of 8 to 10 years for it to be tested in all stages of clinical trials, reviewed, approved, and finally available for people to take – and that’s assuming it’s successful at all. As I’ve written about before, this process can be sped up without compromising on rigor.
A common response I’ve heard is this: if progress in AI massively improves drug discovery, resulting in much more effective candidate drugs and vaccines, shouldn’t that also massively speed up clinical testing? If that’s the case, maybe we don’t have to worry too much about the other bottlenecks at all.
I’m intuitively skeptical of this, but I’m not sure. It is true, after all, that a more effective drug or vaccine will complete clinical trials faster, and take a smaller sample size to detect its effect. And I’ve also been thinking about how the impact of AI in medicine isn’t limited to drug discovery. It could be used in many parts of the process: to help recruit participants faster, identify better biomarkers or surrogate endpoints, or improve remote testing, to name just a few examples. And although AI can carry out complex tasks now (while sometimes getting lost along the way), it seems like a lot of the conversation is still focused on how well it makes predictions and generates images and text.
On the other hand, I wonder how much more important it will become to solve the remaining barriers if drug discovery improves rapidly. For example, Ruxandra Teslo and Adam Kroetsch have written about how high manufacturing standards and slow, sequential regulatory review in the United States slow down the process of starting early-stage clinical trials, which give scientists early knowledge that they can iterate from. In contrast, countries like Australia run these reviews in a more streamlined way, with different review stages running in parallel, meaning clinical trials can get started sooner.
But even later-stage clinical trials are fairly complicated to start, and not necessarily for reasons that have much to do with safety.
Currently, each trial is built from scratch, with its own infrastructure, contracts and administrative overhead. Before a single patient is enrolled, each hospital or research center running the trial must review the protocol, negotiate contracts, obtain ethics approval and train staff; this process takes several months for each trial. Recruiting enough participants is then another challenge: trials seek out patients who meet narrow eligibility criteria, are willing to be randomized to treatment or placebo (meaning they may not receive the treatment at all), and will return to a clinic for follow-up visits sometimes dozens of times over several years. Around half of trials fail to recruit as many participants as planned; only one in five finish on time, with median delays of over a year. On average, it costs $54 million in the US, or around $37,000 per patient, to run a single late-stage trial for a single infectious disease drug.
There are also many diseases for which the data and research to understand their biology and design better drugs is lacking in the first place. Neglected diseases, including tropical diseases and rare diseases, are perhaps the best examples of this, where the financial incentive for research and drug development is severely limited; I doubt general improvements in AI will result in new medicines for those diseases.
But bottlenecks discourse also reminds me that I’ve often written about how the same timeline can be sped up in multiple ways. What if I’m just missing the bigger picture?
So I wanted to step back and think about bottlenecks more deeply. Like, what actually is a bottleneck? Can we measure bottlenecks? And what happens if we shift one of the levers – how much does that speed up the overall timeline? Those questions were the motivation for the game and this blogpost.
An illustration of how to speed up a timeline
I want to start with a quick illustration of the problem. Take the first malaria vaccine, called RTS,S. It has an efficacy of around 30-40%, and is given as four doses during childhood. That sounds fairly low, but consider that just a single malaria parasite reaching the liver is enough to cause the disease. Since the bite of an infected mosquito releases around a thousand parasites into the bloodstream, the bar for a vaccine to have an effect at all is very high.
The modest efficacy of the RTS,S malaria vaccine was one of (multiple) reasons that funders were reluctant to pour more money into clinical trials to test it in the 2000s (even though, given how large the burden of malaria is, a modestly effective vaccine would have been hugely cost-effective in saving lives). Ultimately, the vaccine spent about 23 years in clinical trials and pilot tests before it was recommended and approved for use.
In my view, that process could have been sped up in multiple ways, such as stronger clinical trial infrastructure in Africa (in the early years, the researchers themselves had to fund some of the clinical trial sites to get them up to standards and train staff), more funding for the malaria vaccine, a different regulatory pathway, and a more efficacious vaccine.
From this example, it’s more intuitive that multiple levers might have sped up the timeline, but those levers likely vary in how much of an impact they have. And no matter how effective a vaccine candidate was, it wouldn’t have made it to the finish line and reached children without funding and clinical trial infrastructure to test it in the first place. But which was the bottleneck?
What is a bottleneck?
A common way to think about bottlenecks is that they are something that slows down a process: the rate-limiting step, limiting factor, main constraint, or whatever you want to call it. In a physical process, you might define that by where things ‘get stuck’. Another approach would be to think about which lever, when improved, speeds up the overall process the most. This is a ‘counterfactual’ approach to thinking about bottlenecks and, in my view, is more practically useful, since it tells us what would be most impactful to try to change.
What’s helpful when it comes to clinical trials is that there are fairly straightforward ways to model this. Typically, when designing a clinical trial, researchers might be constrained in a few ways. Say, for example, they want to design a trial for a vaccine to be completed within three years, expect it to have an efficacy of at least 60% against severe disease, and want a certain amount of confidence that its efficacy actually does match or surpass that figure. With those constraints, statistical modeling can help them decide how many participants they should recruit into the trial, how often they should be tested for the disease, and so on.
If we want to model bottlenecks, we can flip this around: trying to estimate, given certain constraints, how long a clinical trial would take. Then we can shift different levers to see how they would affect the timeline.
That’s what I’ve done in this interactive page and game, created with the help of GPT 5.6. I’ve somewhat loosely adapted it from real clinical trials, showing you how different design choices might have sped them up. And for simplicity, I’ve focused on a phase three trial, rather than trying to model many more inputs into drug development.
The clinical trial bottlenecks game
When you play the clinical trial bottlenecks game, the goal is to speed up the timeline of a phase three clinical trial from two years to under a year. You can change many levers to try to do it. And if you succeed, your screen will flood with bubbles – you’ve solved the bottleneck!
That will unlock a bonus game, where you steer a little boat down a river of champagne, avoid bubbles, and collect vaccine vials along the way.
It goes without saying that the models on the site are simplified, and researchers face more constraints and trade-offs while designing clinical trials in real settings. For example, they might be required to follow up the participants for a certain amount of time and test them for additional outcomes. If the disease was seasonal, they might aim to check whether the vaccine worked across multiple seasons, not just one. (Though I wanted to use the malaria vaccine as an example, the disease is strongly seasonal and would’ve complicated the game even more, so I decided to go with a conceptually simpler example with the human papillomavirus.)
In the game, you can make several design choices to shorten the trial, but they would make it more expensive – raising the sample size by recruiting more participants would be expensive, for example, even though the reward to society for getting effective drugs and vaccines available sooner is large.
Since I wanted to give you a sense of how some of these trade-offs might work, I’ve given you coins to ‘spend’ on improving each of the levers. Unfortunately, there’s little data available on the costs of changing these different levers, or how their impact changes with more spending. (More research on these questions would be very valuable.) So instead, the model assumes that each of them faces diminishing returns; it becomes more expensive to continue to improve each lever.

Consider how, at first, it might be easy to improve recruitment by investing in simple campaigns and finding people already fairly willing to participate (think people who are bored and enjoy participating in scientific experiments for fun, like me when I was a university student); but that eventually, you would have to spend more effort trying to recruit people who have higher opportunity costs, like me now… :(
Some insights about bottlenecks
I’m hoping the game will help build an intuition of how bottlenecks work and how to solve them. I recommend playing around with it, but I also want to share some takeaways I had while designing and playing it.
The same timeline can be sped up in multiple ways. You can speed up the same clinical trial by increasing its sample size, recruitment rate, choosing an earlier surrogate endpoint, running the trial in a setting where the disease is more common, reducing dropout, or making other decisions. After you complete the challenge, the game shows you how different people completed it – in a lot of different ways!
Diminishing returns are everywhere. Improving a lever by the same amount eventually saves less time. The statistical goal of a clinical trial is to see enough cases of the disease in the trial that researchers can see if there’s a difference in how they’re spread between in the vaccine group versus the control group. Suppose a trial needed to see 100 cases of the disease. If there were 10 cases per year, it would take 10 years to reach the endpoint. Increasing the rate to 20 cases per year would cut the trial’s duration to five years, saving five years. Increasing it again to 30 cases per year would cut the duration to 3.3 years, saving only another 1.7 years. Though you’re increasing the lever by the same amount, since the target remains fixed, this results in diminishing returns.

Here is a simple model I made to show how the disease incidence rate and sample size affect the length of a vaccine trial. You can recreate this chart with my code on GitHub. Some levers have a much larger impact than others; their impact depends on their level and the curve. As you’ll see in the game, some levers start very high to begin with. Trying to improve them further is expensive and generally has little additional impact. This follows from the fact that they face diminishing returns in terms of both the lever’s impact and in how much you can continue to move it with spending.
We live in a multi-causal world, where causes aren’t independent. When a drug succeeds in clinical trials, you could attribute its success to many different people and interventions. Was it because of the scientists who studied the disease to begin with, or those who developed the drug? The statisticians and staff who designed and ran the trials, compiled the data, and ran the analyses? Or the healthcare workers who ran the tests to detect the disease? Maybe it’s the staff who kept the clinical sites running, or the construction workers who set them up in the first place. Or perhaps it’s the participants in the trial, since they made it possible to test the vaccine at all. My answer is that all of them contributed in some way to making it happen, and you can’t neatly divide up their impact like slices of a pie. But at the same time, some individuals had a much larger counterfactual impact on the drug’s success than others; without those individuals, the drug would have been much more likely to end in failure.
The bottleneck can shift over time. Once one bottleneck is resolved, the process gets rate limited at another part of the process. If recruitment is slow, that may be what slows the trial down initially. But once it’s sped up, you would still have to wait for enough participants to develop the disease (in the control group) to get to a result. My colleague Willow has a neat example of this as it relates to energy and economic growth. She says:
“Effective energy can become the binding constraint either due to supply issues (e.g. the 1970s oil shocks or the current Hormuz situation), or if the labor/capital deployed in energy-using industries increases faster than supply can respond (e.g., due to AI and electrification). However, once the bottleneck resolves and energy is again a small share of GDP, available labor and capital again become the binding constraints.”
I wanted to make this last point clearer, so I added an additional bar chart to help you see exactly which levers were bottlenecks in the example. Once you improve a lever, you can see how it shifts where the bottlenecks are, often moving them to other levers.
There were also a lot of other concepts I started thinking about while designing the game. Unfortunately, they didn’t quite fit, partly because of a lack of data to model them, partly because I didn’t have time. But maybe in the future there’ll be a version 2.0.
Some levers are likely much more expensive to change than others. In the game, because of the lack of cost data, I’ve set all the levers to have the same cost curve. But in reality, different levers might have very different costs. On average, it costs about $37,000 per patient in a clinical trial. Increasing the sample size might then be easy to do, but also very expensive. In contrast, remote-testing – such as with rapid tests – might be much cheaper, and could help you achieve a faster trial. Another option might be to move to a setting where the disease is more common and it takes less time to see cases of the disease.
Different interventions have different costs. Even within the same lever, different interventions can have substantially different costs. A national volunteer registry, where people express their interest in being contacted by researchers for eligible trials, could make it much faster for trials to recruit participants, and it would be fairly cheap to set up. But it would have to be done at the national level. On the other hand, a clinical trial network, advertising, or compensating participants could each improve recruitment, but are likely more expensive. In the game, I’ve given each lever the same illustrative cost curve because there’s limited public data on the costs of these interventions.
Some interventions help speed up the current trial; some help speed up many trials. Just as in the previous example, some interventions would speed up many trials, not just one. They include: clinical trial networks, volunteer registries, and standardized contracts. But researchers running individual trials would likely not find them worth the investment. For a pharma company, a CRO (which operates clinical trials), would likely fulfil some of the same functions, being able to pull from multiple sites they have contracts with.
The costs of interventions aren’t fixed. If you remember life before the pandemic, you might remember that rapid testing for infections wasn’t much of a thing back then. Until 2020, it might have seemed prohibitively expensive and questionable to use remote-testing in a clinical trial; could you really expect participants to perform a PCR test on themselves at home? Things are different now, since rapid tests have been developed, people are used to them, and the tests themselves have been much more validated. But this is surely not the only example. We can likely find cheaper ways to improve other aspects of clinical trials.
Conclusion: Who’s right about bottlenecks?
All in all, the model and game helped shift my views on some fronts. First, improving the efficacy of drugs and vaccines in the pipeline – which you can use as a sort of stand in for better biology and drug discovery – really can speed up the timeline a lot. If our biological understanding of diseases and ability to develop better drugs really did improve substantially, that would be quite significant in shortening clinical trial timelines! It would also shift the bottleneck to other factors, in particular, how long the outcome takes to develop (which you could still circumvent by massively increasing the sample size, at a cost), or other factors that I didn’t include in the game, like funding for research and drug development, clinical trial operations, and manufacturing costs and time.
Second, surrogate endpoints, meaning early proxies for the disease, could also speed up the timeline substantially. This is probably the ‘bull case’ for those who believe AI will help cure diseases fast. If it develops better surrogate endpoints, we can use them to iterate, develop better drugs, and speed up drug development without needing to reform clinical trials.
But then, sadly, my views got tempered again. Getting drugs approved on the basis of surrogate endpoints requires testing and validating those endpoints first – to see if reducing them really does reduce the disease as well – and that process can be very time-consuming.
There’s an interesting example of this in the news right now, with several new drugs that aim to reduce cardiovascular disease. They target lipoprotein(a), a cholesterol-carrying particle that is thought to be particularly harmful. The evidence for this comes from large observational studies, genomic studies, and Mendelian randomization studies. In some ways, I’d say they’re ‘as good as it gets’ when it comes to the strength of evidence from observational data. But it’s still unclear if the effect will be causal. That hypothesis is currently being tested in clinical trials that run for several years with thousands of people, which aim to test not only whether the new drugs reduce lipoprotein(a) but that they also reduce cardiovascular disease, as measured by heart attacks, strokes and other events. If they do, it will likely mean future drugs can more simply be approved by showing that they reduce lipoprotein(a), without having to wait and test their effect on cardiovascular effects each time.
The first major trial testing this, run by Novartis to test their Lp(a) drug called pelacarsen, was recently completed and was unsuccessful. Although the drug lowered lipoprotein(a) by around 85%, it did not significantly reduce cardiovascular events, even after six years in its phase three trial with more than 8,000 participants. Other drugs are still in the pipeline, and I’m still somewhat hopeful. But it’s a reminder that you can’t just assume that endpoints will turn out real even with strong observational and genetic evidence.
Then of course there’s everything else: will there be funding to study all the neglected diseases enough to improve our understanding of their biology and develop better drugs for them? And even if we do develop cures, are they even guaranteed to be available?
Ultimately, I’m skeptical that AI will cure all, or even most, diseases within 10 years. In order to test our hypotheses and iterate on them, we’ll continue to need clinical trials that actually allow us to see the effect of new drugs on diseases. That is bottlenecked by the time it takes for diseases to develop, even if we aim to switch to surrogate endpoints, as well as the funding, infrastructure and regulation surrounding clinical trials.
But I think there’s a lot of room for AI to be useful. Beyond improvements in drug design and discovery, we should consider how it might be used in other areas of clinical trial reform: improving participant recruitment, building evidence for surrogate endpoints, improving remote testing, and more. If AI makes drug discovery much faster, the other constraints won’t disappear; they’ll become stronger. We should be thinking now about how to solve them too.




Creative and impressive. Thanks.
A tangential thought: I tip my hat to the necessary and productive work done by thousands to improve the clinical trial world from within "the paradigm." Undoubtedly my own personal health has a lot to thank for that. I'm also really eager to see some profound disruptions to it. I'm off on my own adventure trying to deliver our own disruption, but more broadly I hope that there are some successful game changers that allow us to re-envision the entire process more fundamentally - without sacrificing safety, efficacy, cost, time, etc. Perhaps those disruptions shift the operating envelope out favorably. Though all processes will still suffer the same constraints/bottlenecks you map out here, perhaps they can do so in an operating space that finds, for at least some subset of health, a new local environment to mine.