Bending the Curve of Discovery: AI in Science Today and Tomorrow
Mapping how scientists use AI reveals how the technology can progress from a tool that accelerates research into an engine that expands the frontiers of knowledge.
Scientific and technological progress has long driven economic growth and led to dramatic improvements in living standards around the world. In his 2025 Nobel lecture, economic historian Joel Mokyr attributed this to a “positive feedback loop”: new technology drives waves of scientific advances that, in turn, fuel further innovation and progress. But he also warns that sustaining this positive loop is never guaranteed.
Recent history illustrates Mokyr’s point. The U.S. and E.U. entered the 1990s with relatively similar economic growth rates, but the two have since diverged substantially. The U.S. maintained roughly 2% annual increases in output per capita, largely powered by successive science and technology booms and the broad diffusion of innovation across the economy, while the E.U. has mostly stagnated.
Today, the type of scientific progress that drives growth faces significant headwinds. Across semiconductor physics, agriculture, and pharmaceutical R&D, researchers have documented a decrease in productivity that suggests ideas may be “getting harder to find.” Sustaining historical rates of innovation now requires exponentially larger research teams and capital expenditure. Decades of accumulated knowledge have led scientists into narrower specializations, longer paths to PhDs, and prohibitive coordination costs for achieving breakthrough innovations across teams. Overcoming these hurdles is critical because there are still so many open questions, along with many questions not yet asked, whose answers are key for unlocking new possibilities for human progress.
There is widespread optimism that artificial intelligence could help counter these headwinds, fulfilling a promise of computational science that researchers have pursued for more than 70 years. The journey began in the 1950s when Alan Turing and John von Neumann first applied foundational concepts of computation to the natural sciences. In the decades that followed, the field progressed from early symbolic and expert systems to machine learning, and then to the specialized and foundation models that power today’s agentic systems.
AI now has cross-domain capabilities that can substantially accelerate scientific productivity (more science and faster). Large Language Models (LLM s) synthesize ideas across vast scientific corpora, reasoning agents orchestrate advanced analyses, and specialized models predict and simulate complex phenomena at scale. In addition to being a general-purpose technology, there is a strong case for AI as an invention of a method of invention (IMI). Like microscopes and telescopes in the past, AI makes it possible to see and explore new problems in completely novel ways, or tackle old ones with fundamentally different approaches (new, different science).
AI is already helping to advance science across disciplines. Its impact spans the life sciences (such as genomics, protein science, connectomics), the physical sciences (for example, physics and astronomy), and the formal sciences (including mathematics and algorithm discovery) – as highlighted recently in the American Academy of Arts and Sciences Dædalus volume on AI & Science edited by one of this paper’s authors (Manyika). Modern AI systems have tackled longstanding problems and overcome decades-old grand challenges, from advancing mathematician Paul Erdős’ 1946 unit distance problem to protein structure prediction with the AlphaFold system, which was recognized with the Nobel Prize. These breakthroughs give us a glimpse of what may be the most distinctive aspect of the AI-for-science era: discovery at unprecedented speed, scope, and scale. We can now model the universe of possible proteins, catalog millions of potential genetic variants, simulate and perturb complex molecular systems in silico, and employ multiple agents to pursue different paths of inquiry. These advances are increasingly being translated into real-world applications in healthcare, food security, natural disaster prediction and more.
At the same time, AI’s impact on science raises concerns. Will scientific institutions be able to withstand an influx of AI-enabled research and papers? Even as it expands the possibilities for discovery, will AI narrow the scope of research to data-rich and verifiable domains? Will scientists “offload” their reading and thinking to agents, potentially leading to de-skilling and a homogenization of scientific thought? The prospect of AI-enabled science also poses foundational questions for the philosophy of science: if an AI system discovers a solution or law that humans can’t grasp or conceptualize, what will it mean for the human motivation that has, at least until now, driven scientific pursuit?
Debates about AI’s role in science range from unbridled optimism to deep skepticism, with little empirical evidence to evaluate either perspective. In our recent study, we confirm the wide breadth and scale of the use of AI in science while pointing to the challenges ahead.
Just as with previous transformational technologies like electricity and the computer, overcoming these challenges will require deep structural and institutional reforms to turn today’s early gains into sustained scientific, technological, and economic progress. And since the impact of AI will likely be greater (potentially far greater), these changes will need to be far more significant than anything we’ve seen before.
How AI is being used in science today
Building a framework for AI’s role in the future of discovery requires understanding how researchers currently deploy these tools. While LLMs dominate public discourse, how AI is being used in science is nuanced as researchers deploy general and specialized models in distinct ways across their workflows. To better understand this division of labor, our team combined data from 15 million interactions across Gemini surfaces, bibliometric data on over 2,600 specialized AI models (ranging from protein structure prediction to weather forecasting and material discovery), and a new survey of more than 600 scientists in the U.S. and U.K. We mapped these interactions against a new taxonomy of scientific tasks developed by MIT FutureTech, which provides a level of granularity previously missing from labor research and is essential to capture the work of scientists at the fast-moving frontier of their fields.
What did we find?
- Scientists use AI more than other professions; their occupations are substantially overrepresented in Gemini usage relative to their share of employment.
- Nearly half of surveyed scientists report using AI every day as part of their workflow.
- Specialized models cover all scientific fields and, in almost half of cases, are in the top 1% of citations for their domains.
- AI adoption in science is geographically concentrated, scaling directly with a country’s scientific workforce.
When we look at the specific tasks performed by LLMs and specialized models, our analysis reveals that LLMs are used for a broad and general set of tasks, such as coding and writing, while specialized models support domain-specific data analysis and modeling. The scientists we surveyed report spending over 30% of their “AI time” using specialized models like AlphaFold (predicting protein folding), Graph Networks for Materials Exploration (GNoME, which predicts possible crystal structures) and MatterGen (inorganic compound design).
Importantly, there is limited task overlap between the two classes of models. They appear to function as economic complements that reinforce each other within the scientific workflow.
Distribution of model labor across scientific tasks
Distribution of scientific tasks in Gemini logs and specialized models. Each point on the green (Gemini logs) or blue (specialized models) lines represents the share of total activity within a given high-level task, categorized according to the MIT FutureTech taxonomy. Tasks are arranged by the narrow task domain that they belong to and ordered from relative dominance of LLMs (left) to specialized model dominance (right). We use a logarithmic scale and exclude tasks that are not present in LLM logs or specialized models. Tasks within domains that account for less than 1% of total activity are grouped into an “Other” category. More information can be found in Codreanu et al. (2026).
Scientists also report substantial productivity gains from AI, saving an average of just under seven hours per week, which is mostly put back into research.
However, while some tasks - for example, quantitative, computational work - are accelerating, scientists report that the bottlenecks to scientific output are shifting downstream into physical experimentation and validation. Like in other sectors, the transition from “micro” level productivity improvements to increased “macro” outputs is not immediate. Scientists increasingly face a backlog of untested hypotheses and are spending substantial time verifying AI outputs.
The implications of AI for scientific creativity are mixed. While most researchers report greater cross-field and interdisciplinary insights, almost 50% also say AI encourages them to focus on incremental questions - a tendency that is particularly concentrated among junior researchers. This echoes a related finding that while AI-enabled science expands researcher impact (tripling publications and quintupling citations), it narrows the collective range of inquiry by directing research toward data-rich, established epistemologies rather than exploring novel, data-sparse frontiers.
From general purpose technology to invention for inventing
Our emerging findings support the idea that AI is a general-purpose technology for science. It is improving rapidly, being adopted for many disciplines and tasks, and creating the need for complementary innovations in scientific institutions like peer review and training. Its generality and power increase demand for inputs like data, compute, or benchmarks, and could transform the organization of scientific teams and workflows.
But scientific progress is not just (or even primarily) about increasing the volume of research output. For AI to bend the curve of discovery, it needs to expand the boundary of what is scientifically possible and extend the knowledge frontier. That may happen if AI turns out to function as an Invention of a Method of Invention (IMI), as some have hoped. While general-purpose technologies reduce the cost of executing existing routines - which may alone increase the space of possibilities - an IMI permanently shifts the production function of discovery itself.
Consider the following example. Experimental crystallographers spent decades resolving roughly 170,000 protein structures. AlphaFold 2 built on this foundation to predict more than 200 million protein structures - virtually all of the proteins known to science. This breakthrough led to the creation of a comprehensive, open database that has been used by more than 4 million researchers in 190 countries.
This shift in the discovery process is unfolding across science:
- In materials science and biochemistry, models like GNoME (Graph Networks for Materials Exploration) have predicted millions of potential new candidate crystal structures, and generative models design synthetic molecules de novo for drug development research, such as in David Baker’s Nobel Prize-winning work.
- In neuroscience and genomics, researchers have used AI to create a complete map of the neuron connections in the adult fruit fly central nervous system and to construct a human pangenome that is more representative of the human population.
- In the physical sciences, deep learning models are helping to control the plasma in experimental fusion reactors and bring unprecedented accuracy to weather forecasting and natural disaster predictions.
Why do these specialized breakthroughs represent an IMI? Because they act as new instruments and platforms for accelerated research and innovation by the scientific community. Models like AlphaFold have enabled datasets and novel methods that are opening whole areas of the natural and biological sciences. The true IMI shift is not just saving scientists years of wet-lab crystallography on individual proteins ("more is more"), but placing prediction inside the inner loop of large-scale search (see CoScientist and ERA) so that entirely new classes of questions can be asked and solutions delivered ("more is different").
Researchers are now clustering all 200+ million protein structures simultaneously to reconstruct deep evolutionary histories invisible to traditional methods of DNA sequence alignment. Others are running exhaustive all-by-all in silico interaction screens to uncover elusive multi-protein complexes, like the molecular bridge governing vertebrate fertilization, or querying AlphaGenome Atlas’ 9 billion counterfactuals to map regulatory networks across the non-coding genome. Indeed, basic research leveraging protein structure predictions that became available through AlphaFold increased by 15-40%, which facilitated work that may otherwise have been ignored or missed.
The friction points
Yet, accelerating science through AI still faces roadblocks. Speeding up the front end of research is triggering an inversion of the research process. Historically, the physical wet lab was the primary site of ideation, while computation served as an auxiliary aid. Today, in silico simulation is becoming a main source of ideation, while the physical laboratory is becoming a downstream mechanism of testing and verification.
Just below half of surveyed scientists report that their primary research constraints have migrated downstream into physical lab testing, wet-bench analysis, and clinical trials. This validation bottleneck creates a rapidly growing backlog of untested hypotheses and computational predictions. A recent study of AlphaFold’s impact found that while less-studied proteins did become the target of basic research, there is so far little evidence that these proteins have contributed to downstream applied research and early-stage drug discovery, with some exceptions in neglected diseases like Chagas disease and leishmaniasis. One likely reason is that understanding a protein’s biological context remains a slower, physically constrained process.
The emerging bottlenecks have at least two dimensions: first, scientists still need to interpret and extrapolate AI-driven insights; second, physical lab work remains limited by a relative lag in automated lab technologies compared to rapid model progress.
As Nobel-winning economist and AI pioneer Herbert Simon famously noted, "a wealth of information creates a poverty of attention." In AI-enabled science, the scarce resource is often no longer hypothesis generation; it is the bandwidth of the theorists who interpret the hypotheses and the experimentalists who test them. Their capacity could become a limiting factor for translating AI information into scientific progress and applications.
Downstream Workflow Frictions and Research Risk Profile
Based on a survey of N = 637 active scientists in the US and UK (Codreanu et al., 2026, Figures 13 and 14). Panel A shows the share of respondents reporting that they spend over 25% of their AI-saved time verifying outputs, that their primary research bottleneck has shifted downstream over the past two years, and that their backlog of untested hypotheses has grown. Panel B compares the share of respondents reporting that AI encourages them to focus on safer, more incremental questions versus higher-risk, more ambitious ones. The unchanged risk profile also includes the researchers who were not sure.
We already observe this in our data. Nearly half of surveyed scientists report that a good chunk of the total time they save with AI is spent on auditing, debugging, and validating AI outputs. Leaving aside the risk of mistakes, verification is important because scientists need to maintain a core understanding of the research process they are conducting. Invoking E.M. Forster’s cautionary novella The Machine Stops, our colleague Pushmeet Kohli warns against epistemic complacency - the danger that human scientists might become passive consumers of black-box models - because this could hinder our ability to understand, develop, and build on scientific output.
The model for training the next generation of scientists is also at risk. Scientific institutions rely on peer review and hands-on apprenticeship to train junior scholars in the organized skepticism and tacit "systems thinking" required to produce original work. If AI automates both paper generation and junior-level lab execution, it risks overwhelming the peer review system while hollowing out the pipeline that trains scientists to evaluate and advance the frontier.
The abundance of AI outputs also creates new cognitive challenges for scientists: when a breakthrough suddenly unleashes millions of viable structural candidates at once, how do human investigators decide which is the most promising? Overwhelmed by choice, they might default to familiar areas or domains where AI is better suited and outputs are more easily verifiable. This might explain why almost half of the scientists in our survey said that AI was steering them toward incremental, safer questions compared to about 30 percent who said it enables them to take on riskier questions.
Perhaps counterintuitively, “more” or better AI could help overcome this inertia, with AI agents exploring data and synthesizing disparate knowledge to transform massive predicted possibility spaces into prioritized and even unexpected hypotheses that scientists could inspect and pursue. In line with this, over two-thirds of surveyed scientists report that AI tools have directly expanded their access to insights outside their primary fields. AI agentic explorers and synthesizers could help counteract the tendency to look for solutions where tools make discovery easier, but overcoming the "streetlight effect" will also require deliberate incentives that reward out-of-distribution exploration.
Science today and science tomorrow
Realizing AI’s potential is an institutional challenge for the global scientific enterprise. Empirical evidence from our model inventory suggests the critical role of broad access to scientific AI models and outputs. Open weights, public benchmarks, and version-stable research APIs allow scientists around the world to audit biases, fine-tune architectures on local laboratory data, and build extensions, all while preventing the reproducibility issues caused by the constant sundowning of unpinned commercial models.
Accessible models and open structural databases also harness scientists’ tacit knowledge because they can be adapted to the specific problems they face. They lower barriers to entry for countries that lack the compute to train frontier models on their own. In line with this, Low- and Middle-Income Countries (LMICs), with the exception of China, are barely represented as developers of specialized AI models in our inventory, but they account for 10 percent of papers citing these available tools. But it’s worth noting that openness is not one-size-fits-all: in dual-use domains like biological and chemical design, the release of raw weights must be balanced against biosecurity risks, making structured access through hosted platforms with built-in safety screening essential.
This illustrates the importance of measuring how different forms of AI (LLMs, specialized models, agents, as well as hybrid systems and harnesses) are developed and adopted by scientists in different disciplines and countries, and what impact AI has on science across the world. Collecting and making this evidence broadly available can inform policies to expand who gets to do science and how to accelerate AI’s benefits for science while mitigating its risks.
Our early insights already suggest some strategic priorities for realizing the potential of AI in science:
- Invest in verification and testing infrastructure to help clear experimental backlogs.
- Incentivize moonshot exploration by reforming grant and peer-review structures to reward high-uncertainty discovery.
- Institutionalize verification and safety norms by establishing global standards for AI interpretability, reproducibility, and security screening. This should include addressing the dual-use nature of advances generally and CBRN (Chemical, Biological, Radiological and Nuclear) risks specifically.
- Reimagine scientific training and global diffusion by increasing access to models, tools, and datasets, as well as updating curricula and apprenticeship so the next generation of researchers are trained for the future of discovery. This includes expanding the open scientific commons by sustaining broad access to structural databases, version-stable research APIs, and open model weights (subject to safety standards).
- Accelerate and scale the translation of advances to address major societal challenges for people everywhere, including solutions to challenges that commercial entities may not take on.
What might the future look like? The division of labor we document offers a glimpse of tomorrow’s scientific workflow. Today, a researcher moves manually between tools: a chatbot or an agent for code and drafting, a specialized biology or materials model for prediction, and a spreadsheet or a lab notebook in between. The next step may be an agent that orchestrates the specialized models directly - planning the experiment, calling the right model for each step, cross-checking outputs, and returning a candidate answer for the scientist to judge.
A recent genome-mining study from Anthropic closely mirrors this framework. In line with the "all at once" approach, the authors used a multi-agent approach to sweep 1.9 billion protein clusters to flag an uncharacterized family of CRISPR-like viral enzymes. While an LLM identified the key DNA repeat anomaly directly in its context window using classical sequence tools, it called specialized models such as AlphaFold 2, ESMFold, and Boltz-2 to characterize how the enzyme and its partner proteins fold and interact. Much like the Virtual Lab study, where LLM agents orchestrated ESM, AlphaFold-Multimer, and Rosetta in an inner loop to design novel SARS-CoV-2 nanobodies, this workflow illustrates the complementarity between general-purpose agents and specialized models. It also highlights the downstream bottlenecks we document here: while agents generated thousands of viable candidates and vastly expanded the hypothesis space, deciding which one to test and running assays in the wet lab remained a collaborative human-AI process. The enzyme’s ultimate biological function is still unknown and requires both verification and further experimentation.
Beyond single-stream workflows, the biggest shift may arrive when artificial intelligence becomes genuinely social. Rather than reducing science to atheoretical data-fitting, early experiments show how swarms of interacting agents and humans can design their own virtual laboratories where novel hypotheses are developed and tested to discover new theories.
As AI systems handle more steps in the research process, the human job in science is expected to become more conceptual: choosing which questions to study in the first place. This shift will move the scientist away from Thomas Kuhn’s traditional "solver of intricate puzzles" and towards an "architect of profound questions." At the same time, this transition raises foundational epistemic questions for the philosophy of science, such as what counts as true scientific understanding when discoveries emerge from opaque models, how researchers extract genuine mechanistic insight from complex outputs, and how to motivate the next generation of scientists to remain curious amid increasingly autonomous exploration.
Our data shows we are not there yet. Resolving downstream verification backlogs and other emerging bottlenecks is a technical problem for researchers. But it is also a broader challenge for the scientific enterprise as a whole. The ambition, in the end, is far larger than time savings or accelerating current workflows. It is to facilitate discovery that is hard, if not impossible, for humans to do alone.
But realizing that potential and turning it into societal benefit and progress is not guaranteed. It requires scientists across academia and industry to work together in clearing critical bottlenecks while setting up institutions to maintain the ability of researchers - together with AI - to push the frontier of what is knowable. Ultimately, AI’s impact will be measured by whether that newly discovered knowledge enables society to tackle some of its greatest challenges.
Original paper: https://ai.google/static/documents/AI-in-Science.pdf