Skip to content
Research

AI Could Transform Science. Which Questions Will It Leave Behind?

9 min read

An October 8 assessment from the DeepMind Institute raises a problem for the AI research boom: faster analysis can favor questions with abundant data. Evidence from biology shows both the risk and a way beyond it.

DeepMind scientists Demis Hassabis and John Jumper at a Nobel Prize press conference
DeepMind’s Demis Hassabis and John Jumper at the Nobel Prize press conference in Stockholm, December 2024.Jennifer 8. Lee / Wikimedia Commons · CC BY-SA 4.0
Key takeaways
  • New research points to a gap between generating scientific ideas and having the resources to test them.
  • Evidence is mixed: some research associates AI with narrower topic selection, while AlphaFold research shows expansion into previously understudied proteins.
  • For research teams and data buyers, the priority is identifying which missing observations would change a real decision.

Artificial intelligence is changing the work of scientists, from searching published research to choosing which experiment should happen next. A new assessment from the DeepMind Institute raises a question that reaches beyond laboratories: when a machine makes some problems much easier to investigate, what happens to the important problems it cannot see?

In an essay published on October 8, Alex Imas and James Manyika argue that faster research will not automatically produce a wider range of discoveries. They point to both growing experimental backlogs and the possibility that researchers will gravitate toward questions with plentiful data and established tests. Their essay is an argument about the direction of science, informed by new research, rather than a prediction that every field will move the same way. Read the DeepMind Institute essay.

The distinction matters for anyone paying for research. An organization can increase the number of analyses it completes while leaving its hardest uncertainty untouched. A model may make an accessible dataset more productive without telling its users that the missing dataset would have been more important.

What Scientists Are Actually Reporting

The underlying September study combines Gemini usage, an inventory of specialized scientific models and a survey of 637 researchers in the United States and United Kingdom. Respondents reported average time savings of almost seven hours a week. Roughly 41 percent reported a growing backlog of untested hypotheses, while about 49 percent said AI encouraged safer, more incremental projects. About 28 percent reported the opposite: greater ability to pursue riskier questions. These are self-reported findings, and the authors explicitly caution against treating the survey as representative of all scientists. Read the research paper.

That is a more complicated picture than either the claim that AI has solved scientific discovery or the claim that it merely produces confident mistakes. Researchers describe a tool that saves time and changes how they allocate it. The important question is what happens to the hours recovered, the candidates produced and the experiments waiting in line.

Consider a hypothetical materials company developing a coating that lasts longer outdoors. Its researchers might use AI to compare published formulations and propose changes. If every proposal requires a month of exposure testing, generating another hundred proposals does not by itself shorten that month. The company needs a way to choose among them, enough testing capacity and measurements that can distinguish a meaningful improvement from ordinary variation.

The commercial value comes from reaching a better decision sooner. Counting proposals can obscure that goal. Counting the number of questions resolved, the cost of resolving them and the performance of the resulting product gets closer to it.

The Data Can Shape the Question

A separate study published in Nature in January examined 41.3 million research papers. It found that AI-associated work was linked to greater publication and citation success for individual researchers, alongside a narrower collective spread of topics. The authors describe a tendency toward fields rich in existing data. This is a broad historical analysis of research patterns, not a randomized experiment establishing that a particular chatbot makes a particular scientist less original. Read the Nature study.

For a research manager, the possible mechanism is easy to recognize. One project comes with a clean database, a familiar method and a score that colleagues already understand. Another requires collecting observations from several locations, agreeing on definitions and waiting for results. The first can produce progress on a predictable schedule. The second may address a more consequential problem, but its first milestone is building the means to investigate it.

Now imagine that AI makes the first project substantially cheaper while doing little to reduce the cost of the second. Choosing the first can be perfectly reasonable for an individual team. If many teams make that choice, however, the collection of projects can become less adventurous even while each team looks more productive.

This is an incentive problem as much as a model problem. A grant, promotion process or quarterly budget that rewards readily counted output will favor what becomes easiest to count. An organization asking for more originality has to make room for work whose value will not be visible in the next reporting period.

Researchers writing in Communications Psychology made a related argument in February: shared AI methods, publication incentives and familiar terminology can reinforce one another until different projects begin to resemble one another. Their article is a commentary proposing an explanation, rather than a new measurement of how much scientific diversity has been lost. Read the commentary.

The Evidence Runs in Both Directions

There is also evidence that AI can help researchers investigate territory they previously neglected. In an April working paper, Ryan Hill and Carolyn Stein studied the effects of AlphaFold on structural biology. They found that basic research increased by 15 to 40 percent for proteins that previously lacked structural information, relative to proteins with known structures. They had not yet found a corresponding shift toward those proteins in early drug-development activity. Read the NBER working paper.

That result is an important counterweight to a story of inevitable narrowing. A prediction can make an unfamiliar subject accessible enough to study. The same class of technology that rewards abundant data in one setting can help compensate for missing information in another. Which effect dominates depends on the task and on what researchers can do after receiving the output.

Google DeepMind’s September 8 release of AlphaGenome Atlas illustrates that possibility at a much larger scale. The resource contains predictions about the molecular effects of nine billion possible single-letter changes in human DNA. Its purpose is to help researchers examine and prioritize genetic variants. Those are model predictions, not nine billion laboratory findings or clinical diagnoses. Read the AlphaGenome Atlas announcement.

A searchable collection of predictions changes the starting point of an investigation. Instead of beginning with an empty page, a team can begin with a candidate explanation. But the team still needs to decide whether the predicted effect matters for its question, whether the relevant circumstances were captured and what observation would count against the explanation.

That is why a claim about AI expanding science needs an answer to a practical follow-up: expanding which part? More possible starting points, more completed experiments and more useful products are different achievements. They can reinforce one another, but the connection has to be demonstrated.

From Candidates to Discoveries

Anthropic offered a recent example on September 23. Its researchers reported that Claude agents identified a previously uncharacterized enzyme system associated with repeating DNA sequences. The company also stated that the system’s biological function was still unknown and that human scientists performed the laboratory work. The finding is a reason for further investigation, not evidence that a new medical tool is ready for use. Read Anthropic’s research account.

Earlier work provides a different example. A study published in Nature in July 2025 described a Virtual Lab of AI agents working with human researchers. It generated 92 candidate nanobodies, small antibody-like proteins, and subjected them to experimental testing. Some showed promising binding results, making them candidates for further investigation. The study ties the computational work to a specific physical test rather than treating a generated design as the finished result. Read the Virtual Lab study.

These examples point to a useful management distinction. A team needs separate budgets for generating possibilities, rejecting weak possibilities and testing the survivors. If the first activity becomes cheap, resources may need to move toward the other two. Otherwise, the main visible result of adopting AI could be a longer waiting list.

For the hypothetical coating company, a good intermediate result might be learning that an entire group of formulations is unsuitable under a particular temperature range. That finding could save future work even if it never becomes a product. Preserving the failed candidates, the test conditions and the reason they failed would let the next project begin with better information.

The Opportunity in Missing Observations

The data question therefore extends beyond the size of a training collection. A useful collection has to contain the observations needed for the decision. Ten thousand measurements taken under nearly identical conditions may answer a different question from one hundred measurements spanning a difficult range of operating conditions. Neither is inherently better. Their value depends on the problem being investigated.

Suppose the coating company has excellent records from indoor tests but little information from coastal installations. A model could perform impressively on the indoor data and still provide limited help with salt exposure. Buying another large collection of indoor measurements would not necessarily fix the gap. The more valuable project might be a small, carefully designed effort to collect the missing observations.

That example is hypothetical, but it suggests a concrete way to evaluate a proposed data purchase. Ask which decision the new information is expected to change. Then ask what is missing from the current evidence, how the seller collected the proposed data and whether the conditions match the intended use. The answer should be specific enough to compare the cost of buying the data with the cost of collecting it directly.

The seller also has work to do. An old file becomes easier to evaluate when it includes units, dates, test conditions and an explanation of omissions. The buyer may need an expert to interpret those details. Preparation, permission to share and continuing support can affect the economics as much as the file transfer itself. A rare dataset is not automatically a valuable one, and rarity alone does not establish a buyer.

For a broker such as SnowRock, the relevant question is whether a particular company can supply information that helps answer a defined request. A broad claim about the growth of AI does not establish demand for every company’s archive. A buyer’s actual research question is a much stronger starting point.

What to Fund Next

A research organization can act on these findings without waiting for agreement on AI’s ultimate effect on science. It can examine its own project list. Which questions became easier after AI arrived? Which remain important but have received less attention? Which experiments are waiting because equipment, staff or source data are unavailable?

It can also compare the kinds of failure its process records. A rejected hypothesis with a clear explanation may be more useful to the next researcher than an attractive demonstration without a follow-up. A dataset collected in difficult conditions may support fewer publications initially while enabling a project that familiar datasets could never answer.

The practical proposal is to evaluate AI projects against the question they were funded to resolve. Track the full cost through analysis, testing and the decision that follows. Leave room for an unsuccessful experiment to produce useful knowledge. Maintain a separate budget for acquiring observations that existing tools cannot supply.

The debate now emerging around AI and science is consequential because these choices accumulate. They shape which materials get tested, which diseases receive attention and which business problems become tractable. Better models expand what researchers can attempt. The decisions about data, experiments and funding determine how much of that possibility becomes knowledge people can use.

All Insights

More from SnowRock

Destani Alvarado and Michael Wallace review a medical record on a computer at Keesler’s medical center.
Governance

When Artificial Intelligence Writes the Medical Record, What Happens to Clinical Judgment?

Sam Altman seated on stage at TED in Vancouver
Strategy

The Intelligence Beyond the Interface.

A Waymo Jaguar I-Pace with roof-mounted sensors driving on a city street
Strategy

Identifying and Scaling AI Use Cases.