The Grind Makes the Scientist

AI can finish a research assignment in minutes. What a student learns from doing it is harder to replace.

Watercolor illustration of two researchers examining a tropical forest growing from an open notebook, with trailing roots, a toucan, and an amber-lit laptop.
Image Generated with Nano Banana 2

For twenty years, by Matthew Schwartz's account, ecologists had an equation they could not solve at scale. Rampal Etienne wrote it in 2005 to test Stephen Hubbell’s proposal that chance alone could explain the changing mix of tree species in a forest. Hubbell's neutral theory said that if you ignored every difference between species and let birth, death, and migration run as a lottery, you would get something that looked a lot like a real forest. Etienne's equation let you test that against actual counts, at least in principle. The computation was too heavy to carry out at scale, although ecologists had other ways of investigating whether the theory described their forests.

Last summer Schwartz, a high-energy theorist at Harvard, worked with Claude to build a toolkit called BootLoops for exact calculations in physics. As the toolkit grew, Schwartz prompted Claude to look for problems it could tackle beyond physics, a search that led to Etienne’s equation. Applied to census data from Barro Colorado Island in Panama, the calculation indicated that the mix of tree species there changes 4.5 times faster than neutral theory allows. By Schwartz's account, he was excited. He took the result to James O'Dwyer, a professor of plant biology and an expert in neutral theory.

O'Dwyer, as Schwartz tells it, was impressed by the technical feat and said the result would likely "be met with a shrug by many ecologists." The field had already observed, less precisely, that a lottery could not keep up with a real forest, and the exact ratio put a number on what ecologists already believed. Then O'Dwyer offered a different use for the same calculation. Subtract the neutral prediction from the data and study what remains to investigate the effects of selection, competition, and differences between species. The same calculation became the starting point for a model of what the neutral theory left unexplained.

Schwartz described the projects in a guest post for Anthropic, where he has been working as a visiting researcher. BootLoops is his project, not Anthropic's, and the ecology model he and O'Dwyer built is one of 36 manuscripts he says are still undergoing exploration and verification.

Using the collaborator he actually had

Last December, Schwartz used Claude as a research assistant on a physics problem and found it could work like a strong graduate student at twenty times the speed. Getting a paper out of the collaboration still meant correcting every sentence it wrote, pulling it back from dead ends, and steering it away from irrelevant threads. The model was capable and the collaboration was exhausting.

This summer, he began choosing problems to suit Claude’s strengths, treating it, as he wrote, “like the collaborator it actually is.” In his assessment, those strengths included broad knowledge, fluent coding, current mathematics and statistics, and the ability to read papers and appendices quickly. It could bring together methods scattered across disciplines, even though he found it unable to help with deep conceptual questions. He called the projects suited to those strengths Claude-shaped problems.

The search took him from scattering amplitudes into population genetics, linguistics, and ecology. Over three months he counts 36 manuscripts in 18 fields, all still undergoing exploration and verification. But the model's ability to find a connection was more reliable than its assessment of what the connection was worth. It favored old, highly cited debates, and in almost every project outside physics, Schwartz found that the technically correct result became interesting only after a domain expert helped steer it.

"I still don't know how to train graduate students," Schwartz writes. He has been saying versions of this since his earlier guest post, and he says the feeling has only grown more acute. The tasks he now hands to the model are also the assignments he might once have given a graduate student.

The uncertainty extends beyond one lab. A Nature paper published in January analyzed 41 million papers across six natural sciences and found that AI-associated papers included about two junior scientists rather than nearly three, a 31 percent difference. Yet its model also estimated that juniors who adopted AI reached project leadership about 1.4 years sooner, using last authorship as the proxy for leadership. The study mostly covers research from before generative AI, so it cannot show what happens when a lab adopts a workflow like Schwartz’s. Its findings suggest that some juniors may advance faster even as papers include fewer junior researchers. But authorship records cannot tell us how those researchers learned to do the work.

Weeks vs. twenty minutes

Before the work in other fields, Schwartz asked Claude to take methods from several of his own papers and related research on scattering amplitudes and combine them in a single framework. Some of those methods were written in Wolfram Language, C++, Python, or Julia, while others appeared in papers without accompanying code. The assignment was to write the missing code and reproduce the published results, the kind of work Schwartz might otherwise have given a graduate student.

The model reproduced the results from his own paper in around twenty minutes, "while the code I wrote to do it took me weeks." In the same breath, the model informed him that he had been doing something very inefficiently and that there was a better algorithm he was unaware of.

Schwartz now regards a Python course for engineers, which he would have recommended two years ago, as unnecessary. In the assignment he describes, the model supplied the missing code, reproduced the results, and found a method he had missed. A student asked to spend weeks doing the same work would reasonably want to know what they were supposed to learn from it.

Schwartz's weeks of coding had not led him to the better algorithm, but that tells us little about what else he learned along the way. Writing and checking the code may have helped him understand the calculation even if the implementation was inefficient. His account doesn't separate that learning from the labor the model can now perform for him.

Schwartz also describes learning something while supervising the new workflow. Claude's estimates of how long a calculation would take were unreliable, so over the course of the projects he developed his own sense of whether it was making progress on a sensible timescale. He learned when to let it continue and when to push it to find a better method. That judgment developed through working with the model, rather than through doing all the calculations himself.

His ability to assess the value of a result varied much more by field. Schwartz says that when Claude claims a result in his own field is fantastic, he can judge whether that is true, and often it isn’t. When it made the same claim about work in another discipline, he found himself agreeing. Recognizing that difference prompted him to seek out experts such as O'Dwyer, who could explain what ecologists already knew and where the new calculation might help.

What a student still has to do

A student using the same tools might compare methods, inspect failed runs, and reproduce results from neighboring fields much faster than before. That could leave more time to test predictions and work out why some were wrong. Schwartz’s experience suggests those opportunities are possible, but he entered ecology with enough scientific experience to recognize when he needed another expert’s help. A beginner would still be learning when to ask for that help. To serve as training, his workflow would have to give students practice in making the judgments he already brings to the research.

Imagine, for example, a student working with a forest model like the one in the opening. Before asking the AI to change a migration assumption, she writes down how she expects the mix of species to respond and why. The calculation comes back with the opposite result. She then has to work out whether she misunderstood the assumption, whether the code implemented it incorrectly, or whether the model exposes a gap in her explanation. She can use the AI to inspect the code and try simpler cases, but accepting its first plausible explanation without checking it would give her no basis for knowing whether it was right.

A supervisor could ask her to explain which check changed her mind and what she now expects in a different forest. Answering would require her to connect the evidence to her explanation and use that explanation to make a new prediction. The calculations could be finished in minutes while understanding why the results came out as they did still took time and effort. An exercise like this is only a possibility, not a demonstrated way to train scientists, but it gives us a way to distinguish the labor AI can save from the learning a research assignment is meant to provide.

Making room to learn

Faster calculations can free up time, but a student still needs opportunities to decide what to investigate, explain a result, and revise an idea when the evidence contradicts it. Those parts of a research assignment take time even when the code does not. A lab hoping to turn the speed of AI into better training would have to make room for them in the work it gives students.

Schwartz hopes humans will keep getting credit for the conceptual part of science. In the forest project, O’Dwyer brought an understanding of what ecologists already knew and what they still wanted to explain. He suggested using the calculation to investigate the effects of selection, competition, and differences between species. A student could ask Claude to follow his suggestion without understanding why those were worth investigating. Giving students time to understand how O’Dwyer chose a more useful question would help them practice choosing questions of their own. The weeks of coding may no longer be needed, but learning to decide what is worth investigating remains part of becoming a scientist.