Can AI really think like a scientist?

by Ashutosh Jogalekar

AI is collapsing barriers in scientific disciplines, most notably in math. An 87-year old conjecture named the Jacobian conjecture was recently solved with help from AI. An AI model solved another 80-year old conjecture, the distance conjecture. Many Erdős problems, named after the prolific late mathematician Paul Erdős, are swiftly succumbing to AI. Some mathematicians like Terence Tao – regarded by some as the greatest living mathematician – have simply come to terms with using AI in their research and think that AI will be as integral to math as were slide rules and calculators. Nor is math the only field that is is seeing striking inroads made with AI. Other theoretical disciplines, including theoretical computer science, are also being rattled. Theorist Henry Yuen sounded a bit shaken to see how many problems in his discipline were being solved by AI, and said that he could “feel the importance of these problems in his bones.”

Given this spectacular progress of AI in math and science, it does not seem unreasonable to ask the next question: how long before AI solves outstanding problems in all sciences: physics, chemistry, biology, astronomy, geology? And then how long before it also solves outstanding problems in other adjacent disciplines: neuroscience, psychology, anthropology, sociology, philosophy? Will we soon live in a post-human world, at least so far as science is concerned, which considering that science is what helps us make sense of the universe, means a post-human world, period?

First of all, let’s state the obvious. Any kind of optimism that AI will solve problems in applied sciences like chemistry and biology anytime soon is misplaced for obvious reasons: unlike math and theoretical computer science and even parts of theoretical physics, these disciplines need experimental verification and the right data. Unlike math, you simply can’t think your way to a new drug or material. In fact if you look at most of the Nobel-winning discoveries in these disciplines, you will realize that while AI could likely have helped accelerate them by generating ideas, most of them involved the serendipitous discovery of new biological molecules or mechanisms. An equal number involved the invention of new tools and techniques. It’s hard to see how AI could have possibly done this by itself. At least in the near future, I don’t see how AI can simply think its way to a new cancer drug, a new mechanism of memory or a new technique to unravel the mysteries of life, hype from AI leaders notwithstanding.

But beyond these specific domains, the entire debate around AI revolves around a more profound question. There are effectively two camps of people when it comes to thinking about AI and scientific discovery. One camp believes that discoveries made by great scientists can never be duplicated by AI, because AI will always miss the spark of human creativity, the leaps of intuition, and because it can only work off data that it has. The second camp acknowledges this, but its basic premise is that a lot of what we call human creativity is just pattern recognition and pattern matching, so there’s nothing in principle that stops AI from duplicating it.

Personally, I think that as tends to happen in many of these cases, the truth is more interesting but perhaps also more disturbing. Let’s first tackle what the second camp says. Is it true that a lot of creativity flows from just pattern matching and pattern recognition? First of all, I think it’s important to acknowledge that those words, “pattern recognition”, are doing a lot of heavy lifting because there’s degrees of pattern recognition and pattern matching. There’s fairly simple pattern recognition and there’s much more sophisticated pattern recognition and matching, which depends on finding higher-order analogies.

There’s a great quote by the mathematician Stanislaw Ulam: “Good scientists find analogies between things, while great scientists find analogies between analogies.” And I think that’s very much true. I don’t want to ding pattern matching and recognition, but I think it is important to acknowledge that a lot of what we call great scientific discovery was, in a sense a lot of clever curve-fitting and data fitting.

Probably the most famous example that I can think of is Watson and Crick’s discovery of the structure of DNA. I don’t mean to imply that Watson and Crick were not brilliant or creative, but I think it is important to realize that the real secret of what made their approach work was that they used data from a variety of different sources to build a model of DNA. For instance, they used Erwin Chargaff’s rules which said that the ratio of purines to pyrimidines is constant in all living organisms. They used data from Jerry Donohue that showed that all four bases of DNA exist in certain chemical forms called the keto forms instead of the enol forms. And they also used data, as is well known, from Rosalind Franklin and Raymond Gosling’s famous X-ray structure. But the point is that the other people were like the blind men and the elephant. Each one was looking at one part of the puzzle, but nobody had all the data to put it all together. Nobody else was as curious. Nobody wanted to really build a model. That’s what Watson and Crick did and why they succeeded.

So the question in my mind is: if an AI were given the three or four critical pieces of data that Watson and Crick had, would it have been able to build a model of DNA? The data was all there by 1953. Somebody should do that experiment (although they should use the data in the exact form that it was available to them). And I think the answer might be uncomfortable to a lot of people, but I wouldn’t be surprised if it’s a yes. If building models was the key to the discovery of DNA structure, it was essentially a combinatorial exercise of fitting those models to a lot of data, something AI generally excels at. The discovery of DNA’s structure was one of the greatest in the history of science. But as uncomfortable as it sounds, I think an AI could have managed to do it in principle.

Let me state another example of a discovery where AI could have managed to find a solution: Planck’s discovery of the quantum. This was what started the whole field of quantum mechanics. Again, undoubtedly one of the most important discoveries in the history of science. What is interesting, though, is that Planck, while a very capable physicist, was not considered a genius like Einstein or Bohr or Dirac or Heisenberg. And that’s because in his own words, what he did was effectively sophisticated curve-fitting, “an act of mathematical desperation”. There was a lot of data on blackbody radiation that was available, and this data showed some anomalies. What Planck did was to come up with a mathematical formula to describe this radiation, and he found out using this formula that it only works if energy comes as little packets and if the energy of each packet is famously, as we know, related to its frequency through Planck’s constant. So a lot of what Planck did was also data fitting. The question then is: Can AI do the same thing if it were given the same data on black body radiation? And again, the uncomfortable answer might be yes and it’s probably worth doing the experiment.

In that sense, the second camp which says that a lot of what we call human creativity and great discoveries is based on fitting a lot of data and pattern matching and data analysis makes a valid point. Probably the ultimate example of data fitting – and this is where I’m not entirely sure AI would be able to do it – is Darwin’s theory of evolution. After coming back from his voyage of the Beagle, he sat on the data and collected new data, both from his friends and colleagues as well as through his own experiments, for twenty-five years. And then he came up with his theory of evolution. The ideas were already in the air. There were a lot of predecessors of Darwin who had proposed some kind of evolution, but he was the first one to really exhaustively look at the data extremely ploddingly, extremely patiently in a way that was unprecedented (as he put it, his mind seemed to be a mechanism for grinding out general laws from specific instances). Nobody then probably could have done it that way and derived the right conclusion about the mechanism of evolution, which is natural selection. So the question again is: If AI had access to all of that evolutionary data, could it have derived some form of natural selection mechanism? And I don’t know the answer to that question, but again, that would be a good case study, if you will, to find out.

Let’s be clear about the claim we are making here because it’s pretty provocative: Some of the greatest scientific discoveries in the history of humanity were made largely by fitting data to a model and abstracting a general set of laws or a theory from it. There is a good chance that AI can, sooner or later, make epochal discoveries using similar data analysis.

One last example from modern times. When I was in middle and high school, the International Math Olympiad was regarded as the epitome of intelligence and creativity. Fellow students who even made it to the final stage, let alone won gold, silver or bronze medals, were regarded as mathematical rock stars. Now we find that the latest AI models can equal or even exceed the performance of students on the Math Olympiad. What does that tell you about how much even complicated mathematical problem-solving is about genuine leaps of insight versus highly practiced pattern matching? This does not take anything away from the diligence and hard work and intelligence of the students who do well on that test, but it does lead to serious questions about how much actual creativity is involved in tackling those kinds of problems.

So, all of this really somewhat argues in favor of the second camp of people who say that even great discoveries are not some kind of mysterious phenomenon, but they can be reduced more or less to pattern matching, even if that pattern matching is very complex and cannot be done by everyone. One of the most disconcerting – if also enlightening – developments in the history of science is how it has regularly dethroned humanity from what we thought was our special place in the universe. The Copernican revolution did that, evolution did it, cosmology did it and now perhaps AI will do it as well (although the ultimate irony of AI is that unlike these other things, it’s our own creation that would have ousted us).

Let’s shift gears. The problem is that the first camp is also right because there are certainly some discoveries which cannot lend themselves to pattern matching. They really seem like they had a very deep degree of intuition associated with them, a kind of magic that still leaves us struggling to explain exactly how they came about. I am going to cite two well-known examples.

One is Bohr’s model of the atom. When Niels Bohr was building his model of the atom, his central tenet was that electrons can exist only in certain stationary states. They cannot exist in any arbitrary energy states or a continuum of energy states. If you look at the history, there’s really not a lot of rational background for why Bohr thought to come up with this hypothesis and test it. In fact, he himself called it a leap of faith. And he was astonished, as was the world at large, when after making this fundamental hypothesis and assumption, Balmer’s formula for spectral lines just popped out of Bohr’s work. Now Balmer’s formula itself is very interesting because it was not based on extensive data fitting. In fact, Balmer only fit eleven spectral lines that the scientist Ångström had measured. He fit those to an equation and obtained his formula. And Bohr then came up with this leap of faith that led to Balmer’s formula just sort of dropping out of the equations. It’s not clear at all that AI can do something like that because 1. The data was lacking, and 2. The mechanism by which Bohr came up with his insight even now seems rather mysterious. And so, until we have insights on that mechanism and can really frame it within a framework of rationality, it’s not clear how AI could duplicate anything like that. That’s the first example.

The second example is, of course, Einstein. The general theory of relativity started with the equivalence principle. Now, the equivalence principle is something that Einstein just sort of dreamt about as a thought experiment about a man falling and not feeling his own weight. He called it the happiest thought of his life. And then it took him five years of sweat and toil and tears for him to discover, helped in no small part by his friend, the mathematician Marcel Grossman, who pointed out that tensors and Riemannian geometry might be the necessary framework with which to describe curved space-time. So, it was a pretty messy process that involved a huge number of flashes of intuition and creativity and false starts. How do you capture that in an AI? Because it really goes toward Ulam’s quote about finding analogies with different analogies. This was not data analysis. Einstein had some data like the perihelion of Mercury that he had to fit, but this was by no means any kind of curve fitting or data analysis.

So, my conclusion is that I do think that there’s great scientific discoveries that were significantly informed by data analysis and curve fitting. There are also clearly discoveries where it seemed like flashes of human brilliance and intuition and creativity really played a huge role. But the fact that there are great discoveries that were based mainly on pattern matching and model building is a tantalizing hint for AI. The reason AI is exciting right now is precisely because of those questions about whether it can recapitulate great discoveries that are in fact based on pattern matching or pattern recognition. And we could run some of those experiments now as I pointed out in the case of DNA and blackbody radiation.

If AI can compress a huge number of patterns into a repeatable rule, it would essentially lead to what we call a law. If AI achieves just that and nothing more, it would already be a revolutionary achievement. We will see what the future brings.