Whoa Now: Cautionary Tales from Materials Science

5 min read Original article ↗

I have a reputation - not undeserved, either - as a skeptic of AI approaches in drug discovery. There is just so much hype, and so much misunderstanding. But that doesn't mean that I am reflexively hostile to computational and machine-learning approaches in general. There are too many counterexamples where it really has proven useful. But as I've said many times, these tend to be in areas with (1) a relatively well-bounded set of questions you're hoping to answer and (2) a large amount of high-quality experimental and physical data for your ML approach to work with. By "high-quality" I mean gathered under conditions which are as internally consistent as possible and which include both positive and negative data (when appropriate to the experimental design). 

Materials science is an area that can fit these criteria in a very attractive way, and there accordingly have been a lot of ML efforts on metal alloys, ceramics, polymers, metal-organic structures, etc. The experimental space in these areas is gigantic, and if you can get a machine learning approach to point you to the more productive parts of it that's a big advantage indeed. Unfortunately, the hype is well established here as well, and there is no more egregious current example than this preprint from an MIT student that came out last fall. The manuscript detailed the experiences at a large unnamed R&D company doing material science research, and found that the scientists using AI assistance were significantly more productive and produced demonstrably more innovative materials that in turn led to more patent filings. A lot of commentary ensued, with plenty of it using the paper's findings to cheerlead AI in general: here was proof, real proof, that AI-driven science was the way of the future and that the future was now. We were just all going to have to get used to it. Interestingly, one of the results that got widely noted was that the already-highest-performing scientists were the ones that got the most benefit from adding AI assistance, although they reported lower job satisfaction as they realized just how much of their work it was able to do for them. You can see how this made a splash.

Those first two commentary links (from the Wall Street Journal and The Atlantic) make for interesting reading now, because the MIT student in question features prominently in both of them, speaking fluently and in detail about the work he did, the results laid out in his paper, and their implications. But it now appears that it may well have all been made up. Every last bit. MIT says that it received allegations about the paper that prompted an investigation, and that the two professors who were cited in the preprint now say that "we want to be clear that we have no confidence in the provenance, reliability or validity of the data and in the veracity of the research". Even at the time, some eyebrows had been raised about the numbers in the paper: just what company was this work done at? Why would they allow an MIT undergraduate so much leeway inside their R&D operation? How were these new materials assessed for novelty, anyway? How could you have so many people involved with only one person to coordinate all the data and author the manuscript? But many people wanted to believe, and to have their existing beliefs confirmed so throroughly.

So what's the real state of the art for AI in materials science? It's running a bit behind the preprint's futuristic take, that's for sure. Here's a paper from a group at Ottawa looking at the metal-organic framework (MOF) field, which would seem ripe for ML approaches. You have a huge range of metals to work with, an even larger list of possible scaffolding materials, and depending on the conditions you use (solvent, temperature, pH, additives) you can even get completely different crystal structures out of a single metal/scaffold pair. God knows that there are many people trying to get a handle on the broader questions in the field, but there are so many of those (synthesis, structure, stability, all sorts of possible use cases), and much of the literature still tends to be in the descriptive "Hey, here's some more of the darn things" mode. It's a huge intruiging mess, and I can tell you from personal experience that it's a lot of fun to work in. But it really does need all the help it can get.

The Ottawa group takes a close look at several open-access databases of MOF structures, which people use to try out their machine learning approaches on. And unfortunately they have found that these are in bad shape. There are way too many listed structures that have metal atoms in the wrong oxidation states (and in some cases, actually impossible ones). The paper describes "alarming error rates exceeding 40% in most databases" and you can imagine what you get when you shovel data of that quality into machine-learning algorithms. Yep, models that are state-of-the-art crap. They go on to show that using these for computing new MOFs gives you structures that themselves are wrong at least half the time, which means that the aforementioned crap is being used to produce still more scientific debris. The authors do identify some lower-error-rate databases (none of which seem to be the most popular), and the whole paper should be a wake-up call to the entire community of computational MOF researchers to slow down a bit and clean things up before hitting the ol' Start button again. That might be useful advice in general. . .