This is really honestly an alarmingly wrong substack post.
1. The claim of the authors is that the scientific literature is on average so dishonest and/or wrong that access to it (pre or post training, I assume?) actually harms LLM accuracy. Hopefully we can all agree this is a claim so extraordinary that would require some very uniquely powerful evidence. After all, we know that LLMs also train on very large corpuses of text with varying degrees of both accuracy and dishonesty - not the least, most of the internet! So the claim of the authors must be that the scientific literature is so bad that it is actually uniquely bad for LLMs, like worse than the general internet.
2. They provide one citation of evidence supporting their claim, which is this paper [0]. If you actually read this paper (which is mostly not about this actual question, but related questions about prefiltering), it provides no evidence for their claims. For example, in Figure 5, removing PubMed, aka the entire biological scientific literature, has by far the strongest negative impact on evaluation of Biomedical questions. It even has a strong negative impact on evaluation of questions in the "Common sense" category! To quote:
"Performance degrades when we remove
domains with close alignment between the pre-
training and downstream data sources: removing
PubMed hurts the BioMed QA evaluations"
3. This post is maybe what you would get if you prompted an LLM: "please provide a citation for the claim that the scientific literature harms LLMs". It's really quite worrying that we're looking at AI generated propaganda designed to convince the reader that the scientific literature is worse than useless.
4. As a scientist, the primary use I want from LLMs outside of code generation is to be a fast and comprehensive search engine. I want to see the papers behind their claims and evaluate them. Literally the only time they are useful in scientific research is when they provide citations for the claims, so I can read the papers and evaluate.
5. Somehow we have gotten to the point where many people, especially people in tech, believe that the scientific literature is mostly junk or mostly useless. This contrasts so starkly with the current rapid pace of genuine, meaningful, society-impacting scientific progress in pretty much every major field. How we have gotten to the point where this myth is so pervasive, I do not know. Is it really spurred just by a few recurring news stories about reproducibility? By just interpersonal bitterness and feelings of anti-'elite' sentiment? (how on Earth your local state university climate scientist is a member of the 'elite' but software engineers making $500k a year or X influencers with millions of followers are not will always simply be beyond me).
[0]. https://aclanthology.org/2024.naacl-long.179/