Part the First: Denial Ain’t Just a River in Egypt. This summer has been one of the hottest since temperature records have been kept. Most of the hottest years ever recorded have been in the last twenty years. The mechanism that caused anthropogenic global warming (AGW) was mentioned by the polymath Charles Babbage in 1835 and explained by Svante Arrhenius (awarded the third Nobel Prize in Chemistry) in a paper in 1896. The founder of physical chemistry (shudder) was off in his prediction because he did not see oil and natural gas coming and extrapolated from the current use of coal.
In the New York Review of Books, Judge Jed S. Rakoff explains how “a new edition of a manual that helps federal judges understand scientific subjects included a chapter on climate change that was withdrawn after attacks by denialists.” And so it goes, as the spawn of the Powell Memo continue to be the gift that keeps on giving:
To help acquaint federal judges with the basics of such subjects, the research and educational arm of the federal judiciary, the Federal Judicial Center (FJC), began publishing the Reference Manual on Scientific Evidence in 1994. It includes chapters on general topics such as the scientific method and others on more specific subjects that a judge may have to confront in a given case, such as epidemiology, toxicology, neuroscience, or computer science. Supreme Court Justice Elena Kagan, in her foreword to the most recent edition, writes, “The manual is the product of close collaboration among highly respected scientists, engineers, judges, and lawyers. It delves into the scientific subjects that judges most often face. It explains scientific approaches and explores scientific uncertainties and limits.”
…
In preparing new editions of the manual, the FJC has partnered since 2010 with the National Academies of Sciences, Engineering, and Medicine, a private institution established by Congress in 1863 to advise the government on scientific issues. That collaboration resulted in the widely praised Third Edition of the manual, published in 2011, and the FJC asked the National Academies to prepare a Fourth Edition, which was published at the end of 2025.
The chapter on AGW has been deleted; the mechanism for this is explained in the article. We have seen this before in our history:
In 1633 the Roman Inquisition convicted Galileo Galilei on suspicion of heresy for daring to assert that the earth was not the center of the universe but revolved around the sun. He was required to abjure his opinion but afterward was rumored to have quietly muttered, “And yet it moves.” We may try to deny that human activity is a major cause of the rapid warming of our planet, which has already begun to show signs of looming catastrophe. And yet it warms
For anyone who wants to read more about Galileo, Stillman Drake is the go-to source. I found a copy of his version of Dialogue Concerning the Two Chief World Systems on a remainder table as a college freshman (alas, lost in a move long ago) and almost changed my major to history. There is some small comfort in the knowledge that some things never change. The problem, though, is that eventually we will reach the Dead End of No Return. And so it goes.
Part the Second: Penalties for Scientific Malpractice? The collapse of authority surrounds us, and it pains me, although I am not surprised (Bayh-Dole Bill of 1980), that science is not much different from politics and economics in this regard (although when a scientist is found to have just “made stuff up,” that is one career ended beyond redemption, even if it sometimes takes too long in practice). It seems that countries with emerging scientific research establishments are making rules at the beginning, according to Nations expanding their research programmes impose sweeping penalties for malpractice (India, Peru, Vietnam):
In a bid to tighten research standards, governments and funders in countries with emerging research systems are increasingly rolling out disciplinary measures for researchers who do the wrong thing. The policies include docking bonuses, restricting access to grants and lowering the rankings of institutions linked to misconduct.
The reforms come as instances of research misconduct are on the rise worldwide. In May, for example, an audit of biomedical papers found that the rate of fabricated citations was 12 times greater in 2025 than in 2023. And later that month, an analysis of 33,000 cancer studies linked roughly 12% of them to suspected paper-mill activity.
The responsibility for cleaning up research malpractice has, for a long time, been a bottom-up affair: journal editors withdraw papers; universities investigate their staff members; and independent sleuths and organizations, such as Retraction Watch, expose what they can. Sometimes researchers will come forwards and correct the record themselves once they realize that they’ve made an error.
But there is a growing movement, particularly in countries with emerging research communities, to take a top-down approach. Funders and governments in some nations are increasingly restricting grants and opportunities for scientists who have been accused of research-integrity malpractice. These changes seek to cover a wide range of integrity issues, from plagiarism and falsified data to image manipulation and peer-review fraud.
We shall see. One can only hope that this actually takes. But it is also true that many of scientific publishers in the online, open-access, pay-to-publish-anything world do business from India. It would seem that horse left the barn some time ago. One notable exception is China, which by most measures has passed the Anglophone and European scientific establishments in research:
China has ramped up measures to combat research malpractice, such as by conducting a nationwide review of misconduct in 2024 that required universities to report papers that had been retracted in the previous three years and investigate suspected misconduct. China has also established a national database of scientific misconduct, and its science ministry has said that it will punish institutions that fail to investigate or sanction serious cases.
In the meantime, American scientists with the option and opportunity are finding other places to do their essential work. As I have mentioned before, one of the most prominent of these is nutrition researcher Kevin Hall, who retired early from NIH after RFKJr objected to his results because they did not comport with his preconceived notions, will be joining the University of Ottawa next year. Dr. Hall’s fundamental research on ultraprocessed foods was discussed here in 2024.
Part the Third: Fake Citations in the Scientific Literature. Nothing new here, but with AI in the offing, this will only get worse. From earlier this year in Nature we have, Surge in fake citations uncovered by audit of 2.5 biomedical science publications. The absolute numbers are small but the increase is large. That is cold comfort in the coming world in which AI will make this more common:
An audit of 2.5 million academic papers has identified nearly 3,000 biomedical-science papers that contain fake references — ones that could not be traced to known publications.
The findings, published in The Lancet on 7 May, are contained in the first academic study to estimate the scale of fake citations in the biomedical literature.
The authors designed an automated pipeline to screen papers from PubMed Central — a database of publicly accessible biomedical articles — published between January 2023 and February 2026.
The findings are “conservative underestimates”, says study co-author Maxim Topaz, an AI researcher at Columbia University in New York. “What we identified is the lower bound of true prevalence. We’re scratching the tip of the iceberg,” he adds.
Lately I have discussed this with several colleagues who are transitioning into medical education from the practice of medicine. They are, to a person, very enthusiastic about the promise of AI in parsing the biomedical literature. As we discussed recently, what these physicians lack is appreciation of the “tacit knowledge” that is essential for understanding the scientific literature. To the novice, bullshit looks perfectly fine. And I have not yet figured out how to get them to understand that the training set will have included papers with fictitious references and how this should dampen their enthusiasm for large language models that underly algorithmic intelligence
‘Tis a mess that gets messier…
Part the Fourth: Hallucinatory Citations. Can researchers stop AI making up citations? This paper is more than a year old, i.e., ancient in context, but the problem remains, based on my limited use of AI in using the scientific literature (that allergy is still strong in this one):
Artificial intelligence (AI) models are known to confidently conjure up fake citations. When the company OpenAI released GPT-5, a suite of large language models (LLMs), last month, it said it had reduced the frequency of fake citations and other kinds of ‘hallucination’, as well as ‘deceptions’, whereby an AI claims to have performed a task it hasn’t.
With GPT-5, OpenAI, based in San Francisco, California, is bucking an industry-wide trend, because newer AI models designed to mimic human reasoning tend to generate more hallucinations than do their predecessors. On a benchmark that tests a model’s ability to produce citation-based responses, GPT-5 beat its predecessors. But hallucinations remain inevitable, because of how LLMs function.
“For most cases of hallucination, the rate has dropped to a level” that seems to be “acceptable to users”, says Tianyang Xu, an AI researcher at Purdue University in West Lafayette, Indiana. But in particularly technical fields, such as law and mathematics, GPT-5 is still likely to struggle, she says. And despite the improvements in hallucination rate, users quickly found that the model errs in basic tasks, such as creating an illustrated timeline of US presidents (funny).
I suppose we have moved beyond these models, but I have not yet seen convincing evidence that hyperscaling will solve the problems. The crap is still in the residue scraped without consideration of its quality. From my point of view, the problem is how to get scientists to pay attention and not be seduced by shortcuts that can lead them into trouble. The overuse of AI in medical education is not encouraging, though. When I have used AI to summarize a subfield that I know a fair amount about, the results are C-plus work at best, but very convincing, as in a simulacrum of authoritative self-assurance.
It may not be an accident that this is congruent with current pedagogical thinking that os overconcerned with something called “the minimally competent student.” When I ask if these brain geniuses want their children’s doctors to be minimally competent physicians, there is a glimmer of recognition there may be a problem in there, somewhere. I respond that our goal should be to prepare the maximally competent student who will become the maximally competent physician. Or so I thought. In any case, I have never concerned myself much with the student who was happy to remain minimally competent. They usually sort themselves into the backwaters of big cities where comparative anonymity is not a problem for them.
Part the Fifth: Low-Quality Biomedical Research Papers. We will wrap up our coffee break today with another paper from 2025 that I intended to discuss but never got back to, AI linked to explosion of low-quality biomedical research papers. When I read this, Alex Trebek came to mind with a Jeopardy contestant saying, “I’ll take ‘No sh*t, Sherlock,’ for $2000, Alex!”
The scientific literature is at risk of becoming flooded with papers that make misleading health claims based on openly available data that are easy to process using artificial intelligence (AI) tools, researchers have warned.
In a study published in PLoS Biology on 8 May, scientists analysed more than 300 papers that used data from the US National Health and Nutrition Examination Survey (NHANES), an open data set of health records. The papers all seemed to follow a similar template, associating one variable — for example, vitamin D levels or sleep quality — with a complex disorder such as depression or heart disease, ignoring the fact that these conditions have many contributing factors.
We have a sudden explosion in publication rates [of papers] that are extremely formulaic that could easily have been generated by large language models,” says study co-author Matt Spick, a biomedical scientist at the University of Surrey in Guildford, UK.
Spick and his colleagues found that the associations in many of the papers did not hold up to statistical scrutiny, and that some studies seemed to have cherry-picked data.
“Imagine you’re trying to pass an exam that has a particular pass rate, and you add as many questions as you want. You see which ones you got right, and you remove the ones that you got wrong. That’s basically what they’re doing,” explains Charlie Harrison, a computational biologist at Aberystwyth University, UK, who also worked on the study.
Publishers and editors are trying to get a handle on this, but the AI detectors are not particularly sensitive. Paper mills are in the crosshairs, but their sludge has been used to train the LLMs, and how this can be removed is difficult to grok. Without sounding too much like a broken record, this has begun to have a malign effect on medical education. To make a long story short, the first medical board exam that is taken after the second year of medical school and is required for the vast majority of students to enter their clinical years in the teaching hospital went from scored with percentile ranks to Pass-Fail several years ago.
No one with half a brain thought this was a good idea, for myriad sound reasons. The most important is that this shift removed the one “objective” criterion for medical student achievement and mastery of the foundations of medicine. This removed the opportunity for students at medical schools other than Harvard-Yale-Stanford-Johns Hopkins to distinguish themselves with a high basic medical science score (H-Y-S-JHU students go where they want to, where some are inevitably found out).
The natural alternative to this change was “research.” Now there are medical students applying for their preferred residency slots who have 10, 20, or 30 research papers in their CVs. No, not so much. Most of these “papers” are statistical sleights of hand, probably written with or by AI and published in journals that would not be fit to line a bird cage, if they existed on paper. It seems residency program directors are catching on, but the urge to count instead of to evaluate is very strong with program directors and their institutions.
Once again we have come to the immovable object in the form of “tacit knowledge” that can only be gained on the long path. The sad thing is that the few medical students I know who took this path, which can no longer happen at our institution, think they are now “researchers.” No, not so much.
Thank you for reading! See you next week. We should remember this is the date 25 years ago that may have finally knocked the world off balance for good. The deeply felt aches of that tragic atrocity remain but we can still rise above “all war all the time” that was “justified” by 9/11. Actually, we have no choice.


