#News
Can AI help overstretched peer reviewers?
Researchers advocate using AI to screen manuscripts, detect fraud, and support editors, while warning of risks to process integrity
Experts argue that AI should serve as a pre-review filter, dealing with technical, verifiable tasks to ease the burden on human peer reviewers | Image: Unsplash
The growing volume of manuscript submissions is placing increasing strain on peer reviewers, who often volunteer their time and are becoming progressively harder to recruit.
Publisher Elsevier alone received approximately 4.2 million manuscript submissions in 2025, of which 795,000 were published, while the remainder were rejected during the editorial process—many before ever reaching peer review.
The 2026 Future of Peer Review report found that editors now need to send out an average of five invitations before a single researcher agrees to review a manuscript.
The widespread adoption of generative AI by both authors and reviewers has sparked debate over whether—and how—the scholarly publishing system should incorporate the technology.
In an interview with Science Arena, physician and researcher Howard Bauchner, former editor in chief of the Journal of the American Medical Association (JAMA), advocates using AI as a pre-review screening tool.
“AI would evaluate the manuscript, and the editor would review that assessment before deciding how to proceed. They could reject the manuscript because the AI identified fundamental flaws, or return it to the authors for revision before sending it out for external review,” he suggests.
The primary goal of AI is not simply to speed up the publication process, but to reduce the number of manuscripts that ultimately reach human reviewers, he says.
How AI-based editorial screening would work
Under this model, AI would be responsible for technical, verifiable, and repetitive tasks, including:
- Compliance with reporting guidelines;
- Study registration;
- Reference accuracy;
- Detection of undeclared AI-generated text;
- Image manipulation.
The last of these—image manipulation—is particularly important in laboratory sciences and is rarely detected through conventional peer review.
Bauchner also points to what he calls ghost reviewers, who use AI even when journal policies prohibit it. Rather than allowing the practice to remain hidden and unregulated, he argues that AI use should be permitted, provided it is fully disclosed.
What AI can already do in editorial screening, according to Howard Bauchner
1. Compliance with reporting guidelines: Verifies that the manuscript meets the journal’s required reporting checklists.
2. Study registration: Confirms that clinical trials and other eligible study designs were properly preregistered.
3. Reference accuracy: Detects incorrect or nonexistent citations, including those fabricated by generative AI.
4. Detection of undeclared AI-generated text: Flags passages that appear to have been written by language models without appropriate disclosure.
5. Image manipulation: Identifies inappropriate alterations to scientific figures—a frequent oversight in human-only peer review.
The limits of artificial intelligence in manuscript evaluation
For now, AI cannot replace human reviewers in every aspect of manuscript assessment. Bauchner explains that AI is still unable to evaluate the broader scientific context of a study—how a manuscript fits within the ongoing research.
“And if you ask most reviewers, that’s exactly what they enjoy commenting on,” says the former JAMA editor in chief.
Jesús Mena-Chalco, a professor at the Federal University of ABC (UFABC) and a researcher in scientometrics—the field that examines the quantitative dimensions of science—identifies a similar limitation. “Determining whether a study is truly original and relevant is a key role of the human reviewer.”
This limitation is also reflected in the lack of critical judgment displayed by AI models. “Current models tend to be affirmative—they’re rarely critical,” Mena-Chalco explains.
The advantages AI offers in the editorial process are not without risks. Perhaps the most pressing is data confidentiality. Manuscripts contain unpublished findings that should not be absorbed into the language models used by reviewers—or by the authors themselves.
“If someone uses a language model, safeguards must be in place to ensure that those data do not become part of the model’s corpus,” Howard Bauchner warns.
Even so, AI may help address a problem that it has itself helped to create: fabricated references generated by language models that invent inexistent citations.
Automatic reference verification—already envisioned as part of the proposed pre-review screening process—has therefore become increasingly important.
Bias: From training data to human reviewers
Like human reviewers, large language models (LLMs) reflect biases. In the case of AI, those biases stem from the datasets on which the models were trained.
Bauchner argues that AI systems could be instructed to disregard authors’ identities, helping eliminate biases associated with institutional affiliation, country of origin, or language.
Mena-Chalco adds that LLMs have one practical advantage over human reviewers: they do not become fatigued.
“A human reviewer who has to evaluate multiple manuscripts may ultimately be more biased than a computer that never gets tired.”
Both arguments, however, describe intended uses rather than guarantees. Biases embedded in training data persist regardless of the instructions the models receive.
Both researchers believe that openly integrating AI into the scholarly publishing ecosystem is only a matter of time. Mena-Chalco notes that many researchers already rely on these tools but remain reluctant to admit doing so.
“Many people feel guilty about using these systems because they worry their work will seem too easy. So where should the line be drawn?” he asks.
For Mena-Chalco, the answer will come through normalization.
“In about 10 years, it will no longer be necessary to disclose AI use,” he predicts.
“I believe that by 2030, AI will have become an integral part of the scientific communication ecosystem,” Bauchner argues.
There are already signs that this transition is underway. The Public Library of Science (PLOS), one of the world’s largest nonprofit open-access scientific publishers, has implemented automated research integrity screening.
Data presented by the organization show that desk rejections increased from 13% in 2021 to 40% in 2025.
A study conducted by the editorial team of the journal Organization Science provides the first full-corpus empirical analysis of this phenomenon within a scientific journal.
Between 2021 and 2026, submissions to the journal increased by 42% following the launch of ChatGPT in November 2022. The rise was driven primarily by manuscripts produced with extensive AI assistance rather than by organic growth in the field.
Manuscripts containing more than 30% AI-generated content—as measured by the Pangram detection tool—were up to 30 percentage points more likely to be rejected during initial editorial screening. Over the same period, writing quality, measured using the Flesch Reading Ease score, declined by 1.28 standard deviations.
Review reports followed a similar pattern. More than 30% now show some degree of AI assistance, with reports that are more difficult to read and narrower in scope, placing greater emphasis on theory and less on data.
For researchers in countries where English is not the primary working language, however, this shift could produce tangible benefits well before that timeline.
“AI can be extremely helpful for translation. If we use these models with a well-crafted prompt, the result is a highly professional, context-aware translation,” Mena-Chalco says.
What remains uncertain are the editorial policies that will govern this transition—and how they will affect the integrity of the peer review process that underpins scientific research.
*
This article may be republished online under the CC-BY-NC-ND Creative Commons license.
The text must not be edited and the author(s) and source (Science Arena) must be credited.