Estimating the predictability of questionable open-access journals

Han Zhuang; Lizhen Liang; Daniel E. Acuna

Science Advances, 11(35), eadt2792 · 2025

Abstract

Questionable journals threaten global research integrity, yet manual vetting can be slow and inflexible. Here, we explore the potential of artificial intelligence (AI) to systematically identify such venues by analyzing website design, content, and publication metadata. Evaluated against extensive human-annotated datasets, our method achieves practical accuracy and uncovers previously overlooked indicators of journal legitimacy. By adjusting the decision threshold, our method can prioritize either comprehensive screening or precise, low-noise identification. At a balanced threshold, we flag over 1000 suspect journals, which collectively publish hundreds of thousands of articles, receive millions of citations, acknowledge funding from major agencies, and attract authors from developing countries. Error analysis reveals challenges involving discontinued titles, book series misclassified as journals, and small society outlets with limited online presence, which are issues addressable with improved data quality. Our findings demonstrate AI’s potential for scalable integrity checks, while also highlighting the need to pair automated triage with expert review.

From the original work, under its Creative Commons license.

Research question

Can journal websites and publication metadata help identify journals that warrant closer investigation?

Finding

At a 50% decision threshold, the classifier flagged 1,437 of 15,191 journals. Manual review estimated that 24% of flagged journals were false positives, supporting automated screening as a guide to expert investigation.

A model flag is not a determination of misconduct. Errors included discontinued titles, misclassified book series, and small society journals with limited web presence.

Corrected original Figure 3A showing annual publication output in journals flagged as questionable, rising through the 2010s with lower output in 2019 and 2020.
Annual output of journals flagged at the 50% threshold; error bars reflect uncertainty in their classification. Zhuang, Liang, and Acuna (2025), corrected Figure 3A, cropped. Source · CC BY-NC 4.0 · Published correction.

Media coverage

Cite this work

Han Zhuang; Lizhen Liang; Daniel E. Acuna (2025). Estimating the predictability of questionable open-access journals. Science Advances, 11(35), eadt2792. 10.1126/sciadv.adt2792.

@article{zhuang2025estimating,
  title = {Estimating the predictability of questionable open-access journals},
  author = {Zhuang, Han and Liang, Lizhen and Acuna, Daniel E.},
  year = {2025},
  publication_date = {2025-08-27},
  journal = {Science Advances},
  volume = {11},
  number = {35},
  pages = {eadt2792},
  doi = {10.1126/sciadv.adt2792},
  url = {https://doi.org/10.1126/sciadv.adt2792}
}
Download BibTeX