SciIntBench: Measuring LLM Compliance with Research Integrity Norms Under Adversarial Framing

Almene De Meran Meguimtsop; Maria Leonor Pacheco; Daniel E. Acuna

arXiv preprint arXiv:2605.29468 · 2026 · Preprint / working paper

Abstract

Large language models (LLMs) are increasingly used to support scientific work, but it is unclear whether they uphold responsible conduct of research (RCR) norms or help undermine them. We introduce SciIntBench, an adversarial benchmark of 810 prompts across ten RCR categories and three scientific domains. Each scenario appears as an Overt Adversarial, Covert Adversarial, and Benign version, allowing us to jointly measure framing-sensitive refusal of misconduct and helpfulness on legitimate requests. We evaluate 16 commercial and open-weight LLMs from six providers (2024--2026), producing 12,960 responses. We find that scientific integrity alignment is strongly framing-sensitive: models refuse explicit misconduct far more reliably than covert violations, especially failing when misconduct is presented as a pressure-driven shortcut. Refusals vary by RCR category, with weaker boundaries around transparency, plagiarism, and fabrication.

From the original work, under its Creative Commons license.

Research question

Do language models uphold research-integrity norms when misconduct is framed as a routine request?

Finding

Across 16 models, refusals were substantially more reliable for explicit misconduct than for covert violations. SciIntBench evaluates both misconduct refusal and helpfulness on legitimate scientific tasks.

These results concern the benchmark's scenarios and model versions; they do not establish safety across every scientific use.

Original SciIntBench chart comparing overt and covert adversarial refusal rates for 16 language models; covert requests generally receive fewer refusals.
Refusal rates for explicit and covert misconduct across 16 models. Meguimtsop, Pacheco, and Acuna (2026), Figure 2A, cropped. Source · CC BY 4.0.

Cite this work

Almene De Meran Meguimtsop; Maria Leonor Pacheco; Daniel E. Acuna (2026). SciIntBench: Measuring LLM Compliance with Research Integrity Norms Under Adversarial Framing. arXiv preprint arXiv:2605.29468. 10.48550/arXiv.2605.29468.

@article{meguimtsop2026sciintbench,
  title = {SciIntBench: Measuring LLM Compliance with Research Integrity Norms Under Adversarial Framing},
  author = {Meguimtsop, Almene De Meran and Pacheco, Maria Leonor and Acuna, Daniel E.},
  year = {2026},
  publication_date = {2026-05-28},
  journal = {arXiv preprint arXiv:2605.29468},
  doi = {10.48550/arXiv.2605.29468},
  url = {https://arxiv.org/abs/2605.29468}
}
Download BibTeX