Social Sciences › Social Sciences › Safety Research
Academic integrity and plagiarism
50 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Neueste Paper
- hacktrace: behavior-supervised detection of reward hacking during code generation
Hao Jiang, Xin Li, Annan Wang, Yichi Zhang, Weisi Lin · 5. Oktober 2026
A coding agent can earn a passing grade by fixing its code, or by deleting the test that exposes the bug. Detecting such reward hacking requires recognizing attempted shortcuts, including those that fail. We release 173,561 annotated multi-turn coding trajectories from Qwen3-8B and show that supervi…
- Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers
Yerim Oh, Young-Jun Lee, Jaewoo Ahn, Gunhee Kim, Dongyeop Kang · 2. Oktober 2026
AI-generated content, often called AI slop, is increasingly common everywhere, particularly in academia. Slop in AI-generated scientific papers, however, has more complex patterns that cannot be easily detected by existing token-based AI detectors. Each part of such a paper looks plausible while the…
- Correctness, Convergence, and AI-Generated Code Detection: A Longitudinal Study of Student and Large Language Model Code in Introductory Programming
Runlong Ye, Jing Fan, Angela Zavaleta Bernuy, Oscar Karnalim, Paul Denny, Juho Leinonen, Michael Liut · 2. Oktober 2026
Large language models can generate plausible solutions to programming assignments, making it tempting to detect their use by matching student code against a reference bank of generated solutions. Yet similar code can also arise when an assignment admits only a few natural implementations, which leav…
- Policy-Conditioned AI-Use Detection: An Evidentiary Framework for Academic Publishing
Jairo Diaz-Rodriguez, Mumin Jia · 1. Oktober 2026
Major venues now publish detailed rules about how authors, reviewers, and area chairs may use AI, and those rules differ by role, by task, and by what must be disclosed. AI detection, the instrument usually proposed to enforce them, estimates something else: whether an AI model wrote the text. We ar…
- DraftTrace: A Multi-View Analytics Environment for AI-Integrated Writing
Divyansh Chandarana, Sandipan De, Vivek Gupta · 30. September 2026
Generative AI has changed how students produce writing assignments. The final artifact is no longer sufficient to understand the process through which it was produced. We introduce DraftTrace, a writing environment that jointly captures three complementary views of writing: the final product, the wr…
- CheatBench: Measuring Reward Gaming in AI Agents
Long Phan, Stephen K. Yang, Jason J. Lim, Mantas Mazeika, Wenyu Zhang, Zheyuan Liu, Richard Ren, Jingxiang Meng, Yaoteng Tan, Weiliang Zhao, Addison Wu, Matei Anghel, Dan Hendrycks · 30. September 2026
Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended. In recent incidents and controlled evaluations across the AI industry, agents trained to maximize reward have accessed unauthorized information, attempted t…
- Early Prediction of AI-Assisted Cheating Risk in Online Exams Through Learning Analytics
G\"okhan Ak\c{c}ap{\i}nar · 30. September 2026
AI-assisted cheating has become an important threat to the security of online exams. This study examines whether the risk of AI-assisted cheating in the final exam can be predicted using students' digital traces in the learning management system (LMS) during the first eight weeks of the semester. Th…
- Argus: Academic Integrity in the Era of Generative AI
David Racovan, Ajay Rawat, Christopher K. May, Jeffrey A. Turkstra · 30. September 2026
The rapid proliferation of large language models (LLMs) in the context of education has introduced significant challenges in enforcement of academic integrity, especially in programming courses. We present Argus, an automated detection system for LLM-assisted student work in undergraduate C programm…
- "Is This Book AI-Generated?" How Authorship Suspicion Manifests in Marketplace Reviews
Victor Dibia · 29. September 2026
As AI becomes part of how books are authored, reader response to suspected AI authorship grows more consequential, yet remains unexamined. We analyze 863 low-star reviews of 78 Amazon bestsellers across 8 categories at three levels of proximity to AI. Suspicion concentrates in Generative AI books (3…
- Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams
Ali Habibullah, Yazan Alshoibi, Mohammad Alshiekh, Salman Khan, Naeemullah Khan · 25. September 2026
One long-form exam in a large course costs hundreds of grader-hours, and qualified graders are scarce; LLM graders are a tempting alternative. To show its pitfalls we grade a practical Computer Vision exam ($570$ dual-graded students) under $171$ configurations spanning closed and open-weights model…
- Mapping the Authorized Boundary: A Comparative Policy-Vignette Study of Generative AI Governance in Australian Higher Education
Biranchi Poudyal · 25. September 2026
Australian universities regulate students' use of generative artificial intelligence (GenAI) through overlapping policies, procedures, guidance, and assessment instructions, but these environments may classify identical conduct differently. We applied 15 standardized student-use vignettes to the pub…
- CleanScore: Black-Box Benchmark Audits with Negative Controls and Sensitivity Bounds
Jeffery Opoku, David Banahene · 22. September 2026
Public benchmark scores may reflect skill, prior exposure to the questions, or both, and for most models the training data are unknown. We present CleanScore, a black-box audit using scored outputs only. Each benchmark question becomes a parent item with one public form and two independently written…
- Error-Supervised Synthetic Learner Writing for Automated Essay Scoring
Duy Anh Nguyen · 22. September 2026
Synthetic essays can help reduce dependence on human-written data in Automated Essay Scoring (AES). However, they often lack realistic errors, limiting their ability to represent authentic human writing, particularly when the target texts are intended to resemble those produced by language learners.…
- AI-written admissions essays are widespread but penalized
Calvin Isley, Johann D. Gaebler, Sharad Goel · 22. September 2026
AI is rapidly transforming higher education, including the application process, yet relatively little is known about its use and consequences. To help close this gap, we analyze nearly 7{,}500 applications submitted between 2020 and 2025 to a large public policy master's program in the United States…
- Does AI Assistance Leave a Temporal Fingerprint? Detecting Overreliance in AI-Assisted Writing and Programming
Eduardo Davalos, Yike Zhang · 17. September 2026
The rapid adoption of generative AI has made final artifacts unreliable evidence of student learning, and AI detectors that examine only the finished product are inaccurate and ethically contentious. Process data offers an alternative, but prior work covers only English essay writing. We ask whether…
- Vision-Language Models for Criterion-Level Grading of Handwritten Examinations in Outcome-Based Education
Asif Hasan Tonmoy, Saad Ahmed, Md Khalid Syfullah, S. M. Jahangir Alam · 16. September 2026
Criterion-level grading connects examination performance to learning outcomes, but manual marking introduces workload and variation between markers. This study evaluates vision-language models (VLMs) for handwritten outcome-based assessment across five dimensions: accuracy, human agreement, repeated…
- IntraGuard: Committee-Side Defenses Against Review Outsourcing to Commercial Chatbots
Oubo Ma, Ruixiao Lin, Jiahao Chen, Yuan Su, Yong Yang, Shouling Ji · 16. September 2026
LLMs become increasingly capable, editorial boards and program committees are growing concerned about reviewers who fully outsource peer review to commercial chatbots. This concern stems from prior findings that current chatbots lack the independent critical thinking and depth of reasoning required …
- Testing Our Foundations: Citation Trends, Errors, and Emerging Hallucinations in the Computing Education Literature
Paul Denny, Gweneth Barbre, Musa Blake, Yan Cathy Hua, Juho Leinonen, Andrew Luxton-Reilly, James Prather, Brent N. Reeves · 16. September 2026
Accurate references are foundational to scholarly work, enabling verification, attribution, and systematic review. However, the rapid adoption of large language models has introduced a serious integrity concern: plausible-looking but fabricated citations. Although hallucinated references are widely …
- When Auditors Fabricate: Batch-Size Degradation and Confident Hallucination in LLM Detection of Planted Document Contamination
Karan Parekh, Sanjana Pendyala Ravinder, Sana Mhapsekar, Medina Maloku · 10. September 2026
Large language models are increasingly proposed as automated auditors of document quality, yet their reliability as detectors of planted errors is poorly characterised. We construct a contaminated corpus of 150 academic papers spanning supply chain management and medical research, injecting 450 know…
- GradeTrap: Authority Cues in Images Shift VLM Judgments Despite Explicit Instructions to Ignore Them
Deep Dessai (The University of Texas at Austin) · 9. September 2026
As vision-language models (VLMs) become increasingly capable and are deployed in consequential real-world settings, they must evaluate evidence independently rather than defer uncritically to human authority. We introduce GradeTrap, a controlled evaluation that places two social cues in direct confl…
- A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment
Mar\'ia Eugenia Curi, Germ\'an Capdehourat, Isabel Amigo, Magdalena Romano, Rosana Serra, Adri\'an Silveira, Andr\'es Peri · 7. September 2026
The integration of artificial intelligence (AI), particularly large language models (LLMs), into educational assessment has opened new opportunities to enhance the efficiency and scalability of grading processes. This study presents the design and validation of an AI-assisted scoring framework for w…
- Who Should Grade My Work? Student Perspectives on Transparent AI-Assisted Writing Assessment in Higher Education
Rayed AlGhamdi · 7. September 2026
The integration of GenAI tools into higher education assessment raises important questions about how students understand, interpret, and respond to AI-mediated evaluation. As instructors increasingly explore AI tools for providing feedback, prior research has examined whether GenAI-generated feedbac…
- The AI Adaptation Gap in Higher Education: Students, Faculty, and Administrative Staff
Yuriy S. Braun, Salavat M. Khafizov · 27. August 2026
The purpose of this study was to analyze patterns of artificial intelligence (AI) use and attitudes toward AI among students, faculty, and administrative staff at a large university specializing in teacher education. The analytical sample comprised 1809 students, 250 faculty members, and 62 administ…
- Multimodal examination answer data with expert-designed Outcome-Based Education rubrics for criterion-level assessment
Jahangir Alam SM, Md Khalid Syfullah, Saad Ahmed, Munira Akter Mou, A K Z Rasel Rahman, A. K. M. Masudur Rahman, Mohammed Sowket Ali · 25. August 2026
This data article describes a multimodal collection of scanned examination answers paired with expert-designed Outcome-Based Education (OBE) grading metadata. The collection contains 485 answer submissions from 415 consenting students at four academic institutions. Eight faculty contributors supplie…
- Expectations and Practices around AI Disclosure in CS Research
Arati Mohapatra, Danish Pruthi · 25. August 2026
As generative AI tools find increasing use in research workflows, ongoing debates on their impact, appropriateness and responsible use have led policymakers to enact policies to disclose AI use at multiple publishing venues. However, are current AI disclosure policies and practices reflective of the…
