Social Sciences › Decision Sciences › Management Science and Operations Research
Psychometric Methodologies and Testing
75 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- United States42% · 13 papers
- Australia23% · 7 papers
- China13% · 4 papers
- United Kingdom6.5% · 2 papers
- South Korea6.5% · 2 papers
- Italy3.2% · 1 papers
- France3.2% · 1 papers
- Czechia3.2% · 1 papers
Across 31 papers on this subject with at least one lab located. 17 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- Toward Embedding-Based Psychometrics: Structural Modeling of Assessment-Item Semantics With Contextual Scores
Jinsong Chen, Shi-Ting Chen · 30 September 2026
Contextual scores represent assessment items through their similarities to reference words in an external corpus. We examine the semantic structure of scores for 40 TIMSS mathematics scored units using a partially specified two-step factor procedure. A search across factor counts identifies a persis…
- Three Ways Classical Test Theory Misleads for LLM Judges
Louis Yiven Zhu · 25 September 2026
An LLM judge scores a bank of responses against a rubric, and the reliability comes back at $0.52$. What has been measured? Judge evaluation has begun borrowing reliability statistics from classical test theory, usually without stating the measurement design each statistic assumes, and we show that …
- Can Large Language Model-Generated Responses Support Assessment Development? A Human-Calibrated Rasch Benchmark
Eunjeong Song, Sehee Hong · 23 September 2026
Large language models (LLMs) are proposed as synthetic respondents for pilot testing, but their usefulness depends on whether they supply the evidence assessment development requires. We calibrated rating scale models on 14 digital-use skill items from 6,245 adults and used the human item parameters…
- Turkish MMLU Pro: Traceable Option Augmentation and Its Validity Limits in Turkish Multiple-Choice Evaluation
M. Ali Bayram · 16 September 2026
Adding answer options can lower multiple-choice scores without improving assessment validity. Turkish MMLU Pro examines this distinction using 12,000 Turkish-source questions across 58 sections. Each question retains its stem, five original options and source key, and receives five options copied fr…
- Dynamical Non-compensatory Multidimensional IRT Model Using Variational Approximation
Hiroshi Tamano, Daichi Mochihashi · 10 September 2026
Multidimensional item response theory (MIRT) is a statistical test theory that precisely estimates multiple latent skills of learners from the responses in a test. Both compensatory and non-compensatory models have been proposed for MIRT: the former assumes that each skill can complement other skill…
- What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
Meera Desai, Sang T. Truong, Hanna Wallach, Alex Chouldechova, A. Feder Cooper, Jean Garcia-Gathright, Daniel E. Ho, Abigail Z. Jacobs, Sanmi Koyejo, Nicholas Pangakis, Angelina Wang · 9 September 2026
Benchmarks play a central role in the development and governance of models, yet it is often unclear whether they actually measure the concepts they purport to measure (e.g., reasoning, refusal). We adapt convergent and discriminant validity from the social sciences into an approach for interrogating…
- Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?
Qiaoyuan Zheng, Yiqu Yang · 2 September 2026
Small leaderboard gaps are often interpreted as evidence that one language model is better than another, but their sign may depend on which benchmark items are included. We test this using item-level responses from five benchmarks and a family-label-free spectral approximation to multidimensional it…
- Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation
Kefeng Duan, Dewu Zheng, Yanlin Wang, Xiwen Wang, Ensheng Shi, Xilin Liu, Yuchi Ma, Jiachi Chen, Mingwei Liu, Zibin Zheng · 2 September 2026
Evaluating software engineering agents on realistic benchmarks is costly, since each task may require multi-step code exploration, modification, and test execution. Existing efficient evaluation methods select representative subsets to estimate full-benchmark performance, but are largely result-only…
- Check The Scoreboard: An Analysis of Scoring Schemes on Multiple-Choice Evaluation
Nishant Balepur, Paiheng Xu, Wei Ai, Eunsol Choi, Rachel Rudinger, Jordan Boyd-Graber · 1 September 2026
Multiple-choice question answering (MCQA) benchmarks in NLP use number-right scoring (accuracy), but in educational testing, the scoring scheme, the combination of the response mode models follow and the rule for grading responses, is a key design choice that dictates which abilities to reward. We e…
- LLM Judges as Raters: A Pre-Registered Audit of Severity, Halo, Reliability, and Version Instability in LLM Essay Scoring on Public Corpora
Veerendra Kumar Sunkavalli · 1 September 2026
Large language models (LLMs) are increasingly used as essay graders in learning analytics, evaluated almost exclusively with agreement statistics. Educational measurement warns that raters also differ in severity, show halo, and drift as instruments. We treat LLM judges as raters and run a pre-regis…
- One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation
Louis Yiven Zhu · 1 September 2026
Frontier-model leaderboards now rank systems based on economic benchmarks, tests of how well models carry out professional tasks from software engineering to banking workflows, and those rankings inform what organisations buy, what regulators scrutinise, and expectations of how work will change. Whe…
- Rating the Raters: Rasch Measurement Theory for LLM Evaluation
Pratik S. Sachdeva, Nathan Boudol · 31 August 2026
LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models' outputs, and raters of human-generated content. Each paradigm can be viewed as a measurement problem, where a latent property of an object is probed with items from an instrument (e.g., benchmark) by…
- What Reaches Expert Review? Representation, Structural Screening, and Candidate-Form Dependence in AI-Assisted Item Development
Christopher Brooks (School of Information, University of Michigan) · 26 August 2026
Between AI-assisted item generation and expert review sits a computational evaluator whose decisions are usually treated as technical preliminaries. Yet representation, structural reduction, and selection policy determine which items and evidence psychometricians ever receive. Across two linked in-s…
- A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments
Jing Huang, Jihong Zhang, Hua-Hua Chang · 26 August 2026
The rapid expansion of large-scale assessments and the growing adoption of automatic item generation have intensified concerns about incidental content redundancy, where construct-irrelevant elements such as wording or contextual framing become unintentionally repetitive across items. Traditional si…
- LLM Evaluation on Unseen Questions: Contextual Multidimensional IRT Model
Ergan Shang, Weijing Tang, Yinqiu He · 25 August 2026
Evaluation of large language models (LLMs) increasingly requires predicting how a model will perform on new questions or tasks before collecting large amounts of new annotations. This problem is challenging because question difficulty, scenario, and underlying capability demands can vary substantial…
- Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses
Alona Strugatski, Licol Zeinfeld, Jason Cooper, Shelley Rap, Gil Schwarts, Giora Alexandron · 19 August 2026
The evaluation of large language models (LLMs) relies heavily on human-designed assessments, implicitly assuming that AI and humans employ similar underlying cognitive constructs. Challenging this assumption, we investigate whether the latent factors governing LLM performance carry the same substant…
- What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations
Xiaonan Xu, Wenjing Wu · 19 August 2026
Context: Software systems that depend on commercial large language model APIs must migrate to successor versions when vendors deprecate older models. Migration decisions typically rely on aggregate benchmark scores, which compress heterogeneous item-level behaviour into a single net figure. Objectiv…
- When Can Text Embeddings Replace Item Calibration? A Geometric Diagnostic for Semantic Loadings in Multidimensional Adaptive Testing
Amirreza Mehrabi · 11 August 2026
Multidimensional item response theory relies on calibrated item parameters, such as discrimination and category threshold values, which are usually estimated from large samples of human test responses. This study investigates whether the directional loadings of these parameters can be recovered dire…
- Automated item evaluation: Predicting item acceptance and rejection using LLM-generated critiques
Hotaka Maeda, Yikai Lu · 10 August 2026
Automated item evaluation (AIE) refers to the use of computational methods to assess item quality without requiring manual expert review or field testing of the items under evaluation. We aimed to build a near-comprehensive AIE model by predicting item acceptance and rejection from item text using h…
- Item Response Theory for AI Safety
Joshua Fonseca Rivera (Independent), Neil Shah (Independent), David Demitri Africa (UK AI Security Institute), Konstantinos Voudouris (UK AI Security Institute) · 6 August 2026
Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to trust and interpret, because benchmarks duplicate one another, correlate heavily, and models may sandbag when they detect evaluation. To address these…
- Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks
Xiao Fei, Yang Zhang, Sarah Almeida Carneiro, Michalis Vazirgiannis · 5 August 2026
Most multiple-choice question (MCQ) benchmarks evaluate Large Language Models (LLMs) only by whether they select the correct answers. This binary scoring treats all incorrect responses alike, even though an LLM's preferences among incorrect options may contain systematic and useful information about…
- Can LLMs Really Understand Item Difficulty Levels? Implications for Automated Item Generation Using LLMs
Xinyi Wang, Hong Jiao, Ming Li, Sydney Peters, Hanna Choi, Tianyi Zhou, Qingshu Xu · 3 August 2026
The estimation of item difficulty plays a key role in both formative assessment and large-scale high-stakes summative assessments. This study explores how large language models (LLMs) perform in predicting item difficulty levels using items from a large-scale Reading and Writing test. The study inve…
- CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation
Mengting Chen, Yanshu Sun, Wanting Liang, Beidi Luan, Rui Sun, Dezhi Chen, Jing Li, Zuo Bai · 3 August 2026
Reliable evaluation of open-ended LLM outputs requires fine-grained rubrics, yet expert curation is costly and difficult to scale. Existing automated pipelines rely on strict judge unanimity and binary variance filters, which cannot distinguish measurable rubrics from informative ones. We introduce …
- Dimensionality and Measurement Precision in HLE's Multiple-Choice Subset
Mayank Sharma, Savira Nadela, Tyler Matteson · 31 July 2026
Humanity's Last Exam (HLE) is widely used to evaluate frontier language models. HLE organizes its questions into eight subject-domain categories, whose subscores are often interpreted as evidence of distinct capabilities. However, no study has assessed whether these labels correspond to empirically …
- Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach
Wenjie Zhou, Yunting Liu, Renjiao Tang, Mark Wilson · 30 July 2026
Psychometric calibration for educational tests typically requires costly human response data. Large language models (LLMs) simulated examinees offer a promising route to early calibration, but their responses are too accurate and too uniform. We propose Cognitive Diagnostic Profiling (CDP), a zero-s…
Other topics in Management science and operations research
The topics the OpenAlex classification attaches to the same theme, most active first.
- Advanced Bandit Algorithms Research697 papers / 12 months+31%
- Stock Market Forecasting Methods391 papers / 12 months+420%
- Forecasting Techniques and Applications300 papers / 12 months+700%
- Data Quality and Management254 papers / 12 months+1650%
- Auction Theory and Applications98 papers / 12 months+100%
- Risk and Portfolio Optimization98 papers / 12 months+233%
