01
AI in Education
Generative language models can now draft plausible test questions in seconds. Plausible is not the same as valid. Our work on automatic item generation (AIG) asks what has to be added to a prompt, a template, or a post-hoc filter before machine-written items behave like items written by a trained developer — stable difficulty, clean discrimination, no construct-irrelevant giveaways. The STAIR-AIG framework, developed with lab members, treats this as a human–AI collaboration problem rather than a prompting problem.
A parallel line looks at automatic machine translation (AMT) for assessment. Translated items are a standing threat to cross-national comparability; we study when machine translation is good enough, whether curriculum-aware prompting closes the gap, and what a human reviewer should be asked to check.
Both lines are oriented toward assessments that develop critical thinking rather than merely rank students, and both have produced patent filings with lab members.
02
Big Data Analytics
International large-scale assessments — PISA, PIAAC, NAEP, TIMSS — produce some of the richest educational datasets in existence, and almost none of it is complete. Every student sees a fraction of the item pool; every scale score is an imputation problem wearing a different hat.
We work on the population models that make those inferences defensible: latent regression and plausible-value methodology, advanced imputation under complex designs, and the consequences of model misspecification for the country rankings that draw all the attention.
A growing part of this work uses process data — response times, action logs, revision sequences — not as nuisance but as evidence. Timing information carries signal about engagement and strategy that the response itself does not, and incorporating it into population modeling changes what the scores mean.
03
Automated Scoring
Constructed-response items measure what multiple-choice items cannot, and they cost far more to score. Our work treats automated scoring as a supervised classification problem with psychometric constraints: the engine must not only agree with human raters on average, it must agree in ways that preserve the measurement properties of the scale across subgroups and languages.
Related work models the humans. Rater severity drifts, rating designs are sparse and often disconnected, and the linkage set that ties raters together determines whether the model is identified at all. We develop rater response models and monitoring procedures — including the use of scoring engines to monitor human raters rather than the other way around.
This line grew directly out of the machine-supported coding system operationalized for PISA constructed responses, and continues under an IEA-funded project on multilingual text scoring.
04
Adaptive Testing
Adaptive testing promises more precision per minute of testing time. Delivering on that promise in a heterogeneous population — eighty-odd countries, wildly different achievement distributions, one item pool — is a different problem from adaptive testing in a single well-characterized group.
We design and evaluate multistage adaptive testing (MST) structures for exactly that setting: how many stages, how wide the routing modules, how much overlap, and what happens to measurement precision at the tails where policy interest is highest. Introducing MST into PISA 2018 Reading and PISA 2022 Mathematics made these questions operational rather than hypothetical.
Ongoing methodological work addresses the diagnostics that adaptive designs complicate — local dependence, item fit, and adjusted residuals under quasi-independence — and the correction of accuracy-rate bias in data collected adaptively.