Research

What we work on

Four connected lines of inquiry, all circling the same underlying question: what does it take for a measurement of learning to deserve the weight placed on it?

01

AI in Education

Generative language models can now draft plausible test questions in seconds. Plausible is not the same as valid. Our work on automatic item generation (AIG) asks what has to be added to a prompt, a template, or a post-hoc filter before machine-written items behave like items written by a trained developer — stable difficulty, clean discrimination, no construct-irrelevant giveaways. The STAIR-AIG framework, developed with lab members, treats this as a human–AI collaboration problem rather than a prompting problem.

A parallel line looks at automatic machine translation (AMT) for assessment. Translated items are a standing threat to cross-national comparability; we study when machine translation is good enough, whether curriculum-aware prompting closes the gap, and what a human reviewer should be asked to check.

Both lines are oriented toward assessments that develop critical thinking rather than merely rank students, and both have produced patent filings with lab members.

02

Big Data Analytics

International large-scale assessments — PISA, PIAAC, NAEP, TIMSS — produce some of the richest educational datasets in existence, and almost none of it is complete. Every student sees a fraction of the item pool; every scale score is an imputation problem wearing a different hat.

We work on the population models that make those inferences defensible: latent regression and plausible-value methodology, advanced imputation under complex designs, and the consequences of model misspecification for the country rankings that draw all the attention.

A growing part of this work uses process data — response times, action logs, revision sequences — not as nuisance but as evidence. Timing information carries signal about engagement and strategy that the response itself does not, and incorporating it into population modeling changes what the scores mean.

03

Automated Scoring

Constructed-response items measure what multiple-choice items cannot, and they cost far more to score. Our work treats automated scoring as a supervised classification problem with psychometric constraints: the engine must not only agree with human raters on average, it must agree in ways that preserve the measurement properties of the scale across subgroups and languages.

Related work models the humans. Rater severity drifts, rating designs are sparse and often disconnected, and the linkage set that ties raters together determines whether the model is identified at all. We develop rater response models and monitoring procedures — including the use of scoring engines to monitor human raters rather than the other way around.

This line grew directly out of the machine-supported coding system operationalized for PISA constructed responses, and continues under an IEA-funded project on multilingual text scoring.

04

Adaptive Testing

Adaptive testing promises more precision per minute of testing time. Delivering on that promise in a heterogeneous population — eighty-odd countries, wildly different achievement distributions, one item pool — is a different problem from adaptive testing in a single well-characterized group.

We design and evaluate multistage adaptive testing (MST) structures for exactly that setting: how many stages, how wide the routing modules, how much overlap, and what happens to measurement precision at the tails where policy interest is highest. Introducing MST into PISA 2018 Reading and PISA 2022 Mathematics made these questions operational rather than hypothetical.

Ongoing methodological work addresses the diagnostics that adaptive designs complicate — local dependence, item fit, and adjusted residuals under quasi-independence — and the correction of accuracy-rate bias in data collected adaptively.

Emerging direction

AI Measurement Science

Height and weight are easy. The competencies that matter most after school — communication, collaboration, creativity, critical thinking, and increasingly the ability to work productively with an AI system — are multidimensional, unobservable, and poorly served by questionnaires and multiple-choice tests.

The lab's longer-term aim is to establish AI Measurement Science as a coherent field. Three threads run through it: measuring capability from the process data learners generate while solving problems rather than from the final answer alone; measuring human–AI collaboration as a construct in its own right; and applying psychometric method to language models themselves, so that choosing a model for a task becomes an evidence-based decision rather than a vibe. There is no systematic program of this kind in Korea yet. Building one is the goal.

Research collaborations

Current research projects

Collaboration with Macat on Critical Thinking Skills

MindScale collaborates with Macat on automatic item generation (AIG) and machine translation (MT) for critical thinking assessments and on Macat's project with the OECD.

LLM-based Multi-Agent System for Automatic Educational Content Generation

MindScale is developing an LLM-based multi-agent system to generate and evaluate educational content for critical thinking skills.

CELLA Global Project

MindScale represents Korea in the global CELLA research project.

EBS 진단-형성평가 문항 현황 분석 및 개선 방안 연구

MindScale collaborates with EBS to evaluate assessment items for elementary and middle school students and develop an item bank.

TIMSS 2023 Longitudinal Study

MindScale collaborates with KICE to examine growth among Korean students in Grades 4 and 8 using TIMSS 2023 longitudinal data.

Assessment operations

PISA psychometrics

Four cycles of the OECD Programme for International Student Assessment, from junior analyst to co-lead of psychometric operations to the assessment's technical advisory body.

  • 2026–present Member, PISA Technical Advisory Group The first Korean researcher appointed to the OECD's technical advisory body for PISA, and currently its only member from Asia. The group of roughly eight to ten researchers reviews the technical standards behind the assessment. Representation of non-alphabetic writing systems — Korean, Japanese, Chinese — in test design and scoring is an explicit part of the brief she brings to it.
  • PISA 2022 Co-lead, psychometric operations With Dr. Fred Robin (ETS) and Dr. Peter van Rijn (ETS Global). Adapted the analysis plan to COVID-19 disruptions across participating countries, extended multistage adaptive testing to Mathematics, and examined innovations in the background questionnaire.
  • PISA 2018 Co-lead, psychometric operations Introduced the multistage adaptive testing design for Reading with Dr. Kentaro Yamamoto (ETS); operationalized the machine-supported coding system for constructed-response items; co-led the “Item Response Theory and Population Modeling in Large-scale Assessments” workshop for PISA National Project Managers with Dr. Eugenio Gonzalez (ETS).
  • PISA 2015 Psychometric analyst Supported the transition to computer-based assessment under Prof. Matthias von Davier (Boston College) and Dr. Kentaro Yamamoto (ETS), and contributed to the technical report.

Funding

Grants

  • 2022–2023 Automatic Scoring of Text Responses in Multilingual Assessment Data IEA Research and Development Tier II project · $99,060
  • 2016–2020 Learning Progression-based and NGSS-aligned Formative Assessment for Using Mathematical Thinking in Science Institute of Education Sciences (IES) · $1,396,496
  • 2020–2021 Innovations in Design, Modeling, and Process Data ETS Internal Research Allocation
  • 2019–2020 Psychometric Advances in Large-scale Survey Assessments ETS Internal Research Allocation
  • 2018–2019 Conditional Independence between Item Responses and Response Time ETS Internal Research Allocation

Technology transfer

Patents

Filed with the Korean Intellectual Property Office, together with lab members and collaborators. Titles are given in Korean as filed.

  • 2025-03-20 자동문항 생성 기술의 최적화된 사용을 위한 전자 장치 및 이의 동작 방법 Electronic device and operating method for the optimized use of automatic item generation technology. Application 10-2025-0036007 · with Euigyum Kim (Sogang University)
  • 2024-12-23 문항반응이론을 활용하여 맞춤형 평가 상황에서 수집한 데이터의 정답률 편향을 교정하는 방법 및 이를 위한 장치 Method and apparatus for correcting accuracy-rate bias in data collected under adaptive assessment, using item response theory. Application 10-2024-0194037 · with Yeonwho Kim (Hongik University Science High School)
  • 2024-11-12 프롬프팅에 기반한 자동문항생성 방법과 장치 Prompt-based method and apparatus for automatic item generation. Application 10-2024-0160356 · with Seewoo Li (UCLA) and Jihoon Ryoo (Yonsei University; CLASS-ANALYTICS)

Recognition

Awards & honors

  • 2023, 2024Excellence in Research AwardSogang University
  • 2024Excellence in Teaching AwardSogang University
  • 2017, 2019, 2021SPOT AwardEducational Testing Service
  • 2020Operational Excellence Service AwardEducational Testing Service
  • 2009–2014Doctoral ScholarshipKorea Foundation for Advanced Studies
  • 2012Travel grant, UCEC InstituteUniversity of California Educational Evaluation Center
  • 2007–2008Graduate researcher scholarshipBrain Korea 21 Academic Leadership Institute

MindScale is a research lab of the AI Behavioral Studies and the Graduate School of Education.