The Evolution of Psychometrics In the Digital Era: From Classical Measurement Theory to Adaptive Assessment and Artificial Intelligence

Authors

  • Loso Judijanto IPOSS Jakarta

DOI:

https://doi.org/10.59890/mjst.v3i8.312

Keywords:

Psychometrics; Psychological Measurement; Validity; Reliability; Item Response Theory; Computerized Adaptive Testing; Process Data; Artificial Intelligence; Algorithmic Fairness; Digital Ethics

Abstract

Psychometrics is undergoing a fundamental change as psychological measurement increasingly shifts from static tests toward adaptive digital systems that are data-driven and supported by artificial intelligence. This article aims to explain how this transformation is changing the theoretical, methodological, and practical foundations of psychometrics, as well as to identify the scientific and ethical challenges that affect its credibility. The study uses a qualitative literature review with a narrative-integrative approach to peer-reviewed journal articles published since 2020. The synthesis shows that classical test theory remains important as a foundation, but item response theory, computerized adaptive testing, factor and network analysis, response time modeling, and machine learning expand the unit of analysis from final scores to response processes and multimodal data. Digitalization increases efficiency, personalization, and prediction accuracy, but it doesn't automatically improve validity. The quality of inferences still depends on clear constructs, proper reliability, measurement invariance, algorithmic fairness, transparency, privacy protection, and cross-cultural suitability. This article proposes a direction for psychometrics-by-design that integrates validity, fairness, data security, interpretability, and human oversight right from the instrument design stage. The implication is that the future of psychometrics isn’t just about using fancier technology, but building a measurement ecosystem that’s adaptive, auditable, human-centered, and backed by ongoing evidence of validity

References

Acikgoz, Y., Davison, K.H., Compagnone, M. and Laske, M. (2020) ‘Justice perceptions of artificial intelligence in selection’, International Journal of Selection and Assessment, 28, pp. 399–416. doi: 10.1111/ijsa.12306. Available at: https://doi.org/10.1111/ijsa.12306

Bao, Y., Shen, Y., Wang, S. and Bradshaw, L. (2021) ‘Flexible computerized adaptive tests to detect misconceptions and estimate ability simultaneously’, Applied Psychological Measurement, 45(1), pp. 3–21. doi: 10.1177/0146621620965730. Available at: https://doi.org/10.1177/0146621620965730

Barry, E.S., Merkebu, J. and Varpio, L. (2022) ‘State-of-the-art literature review methodology: A six-step approach for knowledge synthesis’, Perspectives on Medical Education, 11(5), pp. 281–288. doi: 10.1007/s40037-022-00725-9. Available at: https://doi.org/10.1007/s40037-022-00725-9

Belzak, W.C.M. and Bauer, D.J. (2020) ‘Improving the assessment of measurement invariance: Using regularization to select anchor items and identify differential item functioning’, Psychological Methods, 25(6), pp. 673–690. doi: 10.1037/met0000253. Available at: https://doi.org/10.1037/met0000253

Boehm, U., Marsman, M., van der Maas, H.L.J. and Maris, G. (2021) ‘An attention-based diffusion model for psychometric analyses’, Psychometrika, 86, pp. 938–972. doi: 10.1007/s11336-021-09783-0. Available at: https://doi.org/10.1007/s11336-021-09783-0

Buda, T.S., Guerreiro, J., Omana Iglesias, J., Castillo, C., Smith, O. and Matic, A. (2022) ‘Foundations for fairness in digital health apps’, Frontiers in Digital Health, 4, 943514. doi: 10.3389/fdgth.2022.943514. Available at: https://doi.org/10.3389/fdgth.2022.943514

Bürkner, P.-C. (2022) ‘On the information obtainable from comparative judgments’, Psychometrika, 87, pp. 1439–1472. doi: 10.1007/s11336-022-09843-z. Available at: https://doi.org/10.1007/s11336-022-09843-z

Cai, L. and Houts, C.R. (2021) ‘Longitudinal analysis of patient-reported outcomes in clinical trials: Applications of multilevel and multidimensional item response theory’, Psychometrika, 86, pp. 754–777. doi: 10.1007/s11336-021-09777-y. Available at: https://doi.org/10.1007/s11336-021-09777-y

Camilleri, J.A., Eickhoff, S.B. and Weis, S. (2021) ‘A machine learning approach for the factorization of psychometric data with application to the Delis Kaplan Executive Function System’, Scientific Reports, 11, 16896. doi: 10.1038/s41598-021-96342-3. Available at: https://doi.org/10.1038/s41598-021-96342-3

Carlo, A.D., Barnett, B.S. and Cella, D. (2021) ‘Computerized adaptive testing (CAT) and the future of measurement-based mental health care’, Administration and Policy in Mental Health and Mental Health Services Research, 48(5), pp. 729–731. doi: 10.1007/s10488-021-01123-9. Available at: https://doi.org/10.1007/s10488-021-01123-9

Coghlan, S. and D’Alfonso, S. (2021) ‘Digital phenotyping: An epistemic and methodological analysis’, Philosophy & Technology, 34(4), pp. 1905–1928. doi: 10.1007/s13347-021-00492-1. Available at: https://doi.org/10.1007/s13347-021-00492-1

Cortina, J.M., Sheng, Z., Keener, S.K., Keeler, K.R., Grubb, L.K., Schmitt, N., Tonidandel, S., Summerville, K.M., Heggestad, E.D. and Banks, G.C. (2020) ‘From alpha to omega and beyond! A look at the past, present, and (possible) future of psychometric soundness in the Journal of Applied Psychology’, Journal of Applied Psychology, 105(12), pp. 1351–1381. doi: 10.1037/apl0000815. Available at: https://doi.org/10.1037/apl0000815

Davidson, B.I. (2022) ‘The crossroads of digital phenotyping’, General Hospital Psychiatry, 74, pp. 126–132. doi: 10.1016/j.genhosppsych.2020.11.009. Available at: https://doi.org/10.1016/j.genhosppsych.2020.11.009

DeCarlo, L.T. (2021) ‘On joining a signal detection choice model with response time models’, Journal of Educational Measurement, 58(4), pp. 438–464. doi: 10.1111/jedm.12300. Available at: https://doi.org/10.1111/jedm.12300

Ellis, J.L. (2021) ‘A test can have multiple reliabilities’, Psychometrika, 86, pp. 869–876. doi: 10.1007/s11336-021-09800-2. Available at: https://doi.org/10.1007/s11336-021-09800-2

Epskamp, S. (2020) ‘Psychometric network models from time-series and panel data’, Psychometrika, 85, pp. 206–231. doi: 10.1007/s11336-020-09697-3. Available at: https://doi.org/10.1007/s11336-020-09697-3

Flake, J.K. and Fried, E.I. (2020) ‘Measurement schmeasurement: Questionable measurement practices and how to avoid them’, Advances in Methods and Practices in Psychological Science, 3(4), pp. 456–465. doi: 10.1177/2515245920952393. Available at: https://doi.org/10.1177/2515245920952393

Fleming, M.N. (2021) ‘Considerations for the ethical implementation of psychological assessment through social media via machine learning’, Ethics & Behavior, 31(3), pp. 181–192. doi: 10.1080/10508422.2020.1817026. Available at: https://doi.org/10.1080/10508422.2020.1817026

Fokkema, M., Iliescu, D., Greiff, S. and Ziegler, M. (2022) ‘Machine learning and prediction in psychological assessment: Some promises and pitfalls’, European Journal of Psychological Assessment, 38(3), pp. 165–175. doi: 10.1027/1015-5759/a000714. Available at: https://doi.org/10.1027/1015-5759/a000714

Fulmer, R., Davis, T., Costello, C. and Joerin, A. (2021) ‘The ethics of psychological artificial intelligence: Clinical considerations’, Counseling and Values, 66, pp. 131–144. doi: 10.1002/cvj.12153. Available at: https://doi.org/10.1002/cvj.12153

Goldhammer, F., Hahnel, C., Kroehne, U. and Zehner, F. (2021) ‘From byproduct to design factor: On validating the interpretation of process indicators based on log data’, Large-scale Assessments in Education, 9, 20. doi: 10.1186/s40536-021-00113-5. Available at: https://doi.org/10.1186/s40536-021-00113-5

Golino, H., Shi, D., Christensen, A.P., Garrido, L.E., Nieto, M.D., Sadana, R., Thiyagarajan, J.A. and Martinez-Molina, A. (2020) ‘Investigating the performance of exploratory graph analysis and traditional techniques to identify the number of latent factors: A simulation and tutorial’, Psychological Methods, 25(3), pp. 292–320. doi: 10.1037/met0000255. Available at: https://doi.org/10.1037/met0000255

Gonzalez, O. (2021) ‘Psychometric and machine learning approaches to reduce the length of scales’, Multivariate Behavioral Research, 56(6), pp. 903–919. doi: 10.1080/00273171.2020.1781585. Available at: https://doi.org/10.1080/00273171.2020.1781585

Goretzko, D. and Bühner, M. (2020) ‘One model to rule them all? Using machine learning algorithms to determine the number of factors in exploratory factor analysis’, Psychological Methods, 25(6), pp. 776–786. doi: 10.1037/met0000262. Available at: https://doi.org/10.1037/met0000262

Goretzko, D. and Bühner, M. (2022) ‘Factor retention using machine learning with ordinal data’, Applied Psychological Measurement, 46, pp. 406–421. doi: 10.1177/01466216221089345. Available at: https://doi.org/10.1177/01466216221089345

Guo, J., Xu, X., Ying, Z. and Zhang, S. (2022) ‘Modeling not-reached items in timed tests: A response time censoring approach’, Psychometrika, 87(3), pp. 835–867. doi: 10.1007/s11336-021-09810-0. Available at: https://doi.org/10.1007/s11336-021-09810-0

Hayes, A.F. and Coutts, J.J. (2020) ‘Use omega rather than Cronbach’s alpha for estimating reliability. But…’, Communication Methods and Measures, 14(1), pp. 1–24. doi: 10.1080/19312458.2020.1718629. Available at: https://doi.org/10.1080/19312458.2020.1718629

Henninger, M. and Plieninger, H. (2021) ‘Different styles, different times: How response times can inform our knowledge about the response process in rating scale measurement’, Assessment, 28(5), pp. 1301–1319. doi: 10.1177/1073191119900003. Available at: https://doi.org/10.1177/1073191119900003

Hilliard, A., Guenole, N. and Leutner, F. (2022) ‘Robots are judging me: Perceived fairness of algorithmic recruitment tools’, Frontiers in Psychology, 13, 940456. doi: 10.3389/fpsyg.2022.940456. Available at: https://doi.org/10.3389/fpsyg.2022.940456

Hong, M., Rebouças, D.A. and Cheng, Y. (2021) ‘Robust estimation for response time modeling’, Journal of Educational Measurement, 58, pp. 262–280. doi: 10.1111/jedm.12286. Available at: https://doi.org/10.1111/jedm.12286

Huang, S. and Cai, L. (2021) ‘Lord–Wingersky algorithm version 2.5 with applications’, Psychometrika, 86, pp. 973–993. doi: 10.1007/s11336-021-09785-y. Available at: https://doi.org/10.1007/s11336-021-09785-y

Hussey, I. and Hughes, S. (2020) ‘Hidden invalidity among 15 commonly used measures in social and personality psychology’, Advances in Methods and Practices in Psychological Science, 3(2), pp. 166–184. doi: 10.1177/2515245919882903. Available at: https://doi.org/10.1177/2515245919882903

Huth, K.B.S., Waldorp, L.J., Luigjes, J., Goudriaan, A.E., van Holst, R.J. and Marsman, M. (2022) ‘A note on the structural change test in highly parameterized psychometric models’, Psychometrika, 87(3), pp. 1064–1080. doi: 10.1007/s11336-021-09834-6. Available at: https://doi.org/10.1007/s11336-021-09834-6

Jebb, A.T., Ng, V. and Tay, L. (2021) ‘A review of key Likert scale development advances: 1995–2019’, Frontiers in Psychology, 12, 637547. doi: 10.3389/fpsyg.2021.637547. Available at: https://doi.org/10.3389/fpsyg.2021.637547

Jiang, Z., Seyedi, S., Griner, E., Abbasi, A., Rad, A.B., Kwon, H., Cotes, R.O. and Clifford, G.D. (2024) ‘Evaluating and mitigating unfairness in multimodal remote mental health assessments’, PLOS Digital Health, 3(7), e0000413. doi: 10.1371/journal.pdig.0000413. Available at: https://doi.org/10.1371/journal.pdig.0000413

Johnson, M.S., Liu, X. and McCaffrey, D.F. (2022) ‘Psychometric methods to evaluate measurement and algorithmic bias in automated scoring’, Journal of Educational Measurement, 59, pp. 338–361. doi: 10.1111/jedm.12335. Available at: https://doi.org/10.1111/jedm.12335

Kang, I., De Boeck, P. and Ratcliff, R. (2022) ‘Modeling conditional dependence of response accuracy and response time with the diffusion item response theory model’, Psychometrika, 87(2), pp. 725–748. doi: 10.1007/s11336-021-09819-5. Available at: https://doi.org/10.1007/s11336-021-09819-5

Kim, S.-H., Kwak, M., Bian, M., Feldberg, Z., Henry, T., Lee, J., Ölmez, İ.B., Shen, Y., Tan, Y., Tanaka, V., Wang, J., Xu, J. and Cohen, A.S. (2020) ‘Item response models in Psychometrika and psychometric textbooks’, Frontiers in Education, 5, 63. doi: 10.3389/feduc.2020.00063. Available at: https://doi.org/10.3389/feduc.2020.00063

Köchling, A., Riazy, S., Wehner, M.C. and Simbeck, K. (2021) ‘Highly accurate, but still discriminatory: A fairness evaluation of algorithmic video analysis in the recruitment context’, Business & Information Systems Engineering, 63, pp. 39–54. doi: 10.1007/s12599-020-00673-w. Available at: https://doi.org/10.1007/s12599-020-00673-w

Langenfeld, T. (2020) ‘Internet-based proctored assessment: Security and fairness issues’, Educational Measurement: Issues and Practice, 39(3), pp. 24–27. doi: 10.1111/emip.12359. Available at: https://doi.org/10.1111/emip.12359

Langer, M., Baum, K., König, C.J., Hähne, V., Oster, D. and Speith, T. (2021) ‘Spare me the details: How the type of information about automated interviews influences applicant reactions’, International Journal of Selection and Assessment, 29, pp. 154–169. doi: 10.1111/ijsa.12325. Available at: https://doi.org/10.1111/ijsa.12325

Leutner, F., Codreanu, S.-C., Brink, S. and Bitsakis, T. (2023) ‘Game based assessments of cognitive ability in recruitment: Validity, fairness and test-taking experience’, Frontiers in Psychology, 13, 942662. doi: 10.3389/fpsyg.2022.942662. Available at: https://doi.org/10.3389/fpsyg.2022.942662

Levy, R. (2020) ‘Implications of considering response process data for greater and lesser psychometrics’, Educational Assessment, 25(3), pp. 218–235. doi: 10.1080/10627197.2020.1804352. Available at: https://doi.org/10.1080/10627197.2020.1804352

Lin, Z., Chen, P. and Xin, T. (2021) ‘The block item pocket method for reviewable multidimensional computerized adaptive testing’, Applied Psychological Measurement, 45(1), pp. 22–36. doi: 10.1177/0146621620947177. Available at: https://doi.org/10.1177/0146621620947177

Liu, Y. and Wang, W. (2022) ‘Semiparametric factor analysis for item-level response time data’, Psychometrika, 87(2), pp. 666–692. doi: 10.1007/s11336-021-09832-8. Available at: https://doi.org/10.1007/s11336-021-09832-8

Lozano, J.H. and Revuelta, J. (2021) ‘A Bayesian generalized explanatory item response model to account for learning during the test’, Psychometrika, 86, pp. 994–1015. doi: 10.1007/s11336-021-09786-x. Available at: https://doi.org/10.1007/s11336-021-09786-x

Lu, J. and Wang, C. (2020) ‘A response time process model for not-reached and omitted items’, Journal of Educational Measurement, 57, pp. 584–620. doi: 10.1111/jedm.12270. Available at: https://doi.org/10.1111/jedm.12270

Lubbe, W., ten Ham-Baloyi, W. and Smit, K. (2020) ‘The integrative literature review as a research method: A demonstration review of research on neurodevelopmental supportive care in preterm infants’, Journal of Neonatal Nursing, 26(6), pp. 308–315. doi: 10.1016/j.jnn.2020.04.006. Available at: https://doi.org/10.1016/j.jnn.2020.04.006

Maassen, E., D’Urso, E.D., van Assen, M.A.L.M., Nuijten, M.B., De Roover, K. and Wicherts, J.M. (2025) ‘The dire disregard of measurement invariance testing in psychological science’, Psychological Methods, 30(5), pp. 966–979. doi: 10.1037/met0000624. Available at: https://doi.org/10.1037/met0000624

Man, K., Harring, J.R. and Zhan, P. (2022) ‘Bridging models of biometric and psychometric assessment: A three-way joint modeling approach of item responses, response times, and gaze fixation counts’, Applied Psychological Measurement, 46(5), pp. 361–381. doi: 10.1177/01466216221089344. Available at: https://doi.org/10.1177/01466216221089344

Muravyeva, E., Janssen, J., Specht, M. and Custers, B. (2020) ‘Exploring solutions to the privacy paradox in the context of e-assessment: Informed consent revisited’, Ethics and Information Technology, 22, pp. 223–238. doi: 10.1007/s10676-020-09531-5. Available at: https://doi.org/10.1007/s10676-020-09531-5

Park, J., Arunachalam, R., Silenzio, V. and Singh, V.K. (2022) ‘Fairness in mobile phone-based mental health assessment algorithms: Exploratory study’, JMIR Formative Research, 6(6), e34366. doi: 10.2196/34366. Available at: https://doi.org/10.2196/34366

Pellert, M., Lechner, C.M., Wagner, C., Rammstedt, B. and Strohmaier, M. (2024) ‘AI psychometrics: Assessing the psychological profiles of large language models through psychometric inventories’, Perspectives on Psychological Science, 19(5), pp. 808–826. doi: 10.1177/17456916231214460. Available at: https://doi.org/10.1177/17456916231214460

Provasnik, S. (2021) ‘Process data, the new frontier for assessment development: Rich new soil or a quixotic quest?’, Large-scale Assessments in Education, 9, 1. doi: 10.1186/s40536-020-00092-z. Available at: https://doi.org/10.1186/s40536-020-00092-z

Ramminger, J.J. and Jacobs, N. (2024) ‘Primacy of theory? Exploring perspectives on validity in conceptual psychometrics’, Frontiers in Psychology, 15, 1383622. doi: 10.3389/fpsyg.2024.1383622. Available at: https://doi.org/10.3389/fpsyg.2024.1383622

Röhner, J. and Holden, R.R. (2022) ‘Challenging response latencies in faking detection: The case of few items and no warnings’, Behavior Research Methods, 54, pp. 324–333. doi: 10.3758/s13428-021-01636-z. Available at: https://doi.org/10.3758/s13428-021-01636-z

Sijtsma, K. and Pfadt, J.M. (2021) ‘Part II: On the use, the misuse, and the very limited usefulness of Cronbach’s alpha: Discussing lower bounds and correlated errors’, Psychometrika, 86, pp. 843–860. doi: 10.1007/s11336-021-09789-8. Available at: https://doi.org/10.1007/s11336-021-09789-8

Sorrel, M.A., Abad, F.J. and Nájera, P. (2021) ‘Improving accuracy and usage by correctly selecting: The effects of model selection in cognitive diagnosis computerized adaptive testing’, Applied Psychological Measurement, 45(2), pp. 112–129. doi: 10.1177/0146621620977682. Available at: https://doi.org/10.1177/0146621620977682

Stefana, A., Damiani, S., Granziol, U., Provenzani, U., Solmi, M., Youngstrom, E.A. and Fusar-Poli, P. (2025) ‘Psychological, psychiatric, and behavioral sciences measurement scales: Best practice guidelines for their development and validation’, Frontiers in Psychology, 15, 1494261. doi: 10.3389/fpsyg.2024.1494261. Available at: https://doi.org/10.3389/fpsyg.2024.1494261

Suk, Y. and Han, K.T. (2024) ‘A psychometric framework for evaluating fairness in algorithmic decision making: Differential algorithmic functioning’, Journal of Educational and Behavioral Statistics, 49(2), pp. 151–172. doi: 10.3102/10769986231171711. Available at: https://doi.org/10.3102/10769986231171711

Sukhera, J. (2022) ‘Narrative reviews: Flexible, rigorous, and practical’, Journal of Graduate Medical Education, 14(4), pp. 414–417. doi: 10.4300/JGME-D-22-00480.1. Available at: https://doi.org/10.4300/JGME-D-22-00480.1

Trognon, A., Cherifi, Y.I., Habibi, I., Prudent, C. and Demange, L. (2022) ‘Using machine-learning strategies to solve psychometric problems’, Scientific Reports, 12, 18922. doi: 10.1038/s41598-022-23678-9. Available at: https://doi.org/10.1038/s41598-022-23678-9

Tutz, G. (2022) ‘Item response thresholds models: A general class of models for varying types of items’, Psychometrika, 87, pp. 1238–1269. doi: 10.1007/s11336-022-09865-7. Available at: https://doi.org/10.1007/s11336-022-09865-7

Ulitzsch, E., He, Q., Ulitzsch, V., Molter, H., Nichterlein, A., Niedermeier, R. and Pohl, S. (2021) ‘Combining clickstream analyses and graph-modeled data clustering for identifying common response processes’, Psychometrika, 86(1), pp. 190–214. doi: 10.1007/s11336-020-09743-0. Available at: https://doi.org/10.1007/s11336-020-09743-0

van der Linden, W.J. and Jiang, B. (2020) ‘A shadow-test approach to adaptive item calibration’, Psychometrika, 85, pp. 301–321. doi: 10.1007/s11336-020-09703-8. Available at: https://doi.org/10.1007/s11336-020-09703-8

Wyse, A.E. and McBride, J.R. (2021) ‘A framework for measuring the amount of adaptation of Rasch-based computerized adaptive tests’, Journal of Educational Measurement, 58(1), pp. 83–103. doi: 10.1111/jedm.12267. Available at: https://doi.org/10.1111/jedm.12267

Yadav, D. (2022) ‘Criteria for good qualitative research: A comprehensive review’, The Asia-Pacific Education Researcher, 31, pp. 679–689. doi: 10.1007/s40299-021-00619-0. Available at: https://doi.org/10.1007/s40299-021-00619-0

Zhang, S. and Chen, Y. (2022) ‘Computation for latent variable model estimation: A unified stochastic proximal framework’, Psychometrika, 87, pp. 1473–1502. doi: 10.1007/s11336-022-09863-9. Available at: https://doi.org/10.1007/s11336-022-09863-9

Zhang, S., Wang, Z., Qi, J., Liu, J. and Ying, Z. (2023) ‘Accurate assessment via process data’, Psychometrika, 88(1), pp. 76–97. doi: 10.1007/s11336-022-09880-8. Available at: https://doi.org/10.1007/s11336-022-09880-8

Zickar, M.J. (2020) ‘Measurement development and evaluation’, Annual Review of Organizational Psychology and Organizational Behavior, 7, pp. 213–232. doi: 10.1146/annurev-orgpsych-012119-044957. Available at: https://doi.org/10.1146/annurev-orgpsych-012119-044957

Published

2026-09-03

Issue

Section

Articles