MetaTOC stay on top of your field, easily

From Tasks to Traits: Can LLMs Infer Personality From Brief Assessment Center Behavior?

International Journal of Selection and Assessment

Published online on

Abstract

["International Journal of Selection and Assessment, Volume 34, Issue 3, August 2026. ", "\nABSTRACT\nAssessment centers (ACs) could provide behavioral foundations for trait assessment but demand substantial resources, especially the labor‐intensive human rating process. As organizations adopt large language models (LLMs) for automated assessment, we establish initial psychometric benchmarks for LLM‐based trait inference from AC behavior. We processed recordings from 363 participants in brief simulated online AC group discussions through an automated pipeline: Whisper (speech recognition), Pyannote (speaker diarization), and Gemini 2.0 Flash (trait inference across the Big Five, narcissism, and intelligence). As a human benchmark, eight trained raters (four aggregated per transcript) independently inferred the same traits from a subsample of 100 group discussion transcripts. Aggregating multiple LLM runs yielded high reliability across traits (ICC(3,4) = 0.75–0.92). LLM convergent validity with trait tests was modest but above chance for all traits except conscientiousness, and varied by trait observability: openness and intelligence (r = 0.20), narcissism, extraversion, and agreeableness (r = 0.12–0.16), emotional stability and conscientiousness (r = 0.07–0.11). These findings lend support to AC behavior as a basis for measuring personality even under minimal information, with these validities likely a lower bound on what richer inputs would yield. Most notably, an LLM could infer personality from this behavior, with convergent validities descriptively higher than the human benchmark for most traits. Overall, constrained by discriminant validity, gender‐related adverse impact, and input dependencies (amount and clarity), the LLM trait inferences remained insufficient for standalone selection but can be produced at negligible marginal cost, positioning LLMs as a potential low‐cost contributor to multi‐source assessment pipelines."]