Scoring German Alternate Uses Items Applying Large Language Models

Saretzki, Janika; Knopf, Thomas; Forthmann, Boris; Goecke, Benjamin; Jaggy, Ann-Kathrin; Benedek, Mathias; Weiss, Selina

Zitationshinweis

Bitte beziehen Sie sich beim Zitieren dieses Dokumentes immer auf folgenden Persistent Identifier (PID):
https://nbn-resolving.org/urn:nbn:de:0168-ssoar-102878-3

[Zeitschriftenartikel]

Saretzki, Janika

Knopf, Thomas

Forthmann, Boris

Goecke, Benjamin

Jaggy, Ann-Kathrin

Benedek, Mathias

Weiss, Selina

Abstract

The alternate uses task (AUT) is the most popular measure when it comes to the assessment of creative potential. Since their implementation, AUT responses have been rated by humans, which is a laborious task and requires considerable resources. Large language models (LLMs) have shown promising performance in automatically scoring AUT responses in English as well as in other languages, but it is not clear which method works best for German data. Therefore, we investigated the performance of different LLMs for the automated scoring of German AUT responses. We compiled German data across five research groups including ~50,000 responses for 15 different alternate uses objects from eight lab and online survey studies (including ~2300 participants) to examine generalizability across datasets and assessment conditions. Following a pre-registered analysis plan, we compared the performance of two fine-tuned, multilingual LLM-based approaches [Cross-Lingual Alternate Uses Scoring (CLAUS) and the Open Creativity Scoring with Artificial Intelligence (OCSAI)] with the Generative Pre-trained Transformer (GPT-4) in scoring (a) the original German AUT responses and (b) the responses translated to English. We found that the LLM-based scorings were substantially correlated with human ratings, with higher relationships for OCSAI followed by GPT-4 and CLAUS. Response translation, however, had no consistent positive effect. We discuss the generalizability of the results across different items and studies and derive recommendations and future directions.... weniger

Klassifikation
Allgemeine Psychologie

Freie Schlagwörter
GPT; alternate uses task; German; assessment; automated scoring; creativity; divergent thinking; large language models

Sprache Dokument
Englisch

Publikationsjahr
2025

Zeitschriftentitel
Journal of Intelligence, 13 (2025) 6

DOI
https://doi.org/10.3390/jintelligence13060064

ISSN
2079-3200

Status
Veröffentlichungsversion; begutachtet (peer reviewed)

Lizenz
Creative Commons - Namensnennung 4.0

Zitationshinweis

Export für Ihre Literaturverwaltung

Statistiken anzeigen

Weiterempfehlen