Abstract:
Large language models (LLMs) have demonstrated outstanding performance in natural language processing, prompting extensive discussion in psychology regarding whether they possess human-like cognitive abilities. This study systematically reviewed 101 empirical studies evaluating LLMs across five classical domains of cognition: creative thinking, memory monitoring and metacognition, reasoning, problem solving and executive control, and theory of mind. Building on this review, the study used the latest GPT-5.4 model to conduct replication checks and minor perturbation tests on 14 representative studies, aiming to further assess the robustness of its cognitive abilities. The results showed that mainstream LLMs often achieve scores approaching or even surpassing human baselines on standardized and static classical psychological tasks. However, they also exhibit strong framing effects, with performance declining sharply when faced with changes in task format, minor semantic perturbations, or complex contexts requiring dynamic information updating. These findings suggest that the cognitive abilities currently displayed by LLMs are better understood as explicit behavioral simulations based on linguistic representations, and that LLMs have not yet developed a stable human-like cognitive system with endogenous mechanisms.