This paper reports two studies investigating the emotional reasoning of Large Language Models (LLM). Previous research has suggested that LLMs are surprisingly accurate at predicting human emotions from text descriptions of situations and reason in a way that is consistent with appraisal theory—a leading theory of emotion. Study 1 tests this claim with a large multilingual corpus (English, French, and German) of autobiographical descriptions of emotionally charged events. We confirm that GPT-4, one of the most advanced and widely studied LLMs, shows a remarkable ability to predict emotion and appraisals. We further show this ability is language-independent, with accuracy being consistent across languages and unaffected by the language of the prompt. However, GPT-4 struggles to accurately predict certain emotions (shame, fear, and irritation) and fails to understand appraisal dimensions related to control and power. We repeat the experiments with Gemini-2.0-Flash and find a remarkably similar pattern of strengths and weaknesses, although it consistently outperforms GPT-4. Study 2 examines a possible mechanism for these failures based on the idea of cognitive appraisal bias. In psychological appraisal theory, appraisal bias is the idea that people evaluate situations in biased, often unrealistic ways. By testing both models on a set of situations designed to identify appraisal bias, we find they exhibit strong—but similar—appraisal bias; for example, evaluating situations as if they were a person high in agreeableness and low in power. We further offer evidence suggesting that LLMs could be debiased by incorporating a person's personality in the prompt. This research underscores LLMs' capabilities and limitations in emotional reasoning, though highlights one mechanism underlying this limitation and suggests an approach for addressing these limits.