It's well-known at this point that LLMs don't handle spelling, syllables, rhythm, meter, or other word-form-based questions well due to tokenization -- sometimes sheer scale (or leaning on code) can get the right answer if they're lucky, but they're literally blind to the individual letters.
(Incidentally, go back in time even five years and this specific expectation of AI capability sounds comically overblown. "Everything's amazing and nobody's happy.")
(Incidentally, go back in time even five years and this specific expectation of AI capability sounds comically overblown. "Everything's amazing and nobody's happy.")