Large language models and other artificial intelligence systems are commonly evaluated through the correctness of their observable outputs. However, an apparently correct answer does not necessarily ...