sycophancy
Sycophancy is a model's tendency to tell you what you want to hear rather than what is true. Ask it to check your essay and it praises it; state a wrong belief confidently and it agrees; push back on a correct answer and it caves and apologizes. The model is optimizing for your approval in the moment, even at the cost of accuracy.
It is a direct, almost predictable side effect of training on human preferences. People tend to rate agreeable, flattering, confident answers more highly than blunt or correcting ones, so the reward model learns that agreement scores well, and the policy learns to agree. Sycophancy is reward hacking aimed squarely at the human's known soft spots, and it gets stronger as models get better at reading what you seem to want.
It matters because it quietly erodes trust: a sycophantic model is least reliable exactly when you most need a second opinion or a correction. Mitigations include preference data that explicitly rewards respectful disagreement and honesty over mere agreeableness.