Learning from Human Feedback — RLHF & Reward Modeling

helpful, harmless, and honest

If you had to sum up in three words what a good AI assistant should be, you might land on: helpful, harmless, and honest. That is exactly the HHH framing, a compact statement of the goal that RLHF and Constitutional AI are trying to train toward. Helpful: actually do what the user needs. Harmless: avoid causing or enabling harm. Honest: tell the truth, including about your own limits and uncertainty, rather than bluffing.

The three are deliberately simple, but the interesting part is that they pull against each other. The most helpful answer to 'how do I make a weapon?' is the most harmful one, so harmlessness must sometimes override helpfulness. An honest 'I don't know' is less immediately satisfying than a confident made-up answer, so honesty must sometimes override apparent helpfulness too. A lot of practical alignment work is really about negotiating these trade-offs case by case, and the preferences fed into a reward model are where those trade-offs get encoded.

HHH is useful as a north star and a shared vocabulary, but be honest about its limits. Each word hides enormous unspecified detail (harmless to whom, helpful by whose definition, honest about what?), so it is a direction, not a precise objective. Honesty in particular is the hardest to train and to verify, since a model can produce true-sounding text without any commitment to truth, and RLHF tends to reward what sounds good to raters, which can quietly trade honesty away for agreeableness, the failure we call sycophancy.

Asked to confirm a false 'fact' the user clearly believes, the honest move is a polite correction; a sycophantic model trained only to please might agree instead, satisfying helpfulness while breaking honesty.

The three H's often conflict; alignment is partly about negotiating the trade-offs.

HHH is a direction, not a precise specification: each word hides hard questions (harmless to whom?), and honesty is the toughest to train, partly because RLHF can quietly trade it for agreeableness.

Also called
HHHthe HHH framing有用無害誠實