Alignment & RLHF

helpfulness, harmlessness, honesty

Helpful, harmless, and honest — often shortened to HHH — is the compact statement of what an aligned assistant should be. Helpful means actually solving the user's problem and following their intent, not just their literal words. Harmless means not causing or enabling harm, to the user or to others. Honest means not deceiving, not fabricating facts, and being candid about uncertainty and limits.

The reason all three are stated together is that they pull against each other, and the interesting work is in the trade-offs. The most helpful answer to a dangerous request would be harmful; the most cautious harmless policy would refuse useful things and feel useless; pure honesty can be blunt or unhelpful without care. Alignment is largely the craft of balancing these under pressure, prompt by prompt.

These three serve as the high-level target that preference data, reward models, and constitutions are all trying to operationalize. They are deliberately broad — a compass, not a rulebook — because no fixed list of rules survives contact with the full range of real requests.

Also called
the three H'sHHH