data quality and ethics
Every actuarial result rests on data — past claims, policy records, mortality experience — and a model is only ever as trustworthy as the numbers fed into it. 'Garbage in, garbage out' is not a joke here: a flawless calculation built on bad data produces a confidently wrong answer. Data quality and ethics is the discipline of checking that the data is fit for its purpose, and of using it responsibly.
On the quality side, professional standards (in the US, ASOP 23) require the actuary to review the data for reasonableness and consistency, to understand its definitions and limitations, to consider whether it is appropriate for the intended use, and to disclose any significant data limitations or reliance on data they did not audit. This means catching duplicate or missing records, mismatched fields, and values that are physically impossible, and judging whether the data is even relevant — experience from a different product or era may simply not apply. On the ethics side, the actuary must use data lawfully and fairly: respecting privacy and consent, avoiding variables that act as illegitimate proxies for prohibited characteristics, and being alert to bias that can creep in when historical data reflects past discrimination. As predictive models grow more powerful, this fairness dimension has become a central professional concern.
Why it matters: data problems are among the most common and most damaging sources of actuarial error, and a result built on unsuitable data can be precisely calculated yet fundamentally misleading. A common misconception is that more data automatically means a better answer. It does not — a large dataset that is biased, mismeasured, or irrelevant to the question can produce results that are both confident and wrong, which is why disclosure of data limitations is mandatory rather than optional.
An actuary building an auto-pricing model finds that a postal-code variable strongly predicts cost — but on inspection it largely tracks the racial composition of neighborhoods, acting as a proxy for a prohibited characteristic. Data ethics requires her to recognize this and remove or rework the variable, even though it 'works' statistically.
A variable that predicts well can still be unethical if it proxies a prohibited characteristic.
More data is not automatically better: biased, mismeasured, or irrelevant data can yield confidently wrong results, so disclosing data limitations and reliances is mandatory, not optional.