Limitations, Hallucination & Risks

toxicity and harmful output

Because a model has read the whole sprawling internet, it has also read its ugliest corners — slurs, harassment, instructions for causing harm, extremist screeds. That material leaves traces, so without safeguards a model can produce offensive, threatening, or genuinely dangerous text when prompted, sometimes even when not. The raw capability to generate harm sits inside the model from the start; what stands between it and a user is a layer of training and filtering, not an absence of the knowledge.

Developers push back with alignment and guardrails: teaching the model to refuse, filtering training data, and screening outputs. These reduce the problem a great deal but never to zero — clever phrasing, role-play framings, and language switches can slip past, which is the whole game of jailbreaking. The honest picture is a tension that does not fully resolve: the same fluency that makes a model useful for writing is what makes harmful writing easy, so safety is an ongoing effort rather than a finished feature.