a frontier model
Picture the leading edge of a map that explorers are still drawing: just past it lies territory no one has charted. A frontier model is an AI system sitting right at that leading edge of capability, among the most powerful general-purpose models anyone has built, doing things no previous system could reliably do. The word 'frontier' does double duty: these systems are both the most capable and the least understood.
More concretely, 'frontier model' usually means a general-purpose AI (typically a large model trained on huge amounts of data) that is at or beyond the current state of the art and could pose serious risks if misused or if it behaves in unintended ways. Because 'most capable' is fuzzy, regulators often reach for a rough proxy: the amount of computing power used to train it. A rule might say, for instance, that a model trained using more than some very large number of operations counts as frontier and triggers extra reporting and testing. That line is admittedly crude, but it gives a measurable handle on an otherwise slippery category.
The term matters because much of AI governance is deliberately narrow: rather than regulate every spreadsheet macro and spam filter, policymakers try to focus oversight on the handful of systems capable enough to cause large-scale harm. The flip side is that the frontier keeps moving. Today's frontier model is next year's ordinary tool, and a clever method can reach frontier-level capability with far less compute than a threshold assumes, so any fixed definition risks aiming at yesterday's danger.
A lab announces a new general-purpose model that scores higher than any predecessor on coding, reasoning, and science tasks, and can also help plan a cyberattack; regulators classify it as a frontier model and ask for its dangerous-capability test results before public release.
Frontier models are where capability and uncertainty are both highest, so they draw the most safety attention.
A compute threshold is a convenient legal handle, not a measure of danger. A small, narrowly trained model can be more harmful at one task than a giant general model, so 'frontier' tracks general capability, not risk in every case.