Grok 4.5 just became the AI that hallucinates the least, and it costs a fraction of the other frontier models. xAI dropped it on the one benchmark that rewards correct answers and punishes confident wrong ones, and it ranks number one in the world.
Least hallucination is the real win
Grok 4.5 topped that reliability test because it knows how to say "I don't know" instead of bluffing. In the ranking, it lands ahead of GPT and Claude.
For real work, that matters more than it sounds. A model that admits when it is unsure is far more useful than one that confidently lies. Most benchmarks reward raw accuracy, but this one specifically penalizes the confident wrong answers that quietly break real workflows, and that is exactly where Grok 4.5 pulls ahead.
It is still genuinely smart
This is not a case of a model playing it safe by refusing to answer. Grok 4.5 set a new record on ARC-AGI-2, one of the hardest raw reasoning tests out there. It also scored 93% on GPQA Diamond, a set of PhD-level science questions where actual PhDs score around 65%.
xAI's pitch for the model is simple: the most intelligence per second.
Then there is the price
Grok 4.5 runs at $2 per million input tokens and $6 per million output tokens. That undercuts the other frontier models hard.
For anyone running agents, that is the headline. Agents that burn a lot of tokens all day suddenly cost a lot less to operate, without giving up reliability. If your AI keeps hallucinating, Grok 4.5 is worth a try.