Aurora-2 was measured on the same public benchmarks the whole field reports. Here is how the flagship scored across reasoning, code, and math, how it stacked up against GPT-4o, Claude, Gemini, and Llama, and the scale it ran at in production. Every score is for the base Aurora-2 model unless noted.
Scores are pass@1 with standard few-shot prompting, reported on the public test splits. Aurora-2 Reason lifts MATH to 86.5% and GPQA Diamond to 69.0% with extended thinking enabled.
Aurora-2 ran the same public benchmarks as the models it competed with. Here is how it landed against GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro, and Llama 3.1 405B. The competitor figures are the numbers each lab published.
| Model | MMLU | GPQA Diamond | HumanEval | MATH |
|---|---|---|---|---|
| Aurora-2 | 89.4% | 60.2% | 90.1% | 75.9% |
| GPT-4o | 88.7% | 53.6% | 90.2% | 76.6% |
| Claude 3.5 Sonnet | 88.7% | 59.4% | 92.0% | 71.1% |
| Gemini 1.5 Pro | 85.9% | 46.2% | 84.1% | 67.7% |
| Llama 3.1 405B | 88.6% | 51.1% | 89.0% | 73.8% |
Aurora-2 leads the field on reasoning (MMLU and GPQA Diamond) and stays within a point of the best models on code (HumanEval) and math (MATH). It did not top every column, and that is the point: these were close, honest margins against the strongest models of the day. Competitor numbers are drawn from each model's published model card or technical report.
Aurora did not just look good on paper. It carried real traffic, in real time, for people in almost every country on earth.
Aurora-2 shipped in three variants so teams could trade speed for depth. All three now run inside Qai.
High volume, low latency
The balanced flagship
Extended thinking
Aurora-2 Flash, Aurora-2, and Aurora-2 Reason are all served through Qai. Point your API at Qai and you are talking to the same models.