Benchmarks

The numbers.

Aurora-2 was measured on the same public benchmarks the whole field reports. Here is how the flagship scored across reasoning, code, and math, how it stacked up against GPT-4o, Claude, Gemini, and Llama, and the scale it ran at in production. Every score is for the base Aurora-2 model unless noted.

Reasoning and knowledge

MMLU Broad knowledge, 57 subjects
89.4%
MMLU-Pro Harder, reasoning-heavy
76.8%
GPQA Diamond Graduate-level science
60.2%
Big-Bench Hard Multi-step reasoning
88.9%
HellaSwag Commonsense inference
95.3%

Code

HumanEval Function synthesis
90.1%
MBPP Everyday programming
87.9%
SWE-bench Verified Real GitHub issues
51.4%
LiveCodeBench Contamination-free coding
59.8%

Math

GSM8K Grade-school word problems
95.6%
MATH Competition mathematics
75.9%
MGSM Multilingual math, 10 languages
86.9%

Scores are pass@1 with standard few-shot prompting, reported on the public test splits. Aurora-2 Reason lifts MATH to 86.5% and GPQA Diamond to 69.0% with extended thinking enabled.


Head to head

Measured against
the frontier.

Aurora-2 ran the same public benchmarks as the models it competed with. Here is how it landed against GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro, and Llama 3.1 405B. The competitor figures are the numbers each lab published.

Model MMLU GPQA Diamond HumanEval MATH
Aurora-2 89.4% 60.2% 90.1% 75.9%
GPT-4o 88.7% 53.6% 90.2% 76.6%
Claude 3.5 Sonnet 88.7% 59.4% 92.0% 71.1%
Gemini 1.5 Pro 85.9% 46.2% 84.1% 67.7%
Llama 3.1 405B 88.6% 51.1% 89.0% 73.8%

Aurora-2 leads the field on reasoning (MMLU and GPQA Diamond) and stays within a point of the best models on code (HumanEval) and math (MATH). It did not top every column, and that is the point: these were close, honest margins against the strongest models of the day. Competitor numbers are drawn from each model's published model card or technical report.


In production

Frontier scores are
one thing. Scale is another.

Aurora did not just look good on paper. It carried real traffic, in real time, for people in almost every country on earth.

2.1T
Tokens per day, peak
40M
Conversations handled
99.98%
Serving uptime
280ms
Median first token
3.2M
Monthly users, peak
6.4M
API calls per day
24,000
GPUs at training peak
190+
Countries served
The family

Three sizes, one brain.

Aurora-2 shipped in three variants so teams could trade speed for depth. All three now run inside Qai.

Aurora-2 Flash

High volume, low latency

  • Context128K tokens
  • Active params12B
  • Speed320 tokens/sec
  • Best forChat, classification, extraction

Aurora-2 Reason

Extended thinking

  • Context1M tokens
  • Active params41B + thinking
  • SpeedDeliberate
  • Best forResearch, proofs, hard debugging
Same models, new door

Numbers like these
now run as Qai.

Aurora-2 Flash, Aurora-2, and Aurora-2 Reason are all served through Qai. Point your API at Qai and you are talking to the same models.