Aurora is a general-purpose frontier model. It reasoned, coded, saw, listened, and spoke more than 100 languages, all from one set of weights. Here is the full picture. To put it to work today, you use Qai.
Aurora carried a 256K-token context, roughly 500 pages, and reasoned across all of it at once. It planned before it answered, showed its work when asked, and caught its own mistakes mid-solution. On graduate-level science questions (GPQA Diamond) it scored 60.2%, ahead of GPT-4o and Claude 3.5 Sonnet.
Aurora wrote, refactored, and reviewed code in more than 80 languages. It scored 90.1% on HumanEval and resolved 51.4% of real issues on SWE-bench Verified, editing across whole repositories instead of one file at a time.
Text, images, and audio in one model. Aurora read charts, screenshots, handwriting, and diagrams, transcribed and reasoned over audio, and answered about all of it in a single conversation. No separate vision model, no handoffs.
Aurora reasoned and translated across more than 100 languages, holding tone and nuance rather than swapping words one for one. On multilingual grade-school math (MGSM) it scored 86.9%, close to its English performance.
Aurora called tools, APIs, and functions on its own, chained multi-step tasks, and knew when to stop and ask. It ran as the brain of autonomous agents that searched, booked, and wrote back to real systems without a human in every loop.
Aurora Guardrails filtered harmful output, resisted jailbreaks, and kept the model steerable when prompts got adversarial. Every release went through red-team review before it shipped to a single user.
Aurora is not sold on its own anymore. The models, the multimodal stack, and the agent runtime are all delivered through Qai.