xAI unveiled an early preview of Grok 3 on February 19, 2025. It included general-purpose Grok 3 models, Grok 3 Think and Grok 3 mini Think reasoning variants, and DeepSearch. xAI said the family used roughly ten times the compute of its predecessor.

Release timeline

  • July 2023: xAI announced the company.
  • August 2024: Grok 2 and Grok 2 mini entered beta on X.
  • December 2024: xAI said Colossus used 100,000 NVIDIA Hopper GPUs.
  • February 19, 2025: xAI unveiled Grok 3, Grok 3 mini, reasoning variants, and DeepSearch.

Milestones are not universal model rankings. Valid comparisons require matched prompts, tools, inference budgets, scoring rules, and dates; see LLM evaluation metrics.

xAI described the Think models as trained with reinforcement learning and able to spend additional inference time on a problem. It did not disclose Monte Carlo tree search, a three-pass loop, fixed latency, or a universal accuracy gain. “Reasoning” does not mean human thought, and outputs still require verification; see AI hallucination.

xAI reported image-understanding results and later exposed search tools for web and X retrieval. Retrieval can supply current context without continually updating model weights. The launch materials did not establish a native audio architecture, specific vision encoder, live incremental-training pipeline, or claimed media-training-set sizes.

Infrastructure and undisclosed details

xAI said Colossus used 100,000 Hopper GPUs in December 2024 and later described a 200,000-GPU cluster. The later configuration is not proof of the exact Grok 3 training setup. xAI did not publish Grok 3's parameter or layer count, complete training corpus, cost, energy use, or measured hardware utilization. Peak device specifications are not measured training throughput.

Vendor-reported launch benchmarks

At the highest disclosed cons@64 test-time-compute setting, xAI reported 93.3% on AIME 2025, 84.6% on GPQA, and 79.4% on LiveCodeBench. It also reported a 1402 Elo Chatbot Arena snapshot for an early model. These results are conditional vendor claims, not a universal ranking.

ClaimStatus
Five-trillion parameters or 200+ layersNot disclosed
Ten petabytes of training dataNot disclosed
MCTS reasoning mechanismNot disclosed
Continual learning from live feedsNot established; retrieval is separate
$1 billion cost, 100 MW use, or 99.9% efficiencyNot disclosed or measured

What users could evaluate

Teams could test reasoning against answer keys, verify DeepSearch citations, probe image interpretation, and measure retrieval across long documents. None of this establishes clinical reliability, valid scientific simulation, or safe autonomous control. For broader background, read generative AI and LLMs.

Historical status

xAI announced Grok 4 in July 2025 and Grok 4.6 in August 2026. Readers evaluating current products should consult current documentation. This page remains a dated launch record; DeepSeek's emergence offers related market context.