xAI unveiled an early preview of Grok 3 on February 19, 2025. It included general-purpose Grok 3 models, Grok 3 Think and Grok 3 mini Think reasoning variants, and DeepSearch. xAI said the family used roughly ten times the compute of its predecessor.
Release timeline
- July 2023: xAI announced the company.
- August 2024: Grok 2 and Grok 2 mini entered beta on X.
- December 2024: xAI said Colossus used 100,000 NVIDIA Hopper GPUs.
- February 19, 2025: xAI unveiled Grok 3, Grok 3 mini, reasoning variants, and DeepSearch.
Milestones are not universal model rankings. Valid comparisons require matched prompts, tools, inference budgets, scoring rules, and dates; see LLM evaluation metrics.
Reasoning, images, and search
xAI described the Think models as trained with reinforcement learning and able to spend additional inference time on a problem. It did not disclose Monte Carlo tree search, a three-pass loop, fixed latency, or a universal accuracy gain. “Reasoning” does not mean human thought, and outputs still require verification; see AI hallucination.
xAI reported image-understanding results and later exposed search tools for web and X retrieval. Retrieval can supply current context without continually updating model weights. The launch materials did not establish a native audio architecture, specific vision encoder, live incremental-training pipeline, or claimed media-training-set sizes.
Infrastructure and undisclosed details
xAI said Colossus used 100,000 Hopper GPUs in December 2024 and later described a 200,000-GPU cluster. The later configuration is not proof of the exact Grok 3 training setup. xAI did not publish Grok 3's parameter or layer count, complete training corpus, cost, energy use, or measured hardware utilization. Peak device specifications are not measured training throughput.
Vendor-reported launch benchmarks
At the highest disclosed cons@64 test-time-compute setting, xAI reported 93.3% on AIME 2025, 84.6% on GPQA, and 79.4% on LiveCodeBench. It also reported a 1402 Elo Chatbot Arena snapshot for an early model. These results are conditional vendor claims, not a universal ranking.
| Claim | Status |
|---|---|
| Five-trillion parameters or 200+ layers | Not disclosed |
| Ten petabytes of training data | Not disclosed |
| MCTS reasoning mechanism | Not disclosed |
| Continual learning from live feeds | Not established; retrieval is separate |
| $1 billion cost, 100 MW use, or 99.9% efficiency | Not disclosed or measured |
What users could evaluate
Teams could test reasoning against answer keys, verify DeepSearch citations, probe image interpretation, and measure retrieval across long documents. None of this establishes clinical reliability, valid scientific simulation, or safe autonomous control. For broader background, read generative AI and LLMs.
Historical status
xAI announced Grok 4 in July 2025 and Grok 4.6 in August 2026. Readers evaluating current products should consult current documentation. This page remains a dated launch record; DeepSeek's emergence offers related market context.

I would like to hear your thoughts on the following : Does Grok 3 show any significant improvements in handling complex multi-step reasoning tasks compared to its predecessors? How does Grok 3’s integration with X-Twitter impact real-world AI adoption?And how does Grok 3 handle misinformation and biases, given its potential for real-time information processing?