77.1%. That’s the score Google’s brand-new Gemini 3.1 Pro just achieved on ARC-AGI-2, a benchmark specifically designed to test whether an AI can solve logic problems it has never encountered during training. Not pattern recognition. Not retrieved memory. Genuine reasoning from scratch.
Three months ago, Gemini 3 Pro scored 31.1…



