Grok 4.7 xHigh 在该基准测试中得分58%,Claude Fable 5.1 Max 得分59%。Grok 4.7 xHigh 超越了 GPT-6 Astra、GPT-5.6、Gemini、Kimi、GLM 以及几乎所有其他前沿模型。