Related News
Sep 22, 2026
Grok 4.7 improves agent benchmarks, but task costs complicate the pitch
Grok 4.7 posts stronger long-horizon agent results at unchanged token rates, while independent testing finds higher usage, elapsed time and pay-per-token API cost per coding task.
Sep 21, 2026
OpenAI says its model resolved 100+ math problems, but hasn’t released the list
OpenAI reports that an internal model resolved more than 100 long-standing math problems. Its new independent advisory group can publish recommendations, but it cannot control company decisions or compel the release of the underlying proofs.
Sep 19, 2026
Claude leads 26% of Anthropic's AI research. What does that actually measure?
Anthropic says Claude leads 26% of its measured model-development work. Its task weighting, human review, monitoring boundaries and safety-compute denominators explain what that figure means.