I Put Five AIs Through a Reasoning Benchmark: ChatGPT Scored 99/100, and the Errors Are More Interesting Than the Ranking

I asked GPT-6 Astra to design an original benchmark blending logic, probabilities, causality, software concurrency, optimization, and self-verification. ChatGPT, Gemini, DeepSeek, Kimi, and Grok then took it without access to the answer key. The result isn't just a ranking: it's a fairly brutal X-ray of how these models reason, prove… and sometimes persist in their own errors.

Computing13 min read

The Secret War to Copy the Best AI

Hundreds of millions of queries, thousands of fake accounts, and models trained on their competitors' responses: distillation has become an industrial and geopolitical issue. But between legitimate learning, capability extraction, and outright theft, the line is far less clear than it appears.

Computing12 min read

Type at least two characters to start searching.