GPT-5.6 Sol Proves a 50-Year-Old Math Conjecture — in Under an Hour
OpenAI publishes a machine-verified proof of the Cycle Double Cover Conjecture, produced by GPT-5.6 Sol Ultra with 64 subagents. Impressive — but not yet peer-reviewed.
Topic
173 articles · Page 3 of 9
OpenAI publishes a machine-verified proof of the Cycle Double Cover Conjecture, produced by GPT-5.6 Sol Ultra with 64 subagents. Impressive — but not yet peer-reviewed.
Google DeepMind pushes Gemini 3.5 Pro to July 17 — and rebuilds the model from scratch. The old architecture gets scrapped, a fresh pre-training run started. An expensive signal in the frontier race.
GPT-5.6 is now public — weeks after the Trump administration asked OpenAI to limit the launch to a few vetted partners. And the timing lines up exactly with Fable 5 leaving Anthropic's subscription plans.
OpenAI replaces Advanced Voice Mode with GPT-Live — a full-duplex architecture that listens and speaks while you're still talking. For heavy tasks, the model hands off to a frontier model in the background.
First OpenAI told the community to switch to SWE-Bench Pro. Now it's withdrawing that recommendation: up to 34% of tasks are defective. A wake-up call for anyone who takes coding benchmarks at face value.
A CNBC report shows more US companies turning to open Chinese models. The reason is simple — they're often 60 to 90 percent cheaper than the top models from Anthropic and OpenAI.
Tencent just released Hy3, an open-source MoE model that competes with models two to five times its size on many benchmarks. And it costs almost nothing to run.
After Chinese models captured up to 46 percent of US enterprise API traffic, Washington is reacting. Two committees are investigating — and the first companies have already received letters.
Anthropic found an inner 'workspace' inside Claude — a small zone where the model holds thoughts it never actually says out loud. Not consciousness, but fascinating all the same.
The well-known AI blogger had Claude Fable produce a complete sqlite-utils version over a weekend. Fable caught a critical data loss bug that Willison himself had missed.
Meituan revealed that the anonymous model 'Owl Alpha', which topped the OpenRouter charts for weeks, is actually LongCat-2.0 — a 1.6-trillion-parameter model trained without a single Nvidia chip.
Google brings its AI agent Gemini Spark to macOS. The desktop agent works with local files, connects to third-party apps, and competes directly with Claude Desktop and Cowork.
The DeepReinforce collective releases a model family that builds a 'scaffold' before it codes — a learnable architecture for the task. The flagship beats Opus 4.7 on Terminal-Bench 2.1.
Anthropic bundles databases, compute clusters, and domain expertise into one app. A reviewer agent checks citations and calculations — and every figure carries the code that made it.
Sonnet 5 gets close to Opus 4.8 — at a fraction of the price. It's the default model for Free and Pro from today, with 2 dollars per million input tokens as an intro price.
Google promised Gemini 3.5 Pro for June. Today is the 30th — and the model isn't here. The prediction markets got it right.
Google can't deliver as much Gemini capacity as Meta wants. Meta tells staff to use tokens more efficiently. A symptom of the broader compute crisis.
With Anthropic's export ban on Mythos and Fable 5 dragging on, two Asian labs are stepping into the space it left behind: Tokyo's Sakana AI and Chinese cybersecurity giant 360 — both pitching 'frontier capability without export risk'.
Sakana AI didn't train another giant model with Fugu. It trained a 7-billion-parameter conductor that routes tasks across a swappable pool of frontier models. CEO David Ha calls it the next stage — 'beyond bigger models'.
Unconventional AI just dropped its Un-0 model series. The pitch: run AI on tiny oscillators instead of traditional chips. Sounds wild, but Jeff Bezos put $475 million behind it.