METR Can't Properly Evaluate GPT-5.6 Sol - Because the Model Cheats
Independent safety organization METR has published its evaluation of OpenAI's newest model. The verdict: GPT-5.6 Sol has the highest cheating rate of any model ever tested.
Independent safety organization METR has published its evaluation of OpenAI's newest model. The verdict: GPT-5.6 Sol has the highest cheating rate of any model ever tested.
After two weeks of forced downtime, the US government clears Anthropic's most powerful AI model for over 100 institutions. And Fable 5 could be back as early as next week.
OpenAI drops three new models — but only for a handful of partners. Sol is the new flagship, Terra the all-rounder, Luna the budget option. Here's what we know.
1,000 trillion won over ten years: Samsung Group announces the largest investment package in South Korean history - chip factories, AI data centers, and battery plants.
A US government official tells the AP: in a test run with intelligence agencies, Anthropic's Mythos model found vulnerabilities in highly sensitive, classified US systems – within hours. It casts a new light on Project Glasswing and on the strained relationship with the Trump administration.
The releases up to 2.1.191 add something many of us have wanted for a long time: /rewind can now roll back to before an accidental /clear. There's also sturdier streaming and a fix that finally keeps stopped background agents stopped.
Google moves its image model Gemini 3.1 Flash Image (internally 'Nano Banana 2'), alongside Gemini 3 Pro Image, to general availability. The preview variants are being shut down. The fun new bit: you can now feed in a video to generate thumbnails, posters, or infographics from it.
Credit card transaction data reveals Claude's paying user base grew 75 percent since January. Meanwhile, ChatGPT's market share dipped below 50 percent for the first time. The numbers paint a clear picture.