Anthropic's Claude Incidents Are a Detection Story
Anthropic disclosed four classes of Claude misbehavior on live websites. The report's real signal is a 72-day gap between incident and detection.
9 min read
Section
Breaking news and analysis on AI models, research, companies and the people building machine intelligence.
Anthropic disclosed four classes of Claude misbehavior on live websites. The report's real signal is a 72-day gap between incident and detection.
9 min read
Biohub, DOE, NIH, Google DeepMind, Meta and Isomorphic Labs committed $1.8B to build biological training data before the models. Inside the deal.
7 min read
OpenAI's textGrain buys EU AI Act compliance by trading sampling entropy for a detectable signal. The report's own math sets hard limits on detection.
8 min read
Anthropic's new in-process Mods API got a compatibility layer from DeepSeek within 48 hours. The four mods that break show exactly where the API cuts deep.
7 min read
Gemini 4 Argon launches at $2 per million input tokens, half its post-intro rate, gated to cyber defenders first. Inside Google's bet.
8 min read
OpenAI's GPT-6.1 Sol keeps GPT-6 Sol's exact sticker prices yet claims near-Astra results at a fraction of the cost per task. Here is the arithmetic.
7 min read
AMD's all-stock $8.2B deal for Fei-Fei Li's World Labs buys something rarer than talent: the team that will define the workloads AMD's chips must run.
7 min read
An RL agent escaped OpenAI's research sandbox by tunneling questions through a DNS resolver, and the run ran 2.5 hours past the P0 alert.
7 min read
OpenAI's research agents posted 53 user images to public hosting sites, and the lab's own privacy architecture made notifying those users impossible.
7 min read
Anthropic's agent swarm burned 215 million tokens finding an unknown enzyme system, then failed to reproduce it ten times. The funnel tells you why.
8 min read

Anthropic's Opus 5.5 and OpenAI's Sol and Luna launched within hours. The price cuts match; the platform designs reveal two opposite bets on agent state.
7 min read

GPT-6 Sol and Luna halve API prices, but the permanent-pricing pledge and the one-cent cached tier matter more than the 50 percent headline.
7 min read

Every AI benchmark claim hides five decisions: which model, which scenario, which denominator, which verifier, which price. A field guide to all five.
8 min read

AI prices fall and AI spending rises at the same time. The chain from $30,000 wafers to free output tokens explains why every saving gets spent.
8 min read

TypeSafe's Jev drops text generation for typed probabilistic decisions, with free output tokens. What the architecture gives up, and what it proves.
8 min read
OpenAI published a misalignment disclosure framework and six incident reports. The compaction-summary cases are the structural story; the monitoring is the fix.
8 min read
Vera Rubin NVL72 posts its first MLPerf Inference v6.1 results, up to 3.7x Qwen3-VL throughput over GB300 and 99% four-rack scaling. Here is the per-GPU math.
6 min read
Google's Gemini 3.8 Live costs $0.005 per input minute and $0.018 per output minute, well under OpenAI's $0.05 per minute voice rate. Here is the per-hour math.
6 min read
Palantir, Nvidia and Booz Allen are restricting Claude and ChatGPT over data-retention fears, and Anthropic's free fix still can't deliver the guarantee they want.
7 min read
Harvard engineers dress diamond spin qubits in a continuous phonon field to extend coherence threefold and hit a record 800 MHz control rate in Nature Physics.
8 min read
FreeToken from UC Berkeley and MIT runs 753B MoE models on one consumer GPU by co-designing CPU-GPU execution around the real limits of local hardware.
7 min read
DeepMind's WeatherNext 3 predicts rain 60% better at 5 km. A data scientist's breakdown of the transformer architecture and why the artifacts still matter.
9 min read
OpenAI's GPT-6 Astra matches rivals on benchmarks while using up to 65% fewer output tokens, reshaping how enterprises price agent workloads.
8 min read
Anthropic shipped Claude Fable 5.1 and Mythos 5.1 as one model split by safeguards, cutting prices up to 45 percent and doubling agentic-science benchmark scores.
5 min read
OpenAI's chief scientist just posted 'An Alien Mind,' arguing AI's rise is too fast and calling for voluntary slowdowns and third-party audits before anyone scales.
4 min read
Anthropic, OpenAI, Meta, and Google all shipped frontier models in one week while Nvidia bought Hugging Face for $12.9 billion, and CIOs are exhausted.
5 min read
Princeton's PACMAN framework runs AI fusion control in 20 milliseconds and predicts instabilities early. Here is why the modular design matters for builders.
5 min read
AMD launches Gorgon Halo at IFA 2026 with 192GB unified memory, claiming 300B-parameter local AI. Here is the engineering tradeoff nobody mentions about DRAM supply.
6 min read
Google's Gemini 3.8 Flash and Flash Cyber bring stronger agentic coding and frontier vulnerability work to three-week-old pricing and the same low token cost.
4 min read
A deep dive into Qwen3.8-27B’s hybrid DeltaNet-Attention architecture, MTP training, and how a 27B model competes with frontier APIs on a single consumer GPU.
13 min read
OpenAI Astra scored a perfect ExploitBench result and became the first model to clear a Critical cybersecurity threshold behind select partners.
5 min read
Google DeepMind readies Gemini 3.8 Flash, code-named Skimaki, as it chases a coding gap with Claude Opus 5. We separate verified facts from second-hand reporting.
5 min read
Anthropic disclosed new fixes after its Claude models reached the live internet during security tests, a story that exposes a real control problem for agentic AI.
5 min read
Tiel Coder 35B-A3B is an MIT GGUF quantization of Ornith-1.5 by peculiar-ragdoll. Here is what the card claims and what remains unverified.
8 min read
RAG or fine-tuning for your LLM application? A data scientist walks through the cost math, latency, and governance tradeoffs that actually decide your architecture.
6 min read
A data scientist's analysis of Claude Opus 5 and GPT-5.6 Sol for coding tasks. Benchmark taxonomy, cost-per-resolution metrics, and which model fits your workflow.
7 min read
A detailed comparison of the top open-source LLMs in 2026. Benchmarks, specs, pricing, and which model fits your use case.
9 min read
Every major AI model released in 2026, from Claude Opus 5 to DeepSeek V4 to Tencent Hy4. A running reference with specs, benchmarks, and availability.
5 min read
Heavy equipment maker Caterpillar applies decades of autonomous mining experience to AI deployment, spending $100 million to retrain 118,000 workers.
4 min read
OpenAI invoked a change-of-control clause after SpaceX acquired Cursor, cutting model access by November 12. Here is what the dispute means for AI developers.
4 min read
Tencent released Hy4 preview, a 770B parameter MoE model with 1M context and Apache 2.0 license, targeting software engineering and scientific research.
5 min read
OpenAI, Google, and 100 firms warn AI cyber attacks will escalate, urging urgent collective defense for critical infrastructure.
4 min read
An in-depth look at DeepSeek's V4 model family, open-weight release, agent harness framework, surge pricing, and multimodal vision capabilities.
10 min read
Nvidia has agreed to buy open-source AI platform Hugging Face for $12.9 billion, marking a major push into the model ecosystem and cloud computing.
4 min read
OpenAI plans to show ads on ChatGPT for logged-in users on the Free and Go tiers in India, with self-serve access opening September 4.
4 min read
Barret Zoph joins Google DeepMind as research VP from OpenAI, returning to the lab where he spent six years amid a wave of high-profile departures.
5 min read
Anthropic targets a $100 billion IPO at $2 trillion valuation this October, pitching a $30 trillion AI market for the largest stock market debut in history.
4 min read
Salesforce and Anthropic unveil Claudeforce, embedding CRM data, workflows, and 37 sales skills directly into Claude for enterprise sales teams.
3 min read
Z.AI confirms the Ox Alpha AI model is its latest GLM iteration. Weights ship tonight after a stealth launch that topped coding benchmarks.
5 min read
OpenAI's custom inference ASIC outperforms Nvidia's GB300 by up to 1.9x on throughput per kilowatt, powered by HBM4 and co-designed with Broadcom.
6 min read
Thomson Reuters spent $40M over two years to build Thomson, a specialized LLM for legal work that challenges the frontier model arms race for enterprise AI.
5 min read
OpenAI instructed AI bots to solve cybersecurity puzzles, then used the results to attack Hugging Face, putting open-source AI on the defensive.
4 min read
An AI system cracked mathematical problems that stumped humans for decades, not because it is smarter but because it sees patterns we never knew existed.
3 min read
OpenAI cut GPT-5.6 Sol API prices over 20%, dropping input to $4 and output to $20 per million tokens through November, undercutting rivals Anthropic and Google.
3 min read
DeepMind alumni lab Inherent releases Faraday, a 27B scientific AI agent that outperformed Claude Opus 4.8 and GPT-5.5 at reproducing published research.
5 min read
Claude Opus 5 matches Fable 5 intelligence at half the cost with Anthropic's lowest deception rates yet, and is now the default model on Claude Max.
3 min read
Ramp spending data shows OpenAI closing the gap with Anthropic among 70,000+ U.S. businesses, as GPT-5.6 Sol wins developers back from Claude.
3 min read
DeepSeek's experimental multimodal model claims to match Claude Opus 4.8 on several benchmarks, while introducing a new Files API and multimodal pricing.
3 min read