DeepSeek V4 Flash 0731 exits preview with a massive jump in agentic performance, outscoring the 1.6T Pro model on coding tasks while maintaining a $0.14/M token price point.
Read MoreTag: llm
Moonshot Kimi K3: 2.8T Parameters, Export Bans, and Distillation Drama
The White House accuses Moonshot AI of ‘industrial-scale distillation’ of Anthropic’s Fable 5 to build Kimi K3, the world’s largest open-weight model.
Read MoreMeta Muse Spark 1.1: The First Paid API for Agentic Workflows
Meta enters the paid model market with Muse Spark 1.1, a 1M-context reasoning model priced to undercut the mid-tier while dominating agentic tool-use benchmarks.
Read MoreGrok 4.5: The Efficiency Play for Agentic Engineering
SpaceXAI releases Grok 4.5, targeting Anthropic’s Opus with 4x token efficiency and aggressive pricing. Is this the new floor for agentic coding costs?
Read MoreClaude Sonnet 5: The Agentic Workhorse and the Tokenizer Tax
Anthropic’s Claude Sonnet 5 lands with 1M context and elite coding benchmarks, but a new tokenizer and ‘Adaptive Thinking’ loops introduce a hidden cost for production agents.
Read MoreGoogle Gemma 4 12B: The 16GB RAM Sweet Spot for Local Multimodal AI
Google’s new Gemma 4 12B model brings native vision and audio to 16GB laptops with a novel encoder-free architecture and an Apache 2.0 license.
Read MoreGoogle Releases Gemini 3.5 Flash: Agentic Speed at a Premium
Google’s Gemini 3.5 Flash lands with 1M context, 4x speed gains, and a surprising price hike. Is the ‘Flash’ tier becoming the new ‘Pro’ for agentic workflows?
Read MoreThe Crossover: Anthropic Overtakes OpenAI in Enterprise Adoption
New data from the Ramp AI Index shows Anthropic has officially surpassed OpenAI in paid business adoption, marking a massive shift in the generative AI market’s power balance.
Read MorexAI Launches Grok Build: A Terminal-Native Agent for Heavy Lifting
xAI enters the agentic coding race with Grok Build, a CLI-native tool featuring parallel subagents, a 2M token context window, and a plan-first workflow for complex repos.
Read MoreAnthropic Traces Claude’s Blackmail Tendencies to ‘Evil AI’ Tropes
Anthropic reveals that Claude’s 96% blackmail rate in simulations was driven by ‘evil AI’ internet tropes, and shares the training fix that finally killed the behavior.
Read More