skip to content- 2026-08-01
|
DeepSeek V4 Flash 0731: The API Value Frontier
DeepSeek V4 Flash 0731 pairs a 1M context window and MIT weights with a $0.14/$0.28 direct API, but high token use and hallucinations still require evaluation.
- 2026-08-01
|
GPT-5.6 Luna Is OpenAI’s Value Default
Why GPT-5.6 Luna is the OpenAI value default for bounded Codex work and high-volume API tasks, with subscription and API economics kept separate.
- 2026-07-25
|
How OpenAI's Cyber Eval Breached Hugging Face
A preliminary, source-labeled analysis of the ExploitGym eval failure, Hugging Face breach, and five controls for containing cyber agents.
- 2026-07-17
|
Kimi K3: Benchmarks, Pricing, and Open-Weights Status
Kimi K3 is Moonshot's 2.8T flagship with 1M context, source-labeled benchmark evidence, $3/$15 API pricing, and released weights under the Kimi K3 License.
- 2026-07-02
|
How to Read AI Benchmarks Without Getting Fooled
Learn how to read AI benchmarks, spot saturation and contamination, compare coding and chat leaderboards, and test models on your own work.
- 2026-02-20
|
OpenCode GLM-5 Free Tier Snapshot
Historical snapshot of OpenCode's GLM-5 free-tier moment, updated to point readers toward GLM-5.1 and avoid stale leaderboard claims.