skip to content- 2026-08-11
|
Agents Running Agents: Codex Meets OpenCode and DeepSeek
A hello-world DeepSeek request became a safer agent-to-agent runner, a dated route preflight, and a Cost per Accepted Result protocol.
- 2026-07-26
|
Prompt Caching: Cut AI Agent Token Costs
Learn how prompt caching changes LLM token costs, why coding agents resend context, what compaction can waste, and how to measure Codex, Claude Code, and OpenCode.
- 2026-07-25
|
Claude Opus 5: Cost, Benchmarks, and When to Use It
Claude Opus 5 pricing, availability, effort levels, independent cost-per-task evidence, migration changes, and practical routing against Sonnet 5 and Fable.
- 2026-07-25
|
How OpenAI's Cyber Eval Breached Hugging Face
A preliminary, source-labeled analysis of the ExploitGym eval failure, Hugging Face breach, and five controls for containing cyber agents.
- 2026-07-17
|
Kimi K3: Benchmarks, Pricing, and Open-Weights Status
Kimi K3 is Moonshot's 2.8T flagship with 1M context, source-labeled benchmark evidence, $3/$15 API pricing, and released weights under the Kimi K3 License.
- 2026-07-02
|
Fable 5 Returns as Claude Sonnet 5 Becomes Default
Fable 5 is back with tighter safeguards while Sonnet 5 becomes Claude's default. Here is the practical routing and cost decision.
- 2026-07-01
|
AI Agents
Navigate the AI agent landscape: autonomous platforms that run independently, coding assistants that amplify your workflow, and the models that power them all.
- 2026-07-01
|
Claude Sonnet 5 Guide
Claude Sonnet 5 pricing, 1M context, API migration changes, tokenizer cost implications, availability, and source-labeled launch evidence.
- 2026-06-28
|
GPT-5.6 Sol, Terra, and Luna Guide
Current GPT-5.6 Sol, Terra, and Luna access, short- and long-context pricing, independent benchmark signals, and deployment controls.
- 2026-06-28
|
GPT-5.6 Sol Preview: Access and Agent Risks
GPT-5.6 Sol is a restricted API and Codex preview shaped by a government request. Check access, agent risks, pricing, and safeguards before routing work.
- 2026-06-20
|
GLM-5.2: July Value Coding Pick
GLM-5.2 is the July best value coding model to test: 1M context, AA Index 51, $1.40/$4.40 API pricing, and Opus 4.8 comparison math.
- 2026-06-20
|
Kimi K2.7 Code: Coding Model Guide
Kimi K2.7 Code is Moonshot's cheaper routine coding lane after K3: 256K context, base/HighSpeed pricing, source-labeled benchmark deltas, and eval caveats.
- 2026-06-05
|
LLM App-Hacking Field Test: $1,500 Takeaways
A cautious read of Kasra Rahjerdi's informal $1,500 field test on whether LLM agents could exploit a deliberately vulnerable app.
- 2026-06-01
|
MiniMax M3: Value Coding Model Guide
MiniMax M3 is a value coding model candidate with 1M context, multimodal input, Token Plan economics, and vendor-reported benchmark strength. Test it before replacing premium lanes.
- 2026-05-18
|
GLM-5.1: Prior Low-Cost Coding Model Guide
GLM-5.1 is Z.AI's prior low-cost coding model context. Use GLM-5.2 for the current 1M-context GLM coding-model guide.
- 2026-05-18
|
Z.AI GLM Coding Plan Guide
Z.AI's GLM Coding Plan gives supported coding tools a GLM-5.2 lane to test. New credit metering, account-relative legacy migration support, referral disclosure, and caveats.
- 2026-03-02
|
OpenClaw Provider Policy Check 2026: Complete Guide
Find out if your AI provider allows OpenClaw usage. 10-second status check for Anthropic, Google, OpenAI, Kimi, and 10+ providers with migration paths if you're at risk.
- 2026-02-19
|
Anthropic OAuth Policy Feb 2026: What Changed
Anthropic's official Claude Code compliance docs explicitly prohibit OAuth tokens in third-party tools, including the Agent SDK. Here's what's new, why OpenCode broke, and where the OSS ecosystem goes next.
- 2026-02-17
|
UltraThink → UltraQuiet: Why Devs Want Receipts
Claude Code started hiding file paths and search details behind Ctrl+O. Here's a working fix: the AGENTS.md Audit Ledger pattern.
- 2026-02-15
|
Kimi Claw: Managed OpenClaw Guide
Current Kimi Claw guide covering managed OpenClaw deployment, K2.6 Thinking, membership requirements, credits, and security tradeoffs.