← All digests
✦ AI News for Builders

Hard budget caps are finally showing up on AWS and Google Cloud, DeepSeek Harness ships desktop apps with a breaking-changes warning, Aleph Alpha's Kolibri fits 78B on one H200, and Prime Intellect opens an OpenAI-compatible endpoint with no price list

Sunday, October 4, 2026·8 min read·4 stories

Sunday's theme is the bill. Agents make it cheap to start things and expensive to forget them, and two clouds finally ship a kill switch. Then a free agent harness you should run in a sandbox folder, an open model that does most of its thinking with 4% of its weights, and a new inference host that hasn't told anyone what it costs yet.

Story I

Simon Willison wants hard budget caps on by default. AWS and Google Cloud just shipped them, with caveats

Simon Willison's post from Saturday night makes a simple case: any pay-by-usage API should let you say "after $X a month, return errors," and that should be the default. A warning email at midnight doesn't help when an agent-built service burned through a few thousand dollars while you slept. His proposed opt-out is a checkbox that says, in plain words, that you accept responsibility for whatever comes after the cap.

The caps exist now, sort of. AWS announced project-level spend limits on September 16: when usage hits the limit the project pauses and its resources stop. The minimum is $20 or a conservative estimate of your spend, whichever is higher, and optional early controls can block new launches about 7 days out or pause top cost drivers, Bedrock and Lambda included, about 4 days out. The docs say it's rolling out to "a limited number of customers," and a paused project you ignore for 90 days gets deleted. Google Cloud's Spend Caps, in public preview since July 29, cover the Gemini API, Agent Platform, Cloud Run and Cloud Run Functions, cut usage within minutes, and delete nothing.

For builders

Today: in Google Cloud Billing, open Budgets & alerts, create a budget with type Spend Cap on every project that holds a Gemini API key. On AWS, check settings.aws.com under Billing for "Set limit"; if you have it, turn on "Pause top cost drivers" for anything an agent can deploy into. If a provider you depend on only offers alerts, move agent experiments to a prepaid key with a small balance, and put the 90-day deletion rule in your runbook before you need it.

Story II

DeepSeek Harness v0.2 ships desktop apps under MIT. The repo's own warning is in capital letters

DeepSeek released v0.2 of its open-source agent harness with official desktop apps for Apple silicon Macs and 64-bit Windows, plus the web UI you can start with npx @deepseek-ai/dsh web, which serves on 127.0.0.1:3080. The design is "everything is a plugin": model adapter, tool registry and agent loop included. New in this release are a plugin manager, a file and diff review sidebar, scheduled "Automation Task" prompts with run history, and, per MarkTechPost, an experimental compatibility layer for Claude Code mods. It takes DeepSeek models, other providers, or any OpenAI-compatible endpoint.

The download page calls it a preview. The GitHub README says "THERE WILL BE COMPATIBILITY-BREAKING CHANGES" and asks you to read SAFETY.md first. We didn't find a documented sandbox or approval mode in the material we read. That matters more this week than last, because Apple just said it will tighten Full Disk Access specifically because of agents.

For builders

If you try it, run npx @deepseek-ai/dsh web from a throwaway directory as a user without Full Disk Access, and point it at a dedicated API key with a hard cap (see above) before enabling any Automation Task. Read SAFETY.md and pin the exact version you test, since the next release can break your plugins. If you were planning to reuse Claude Code mods, treat the compatibility layer as experimental and diff behavior on one mod before porting the rest.

Story III

Aleph Alpha's Kolibri-1 is a 78.1B MoE that activates 3.46B per token and fits on one H200 in FP8

Aleph Alpha released Kolibri-1 under Apache 2.0: 78.1B total parameters, 3.46B active per token, 384 experts per layer with 6 routed. Native context is 262,144 tokens, extendable to about 1M, though the model card recommends staying at or under 262K for complex tasks. Weights ship in FP8 at roughly 78GB, so the stated minimum is one H200, B200 or B300, or two A100 80GB or H100 SXM5 cards. It's English and German, with a June 18, 2026 knowledge cutoff, and reports 75.5 English overall, 70.8 German overall and 89.3 on its English code average.

The card is unusually frank: it says the model hallucinates, may carry political bias, and is "not designed for unsupervised autonomous systems." Same week, Aleph Alpha published a study claiming Chinese open models answer sensitive political prompts in a balanced way only 17% to 41% of the time. It sells sovereign AI to governments, so read both as a pitch and a dataset.

For builders

If you serve German-language users or need a permissive license for on-prem RAG, it's a one-command test: vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 --reasoning-parser kolibri1 --tool-call-parser kolibri1 --enable-auto-tool-choice. Run your own eval set at 32K and at 200K context before trusting the 1M figure. And if you fine-tune on synthetic data from Chinese models, the study's finding of about 3,500 contaminated rows in NVIDIA's 9.3M-row Nemotron SFT set is worth a grep of your own data.

Story IV

Prime Intellect launches Prime Inference: OpenAI-compatible, GLM-5.3 first, prices not published

Prime Intellect opened Prime Inference, with serverless endpoints for spiky traffic and reserved capacity for steady load, both behind an OpenAI-compatible base URL, https://api.pinference.ai/api/v1. The first public model is GLM-5.3, which it has been serving on OpenRouter since September 22 on NVIDIA Blackwell, with Vera Rubin "coming soon." The company claims a target of 100 end-to-end tokens per second per user, says disaggregated prefill and decode cut p90 inter-token latency by nearly 40%, and reports a "near-zero" tool-call error rate.

What it hasn't done is publish per-model pricing; MarkTechPost notes it isn't in the docs yet. Every performance number above comes from Prime Intellect's own post.

For builders

Since it speaks the OpenAI API, swapping it in is a base-URL change: point an existing client at https://api.pinference.ai/api/v1 behind a feature flag and replay a day of real GLM-5.3 traffic. Log tokens per second and tool-call failures next to your current provider. Don't move production until you have a written price, and if you test, use a capped key, because "pricing not published" plus a load test is exactly how surprise bills start.

Sources

  1. Simon Willison — the case for default hard budget caps
  2. AWS Docs — spend limits: $20 minimum, early controls, pause behavior, 90-day deletion, limited rollout
  3. Google Cloud — Spend Caps preview: covered services, enforcement timing, setup
  4. GitHub — DeepSeek Harness README: install, MIT, breaking-changes warning, SAFETY.md
  5. DeepSeek — Harness download page: platforms, preview status
  6. MarkTechPost — Harness v0.2 features, Automation Task, Claude Code mods layer
  7. Hugging Face — Kolibri-1 model card: architecture, context, hardware, vLLM command, caveats
  8. MarkTechPost — Kolibri benchmarks and training token counts
  9. Aleph Alpha — Chinese model alignment study, Nemotron contamination rows
  10. Prime Intellect — Prime Inference launch: endpoint, GLM-5.3, latency claims
  11. MarkTechPost — Prime Inference: pricing not yet in docs

— The Vibe Gate news desk. We read the firehose so you can keep building.

← All digests  ·  The blog