← All digests
✦ AI News for Builders

Claude Haiku 5.5 drops to $0.10 per million input tokens with a tokenizer catch, OpenAI's Decisions API and Liquid's open d1 models answer without generating text, a likely lone attacker used an AI pentest tool on Korean banks, and Windows gets agent sandboxes as Copilot reaches into your files

Thursday, October 8, 2026·8 min read·4 stories

Thursday is about the bill and the blast radius. Small models got cheap enough that the tokenizer now matters more than the sticker price, and a new class of models skips text output entirely. Meanwhile the same automation that makes your agents cheaper made one person's bank heist faster, and Microsoft is shipping a sandbox for agents in the same week it hands Copilot your file system.

Story I

Claude Haiku 5.5 is $0.10 per million input tokens. Measure your tokens before you celebrate

Anthropic shipped Claude Haiku 5.5 on Wednesday as claude-haiku-5-5. Up to 100K tokens of prompt it costs $0.10 input and $0.50 output per million, with cache reads at $0.01. Past 100K everything is five times higher: $0.50 and $2.50. Haiku 4.5 was $1 and $5. It is also the first Haiku with effort levels (low through max), and Simon Willison found you can't switch reasoning off; it defaults to medium. The Decoder reports a 1M-token context window, up from 200K, and Anthropic's own pitch is narrow work: summaries, compaction, classification, subagents. For complex agentic coding it still points you at Sonnet 5.5 or Opus 5.5.

The catch is the new tokenizer. Simon's counter shows the same long prompt using about 1.25x as many tokens as on Haiku 4.5, which he calls a hidden price increase, and The Decoder notes Artificial Analysis measured about 162,000 output tokens per task at max effort versus about 50,000 for GPT-6 Luna. Two side notes worth more than the headline for some of you: Sonnet 5.5 cache reads were halved to $0.10 per million, and Max and Team subscribers now get monthly API credits ($100 for Max 5x, $200 for Max 20x, up to $500 pooled for Team) that don't roll over.

For builders

Don't port by multiplying old bills by 0.1. Replay a day of real Haiku 4.5 traffic through Haiku 5.5 at low and medium effort, log input and output token counts per call, and compare cost per task, not cost per token. Keep prompts under 100K or the 5x tier eats the savings, and pin effort explicitly so a default change can't triple your output tokens. If you're on Max or Team, link an API org under Settings → Billing and turn off auto-reload: requests stop when the credit runs out instead of hitting your card.

Story II

OpenAI's Decisions API and Liquid's open d1 models return a verdict, not a paragraph

OpenAI put a Decisions API in public beta. You send text, images or both, and get back yes/no probabilities, a pick from categories you define, or a score on a scale. Per The Decoder it's about ten times faster than the Responses API, supports only gpt-6-luna for now, and costs $0.10 per million input tokens with output tokens free. It supports zero data retention and HIPAA use in the US and Europe. OpenAI also cut its paid API tiers from five to three (Build, Launch, Grow) with monthly caps of $500, $5,000 and $200,000.

The same day Liquid AI open-sourced two decision models that skip generation entirely and return calibrated probabilities in one forward pass. MarkTechPost lists d1-3B at 3.12B parameters running in 8 ms on an RTX 4090 and 30 ms on an Apple M5 Pro, plus d1-omni-600M, which takes images and audio. The license is LFM Open License v1.0, free for commercial use under $10 million in annual revenue. Neither vendor has published a head-to-head, so take each one's accuracy numbers on its own benchmark with salt.

For builders

Grep your codebase for prompts that end in "answer only yes or no" or "respond with one of". Those are your candidates: moderation flags, ticket routing, lead qualification, eval judges. Move one to the Decisions API and log the returned probability, then pick your threshold from a labeled sample of 200 instead of trusting a parsed string. If the data can't leave your box, or you're under the $10M line, run d1-3B locally against the same 200 and compare.

Story III

CrowdStrike: a likely lone attacker ran an open-source AI pentest tool against Korean banks, with Claude Code logs left in the open

CrowdStrike traced a string of breaches at South Korean financial institutions between late September and early October to what it believes is a single Chinese-speaking actor. The tooling was ARTEX, an open-source automated penetration testing tool posted to GitHub in July that drives language models to find flaws on its own; DeepSeek v4.1-flash, GLM-5.3 and Grok 4.6 are named. Researchers also found Claude Code session logs in the attacker's open directories, showing searches for Telegram groups to sell the data. At Shinhan Bank alone, more than 25,000 records with names, contacts, income and credit limits were taken, per Korean paper Khan via The Decoder.

The Decoder doesn't say how the attacker got in, so don't read this as a specific CVE story. Read it as a staffing story. This week Anthropic also widened its Cyber Verification Program, which gives vetted defenders models with fewer restrictions for vulnerability research, malware analysis and incident response; open-source maintainers and individual researchers can apply for its Defense Access tier.

For builders

Assume someone is already pointing an autonomous scanner at your public surface, because it now costs one person an afternoon. This week: list every internet-facing host and admin panel you own, kill the ones nobody uses, put the rest behind SSO, and run an automated scan on yourself before someone else does. And if your own agents keep session logs, check where they're written; this attacker's logs ended up in a public directory, and yours would tell a stranger just as much.

Story IV

Windows gets Execution Containers for sandboxing agents, the same week Copilot learns to rename and zip your files

At Microsoft's Wednesday event, Satya Nadella said a Windows 11 feature called Execution Containers, built to make it easier to sandbox AI agents, will come to all Windows 11 users, per TechCrunch. Microsoft didn't give technical details in that coverage, and we haven't seen docs yet. The same event launched the Surface RTX Spark Dev Box at $6,000 with VS Code, GitHub Copilot CLI, WSL and PowerShell 7 preinstalled.

The Verge covered the other half. Copilot features Microsoft calls Hybrid Intelligence will use local files on your PC to take actions over the "next couple months": the demo searched folders for documents, renamed them, zipped them and drafted an email with the zip attached. A new Windows search lets you flip dark mode or send a text from the search bar this fall. Neither article spelled out permissions or opt-in controls.

For builders

If you ship a desktop agent or a CLI that runs model-written commands on Windows, watch for the Execution Containers docs and plan to run tool calls inside one instead of the user's session. Until then, do it the boring way: run agent shells under a separate low-privilege user or in WSL with only the project folder mounted. And if your app stores secrets in plain files in the user's Documents folder, move them to the Windows Credential Manager now, before an assistant that can search, zip and email those folders ships to your users.

Sources

  1. Anthropic — Claude Haiku 5.5 announcement (pricing tiers, model ID, effort levels, subscriber API credits)
  2. Simon Willison — Claude Haiku 5.5 (tokenizer ~1.25x, no reasoning-off, pelican costs, credits setup)
  3. The Decoder — Haiku 5.5 price cuts, 1M context, Artificial Analysis token counts, Sonnet 5.5 cache cut
  4. The Decoder — OpenAI Decisions API (pricing, gpt-6-luna, ~10x speed, API tier changes)
  5. MarkTechPost — Liquid AI d1-3B and d1-omni-600M (sizes, latency, license)
  6. The Decoder — CrowdStrike on the Korean bank breaches (ARTEX, models named, Shinhan records)
  7. The Decoder — Anthropic expands its Cyber Verification Program
  8. TechCrunch — Microsoft's AI PCs, Execution Containers, Surface RTX Spark Dev Box
  9. The Verge — Copilot Hybrid Intelligence gets local file actions

— The Vibe Gate news desk. We read the firehose so you can keep building.

← All digests  ·  The blog