Calling a frontier model got cheaper again this week, and one model right now costs nothing at all. None of that tells you the part that actually binds you — who operates the endpoint, what the licence permits, and which privileges the thing runs with when nobody is watching. Four stories about the distance between a sticker price and the terms.
No sponsored or affiliate links in this digest — the links below are sources only.
Story I
OpenAI cut GPT-5.6 Sol by 20% on input and a third on output
On August 21 OpenAI dropped Sol from $5 to $4 per million input tokens and from $30 to $20 per million output. That asymmetry is the story. A 20% input cut is a line item; a 33% output cut changes which workloads are worth running at the top of the ladder at all. Cached input sits at $0.40 and cache writes at $5.00, and the full family now reads Luna at $0.20/$1.20, Terra at $2.00/$12.00, Sol at $4.00/$20.00.
The reduction applies to the pay-as-you-go API, Codex credits, and eligible ChatGPT Work plans. Consumer Pro, Plus, and Business subscriptions do not change. It is the second Sol cut inside a month, after Luna and Terra were reduced in late July, and coverage frames it as undercutting Claude Opus 5 on both sides of the meter.
One word deserves suspicion: OpenAI calls this promotional, and commits only to holding it through at least November 21, 2026. "At least" is not a floor anyone is obligated to keep. Put this beside the last four days — DeepSeek raising prices and inventing peak hours on Thursday, Anthropic quietly cancelling the Sonnet 5 increase on Friday, Gemini 3.7 Flash arriving cheap-with-an-expiry on Saturday — and the pattern is not "AI is getting cheaper." It is that inference pricing has become a thing vendors move without notice, in both directions, on their own calendars.
For builders
Pull your last thirty days of usage and split it into input tokens and output tokens before you decide anything. Most people carry a rough cost intuition built when output was six times input; at $4/$20 the ratio is five, and if your workload is output-heavy — code generation, long-form drafting, verbose tool responses — your real per-task cost just fell by more than the 20% headline. While you are in there, check your cache hit rate: cached input at $0.40 is a tenth of fresh input, and a system prompt over a few thousand tokens that is missing the cache is a bigger leak than the price cut is a win. Then set a calendar reminder for November 14. Promotional rates that lapse quietly are how a Q4 forecast breaks in December.
Story II
A model with no company name on it went free on OpenRouter, and the window shuts Thursday
OX Alpha appeared on OpenRouter on August 20 with no vendor attached: a 1,048,576-token context window, 131,072 tokens of maximum output, text, image and video input, tool calling, and a price of zero until August 27. After that, unstated. Independent analysis puts it at roughly 744 billion parameters total with about 40 billion active in a mixture-of-experts layout.
The viral number is 80% Pass@1 on DeepSWE, against 65% for Claude and 52% for GPT-5.6. Treat it as a rumour, not a result. It came from a ten-task user test rather than an audited leaderboard, and DeepSWE is not the SWE-bench Verified harness where frontier labs quote 96%. Those two figures cannot be placed on the same axis, and plenty of write-ups this weekend did it anyway. Researcher Ben Davis says he is 99% certain the model belongs to Zhipu's unreleased GLM-5.x line, reasoning from video-encoder token consumption that matches GLM-5V-Turbo and tokenizer alignment with GLM-5.3. Nobody official has confirmed it, and we found no preview-specific terms of service or data policy published anywhere.
For builders
One thing to do and one thing not to do, and the deadline on both is Thursday. Do: run your own evaluation while it is free, because a no-cost million-token window is a rare chance to find out whether your long-context prompt actually degrades past 200K or whether you have just been assuming it holds. Ten of your real tasks, scored by your own rubric, beats any leaderboard. Do not: send it proprietary code, customer records, or anything under an NDA. A stealth model shipped without a company name is collecting its traffic — that is close to the whole point of launching one anonymously — and there is no named counterparty you could hold to a data agreement afterwards. Evaluate it on a public repo and synthetic records, and keep the interesting prompts for a provider that signed something.
Correction (August 27, 2026): Two figures in this story were wrong. OX Alpha was revealed on August 26 as Z.ai's GLM-5.3-Flash, a 320B-total, 18B-active mixture-of-experts model — not the roughly 744B total / 40B active that independent analysis estimated at the time and that we relayed here. The free window also closed on August 26, a day earlier than the August 27 date stated above, and the stealth slug was removed from OpenRouter with no documented alias to z-ai/glm-5.3-flash. Full details in the August 27 digest.
Story III
Somebody read all thirty open-weight licences. Eleven of them have conditions.
An audit published on August 16 went through 30 open-weight models from 17 organisations and read the actual licence files. Seventeen were genuinely permissive under Apache-2.0 or MIT. Eleven carried commercial conditions. Two had no public repository to check at all. That last category is its own kind of answer.
The conditions are more varied than the usual "you can't use it if you're Google" clause. Moonshot's Kimi K2.5 and K3 require "Kimi K2.5" branding in your interface once monthly revenue reaches $20M. MiniMax requires written authorisation above $20M annual revenue, and MiniMax-Music3 adds a nineteen-category acceptable-use prohibition list on top of a branding requirement. Meta's Llama 4 Scout and Maverick need manual approval past 700 million monthly active users. Lightricks' LTX-2.3 and Liquid AI's LFM2.5-8B require a paid licence above $10M annually. Tencent's Hy3 preview excluded the EU, UK and South Korea before the final Apache-2.0 release dropped the carve-out. And FLUX.2's non-commercial models state that outputs are not considered Derivatives — meaning you may commercialise the images even though the weights are restricted, which is the opposite of what most people assume.
The detail worth remembering is the smallest one. Black Forest Labs shipped FLUX.2-klein-4B under Apache-2.0 and FLUX.2-klein-9B under a gated non-commercial licence eighteen minutes apart on January 14, 2026. Same family, same launch, opposite terms.
For builders
List every open-weight model in your stack and open the LICENSE file in the repository you actually pulled from — not the family's model card, not the announcement blog, not the aggregator page. Eighteen minutes is how long a sibling model can take to arrive under contradictory terms, and no marketing page will tell you which one you downloaded. Then sort what you find by trigger type rather than by fear: revenue thresholds and MAU caps at $10M or 700 million are irrelevant to almost everyone reading this, but branding clauses fire at the same thresholds and are the only category you cannot retrofit cheaply, because they require changing a shipped interface. If you are near any of those numbers, that is a lawyer's afternoon, not a grep. If you are not, write the trigger down next to the model name and move on.
Story IV
Snowflake's scheduled agents run as every role you hold, and the permission is granted to PUBLIC
On August 21 Snowflake put CoCo automations into preview in the CoCo CLI and in Snowsight. A prompt becomes a recurring unattended run inside a Snowflake-managed sandbox, executing on schedule with your terminal and browser closed. Minimum frequency is one hour, schedules are fixed times only with no event triggers, run history is retained for two months, and it is available in commercial regions on AWS, Azure and Google Cloud but not in government or FedRAMP deployments. During preview each run bills standard Snowflake task charges on top of CoCo token consumption.
The part to read twice is in Snowflake's own documentation, not in anyone's threat report. The active role you had when you created the automation is not the role it runs with. Each run opens a task session whose primary role is your user's default role, with your user's default secondary roles activated. An analysis published on August 22 connects that to two other defaults: EXECUTE AGENT TASK is granted to PUBLIC, so any user in the account can create one of these, and behaviour change bundle 2024_08 set DEFAULT_SECONDARY_ROLES to ('ALL') for most users. Stack the three together and a scheduled run can reach any object its creator can reach through any role they have ever been granted.
Nothing here is a vulnerability, and Snowflake documents the model plainly — it is caller's rights, and RBAC, row access policies and masking policies all still apply. That is exactly what makes it easy to miss. The gap is between the privileges you were thinking about when you wrote the prompt and the union of privileges the prompt will actually execute with at 3am.
For builders
Three statements, five minutes. Run SHOW GRANTS TO ROLE PUBLIC; and look for EXECUTE AGENT TASK. If it is there and you did not put it there, run REVOKE EXECUTE AGENT TASK ON ACCOUNT FROM ROLE PUBLIC; and grant it back only to the teams that need it. Then run SELECT NAME, DEFAULT_ROLE, DEFAULT_SECONDARY_ROLE FROM SNOWFLAKE.ACCOUNT_USAGE.USERS; and set DEFAULT_SECONDARY_ROLES = () for every service account and contractor in the list. The general rule outlives Snowflake, and it pairs with Sunday's AgentCore payments story: an agent that runs on a schedule inherits its authority from an identity, and the authority it inherits is almost never the narrow slice you had in mind. Before you schedule anything unattended, write down which identity it runs as and what that identity can reach at its widest.
— The Vibe Gate news desk. We read the firehose so you can keep building.