Saturday's theme is permission. Apple is about to make the broadest permission on the Mac harder to get, and it named AI agents as the reason. Two speech stories show what you get when you narrow a model's job instead of widening it. And Meta wants its agent on a microcontroller on your desk, with a warning label it wrote itself.
No sponsored or affiliate links in this digest — the links below are sources only.
Story I
Apple will put Full Disk Access behind "very explicit user action." If your agent asks for it, plan the new flow now
Apple posted a developer notice on October 2 saying Full Disk Access exists so backup apps work, and that it "largely sidesteps" macOS privacy controls to do that. Some developers, it says, are using it in ways that expose files, mail, messages and browsing history without users understanding what they agreed to. Going forward Apple will add controls so the permission can only be granted with "very explicit user action," and it says the risk "will grow substantially" as agents get more autonomous. No date, no API details.
The backdrop is Meta's Muse. Inc. columnist Jason Aten reported that Muse surfaced a private Messages thread he says he never gave it. Meta's CTO replied that Muse needs both FDA and its Messages connector enabled. Patrick Wardle told Ars Technica that, technically, any app with FDA can read any non-root file, browser history and chats included. Apple didn't name anyone. It didn't have to.
For builders
Run grep -rn "Full Disk\|kTCCServiceSystemPolicyAllFiles" . on your Mac app and onboarding copy and list every feature that depends on FDA. For each one, write down the scoped alternative: a user-picked folder via NSOpenPanel with security-scoped bookmarks, or the Calendar and Contacts frameworks. And check your own machine: if you granted FDA to Terminal or iTerm so a CLI agent could work, every process launched from that terminal runs with it.
Story II
NVIDIA cut Nemotron 3.5 ASR's Saudi dialect error rate from 55% to 30% in 4.5 hours on two GPUs. The recipe is the story
NVIDIA's developer blog published a full walkthrough of adapting Nemotron 3.5 ASR, a 0.6B streaming model, to Najdi and Hijazi Arabic using SADA, the dataset released by Saudi Arabia's SDAIA. After light curation (103,559 of 125,490 clips kept, 133.7 hours) and 12,000 steps on two RTX PRO 6000 Blackwell cards, WER on the target split went from 55.05% to 29.96%. English on FLEURS improved too, from 11.04% to 10.42%, because they replayed real data: 90% SADA, 7% FLEURS English, 3% FLEURS Arabic, declared as weights rather than concatenated.
The details are the useful part. Setting num_buckets without use_bucketing=True is a silent no-op. Updating all 24 encoder layers beat freezing the bottom 16 by 2.4 points at this data volume. And switching decode to a [56, 13] attention context with beam-8 MALSD cut another 2.71 points with no retraining, at roughly 800ms of extra latency. That is fine for call archives and wrong for live captions.
For builders
If you ship transcription for any accent your vendor calls "supported," pull 30 minutes of your own audio and measure WER before you trust the label. Then open NVIDIA's notebook in nvidia-riva/tutorials and copy two things even if you never fine-tune: the weighted replay mix as a guard against forgetting, and the decode-config sweep, which is free accuracy for batch jobs.
Story III
Microsoft's MAI-Transcribe-2-Streaming is first of 38 on streaming accuracy. It also costs three times what Meta charges
Microsoft AI released MAI-Transcribe-2-Streaming on October 1, its first streaming speech-to-text model. On Artificial Analysis's streaming index it posts 2.5% WER with the final transcript 0.13 seconds after end of speech, first of 38 models, and its first partials score the same 2.5% at 0.12 seconds. That second number matters more than it looks: an agent can start reasoning or calling tools mid-sentence without acting on a guess that later changes. It covers 60 languages with continuous language detection.
The price is $0.54 per audio hour, introductory through the end of 2026. MarkTechPost lists Grok Voice Transcribe 2.0 at $0.20 and Muse Voice Transcribe at $0.18, at 2.7% and 3.1% WER. It ships as a public preview with no SLA and no open weights, through a Realtime-compatible WebSocket API, the Azure Speech SDK, and Vercel.
For builders
Do the math on your volume before switching: 1,000 streaming hours a month is $540 here versus $180 on Muse. If your voice agent already speaks an OpenAI Realtime-compatible WebSocket, point a 5% traffic slice at MAI for a week and log barge-in errors and turn latency, not just WER. And don't put a no-SLA preview on a path that pages you at night.
Story IV
Meta open-sourced an SDK for building your own Muse hardware on ESP32 or a Raspberry Pi. Its own advice: "Proceed at your own risk"
Meta released Muse Gadgets on October 2: firmware for ESP32 boards and a Linux SDK, Apache 2.0, in facebookincubator/muse-gadget-sdk, for wiring its Muse agent to displays, buttons, sensors and actuators. Suggested builds include a color E Ink reminder board and an HDMI stick that puts Muse on a TV. Meta also made 5,000 units of its own Muse Home Link, a USB-C device that lets Muse control anything on your network with an HTTPS interface. The Decoder says they ship within weeks, free for Muse subscribers while supplies last.
Read this next to story one. Eleven days before Apple's notice, Wardle disclosed a Muse configuration that let any code running on a Mac take control of the assistant, per Ars Technica. An agent with network reach into your home is the same bet with a different blast radius.
For builders
If you build one, put the board on a guest network or its own VLAN, and expose only the specific HTTPS endpoints it needs, with a token, behind your router. Don't hand it a Home Assistant long-lived admin token. A scoped user with access to two devices is enough for a demo, and it's what you'll wish you had done when the next disclosure lands.
— The Vibe Gate news desk. We read the firehose so you can keep building.