Originally created by: imshaikot
Grok Build (grok, xAI) joins Claude Code, Codex and Antigravity as a side-panel runner, in beta. It is the first runner contained in both directions at once: a per-run tool list and a kernel sandbox. Closes [#8].
Branched from main, independent of [#10] (Mistral Vibe). Neither of Vibe's hooks is needed here: Grok's stream always closes with an end event, and its permission rules take globs.
grok 1.0.40 actually does, and how that changed the design in [#8]The installed CLI, its bundled docs and grok inspect answered every open question in the issue, and several answers moved the design:
| [#8] assumed | 1.0.40 | So |
|---|---|---|
headless needs --always-approve, so pair it with a sandbox |
--permission-mode dontAsk runs only what was allowed up front |
never --always-approve (now in FORBIDDEN for every runner); one allow rule, MCPTool(browsentic__*) |
containment lives in .grok/config.toml |
--tools, --allow/--deny, --rules, --no-subagents, --sandbox are all flags |
only the MCP server is a file; everything else is argv vetPlan can see |
--session-id names a resumable session |
-s only creates a session; --resume continues one |
-s <uuid> on the first turn, --resume after, in a folder named after the uuid |
| project config might override user config | [permission] rules merge across every source, including ~/.claude/settings*.json; Grok also loads Claude Code's and Cursor's MCP servers by default |
containment is deny rules and --tools, never allow rules; GROK_CLAUDE_MCPS_ENABLED=false / GROK_CURSOR_MCPS_ENABLED=false drop those servers |
a custom sandbox profile with **/.env deny globs |
relative globs are anchored at the workspace; read-only blocks child network on Linux, and the MCP server is a child |
built-in workspace for runs, read-only for tasks; reads are closed by --tools and --deny Read |
Two more things the docs do not say:
--tools "" switches on all 27 tools, shell and write included. A run that needs no built-in passes todo_write, and vetPlan refuses a missing or empty list.GROK_FOLDER_TRUST=0 for their own process rather than --trust, which would write every conversation folder into the user's ~/.grok/trusted_folders.toml.GROK_MEMORY=0 as well, so a page cannot leave anything in the user's next Grok session.
grok -p <instruction> --output-format streaming-json
--permission-mode dontAsk --allow 'MCPTool(browsentic__*)'
--tools todo_write|web_search,web_fetch
--deny Bash --deny Edit --deny Write --deny Read
--no-subagents --sandbox workspace --rules <system prompt>
-s <uuid> | --resume <uuid> [-m …] [--reasoning-effort …]
Tasks write no files: --deny MCPTool refuses every MCP call from any server, --sandbox read-only.
Grok reaches MCP tools only through search_tool/use_tool as browsentic__<tool>, so the rules carry one sentence mapping the skills' bare names. The reader skips search_tool and Browsentic calls (the daemon already draws those), reports use_tool to any other server under that server's tool name, and reports usage per model response with the same arithmetic as the Claude reader.
Tripwire. Grok lists its toolset before the model is called. A run offered anything beyond search_tool, use_tool, todo_write and the web tools fails AGENT_UNSAFE before the model sees it. vetPlan proves what we asked for; this proves what the CLI did.
spawn.ts learns three things, none of which changes the existing runners' results:
--flag value pair must hold for every occurrence of the flag (a CLI takes the last one, so a later --sandbox off was invisible);allows requirement: the allowlist flag present, non-empty, comma lists split, nothing beyond the runner's set;env requirement, for switches a CLI only takes from its environment.FORBIDDEN gains --always-approve and --sandbox=off. XAI_/GROK_ are Grok's kept prefixes and sealed away from everyone else.
When a reader failed a run, the process kept going: its MCP calls got RUN_INACTIVE, but it went on spending tokens and holding its tools. drive.ts now stops it. The tripwire needs that; the other runners get it too.
launch()vetPlan;--sandbox workspace and read-only, with the run id in its environment;~/.grok/trusted_folders.toml, ~/.grok/config.toml and ~/.grok/memory-v2 are untouched afterwards;Not verified: anything past the model call. The account this was built on is on the Free tier, and every request was rate-limited; Grok retries a 429 quietly for about six minutes before giving up with a 500. So use_tool approval under dontAsk, the stream shapes after available_commands, and --resume are unproven, and the stream fixtures for them are hand-written from Grok's own headless docs. That is why the docs call it beta.
yarn check passes (1,260 tests), both extension builds pass, the two Swift edits parse. src/daemon/test/browsers.test.ts › a drifted build is served what it reports fails about 4 runs in 6 on clean origin/main too, so it is not from here.
XAI_API_KEY) Grok account: the popup shows ready, an instruction drives the page, a follow-up resumes, a research run shows web_search, a title still generates.fixtures/grok/*.hand-written.jsonl from that run.AGENT_KINDS, RUNNERS, CONTAINMENT, the docs tables, and two copies of the same allows rule in spawn.ts).🤖 Generated with Claude Code
Ticket changed by: imshaikot