Originally created by: imshaikot
Cursor CLI (cursor-agent, Anysphere) joins Claude Code, Codex, Antigravity, Mistral Vibe and Grok Build as a side-panel runner, in beta. Closes [#25].
Stacked on [#24], not on main: conversationDir() arrives with the Vibe fix there, and this runner needs it. Retarget to main once [#24] lands.
[#8] weighed Cursor and passed on it — "no sandbox flag, so it lands in the same host containment class as Antigravity". That was true when it was written and is not true now, which is what makes this worth doing.
cursor-agent 2026.09.18-9a7762b actually does, and how that changed the design in [#25]The installed CLI answered every open question in the issue, and three answers moved the design:
| [#25] assumed | 2026.09.18 | So |
|---|---|---|
--trust is an escape hatch, to ban beside --force |
headless refuses to start in a folder nobody trusted, and there is no env escape — the whole CURSOR_* namespace has no trust key |
--trust is required. It grants no tool permission, and the folder is one Browsentic made and wrote every file in |
whether .cursor/cli.json survives headless -p was unknown |
a run asked to echo a marker got permissionDenied, twice, and gave up |
the permission file is the containment; --sandbox enabled is a second layer we do not lean on |
| no caller-supplied session id, so scrape it | create-chat returns one — but hangs when unauthenticated |
still scrape it from system/init; not worth an extra spawn and a timeout |
| assistant events are the only text | thinking events exist, with text at the top level rather than in message.content |
a reader keying on any text field prints the model's reasoning into the panel |
| no token usage | result carries {inputTokens, outputTokens, cacheReadTokens, cacheWriteTokens} |
Cursor feeds the context card, unlike Antigravity and Vibe |
tool_call holds one tool |
it holds toolCallId, startedAtMs and hookAdditionalContexts as siblings |
the tool is the key ending in ToolCall; taking the first key was luck |
The forum's dedup rule is confirmed exactly: deltas carry timestamp_ms, and the closing flush repeats the whole message with neither timestamp_ms nor model_call_id.
cursor-agent -p --output-format stream-json --stream-partial-output
--sandbox enabled --trust [--resume <chatId>] [--model …]
-- <instruction>
Four files into a per-conversation workspace, because Cursor reads all of them off disk:
.cursor/mcp.json — the browser server, carrying BROWSENTIC_AGENT_RUN.cursor/cli.json — allow: ["Mcp(browsentic:*)"], deny Shell(*), Write(**), Read(**), and WebFetch(*) off a research run.cursor/sandbox.json — workspace_readonly, network denyAGENTS.md — the system promptThere is no per-invocation MCP flag and no way to scope which servers load, so a project config does not replace the user's global ~/.cursor/mcp.json. Their own browsentic entry would otherwise reach the browser outside a run's gate. The plan reads that file — never writes it — and denies every other server by name.
Tasks write no mcp.json at all and deny Mcp(*) outright, so a one-shot cannot reach the browser from any server.
vetPlan now reads what a workspace file says. Cursor is the first runner whose containment lives in a file rather than in argv, and Requirements.files only proved a path was written — an emptied deny list would have sailed through. fileContains closes that, and retroactively covers Antigravity, Vibe and Grok, whose TOML and JSON containment was unasserted.
A per-kind forbidden list. --approve-mcps and --auto-review must never reach Cursor, but --trust cannot join the global list because Vibe requires it. -f/--force is safe globally and is now there beside --yolo.
Ten tampering cases cover these, including an emptied deny list, a writable sandbox profile, and a task whose Mcp(*) deny was narrowed.
Fixtures are recorded from the real CLI — a turn, and the shell-denial loop — with the working directory rewritten and nothing else edited. 35 runner tests and 189 guardrail tests pass; yarn test is 1368 green.
Not yet proven, hence beta: a whole side-panel conversation — the browser tools reaching a page through the run's own MCP server, and a follow-up turn resuming. A scratch process cannot test it, because a run is only gated when the daemon itself started it.
One cost worth recording: a --trust run adds an entry under ~/.cursor/projects/<sanitized-cwd>. A conversation keeps one folder, so that is one entry per conversation, and it ages out with the workspace sweep.
The README badge row points at browsentic.com/icons/cursor.svg, which still needs adding to the site repo.
🤖 Generated with Claude Code
Ticket changed by: imshaikot