Menu ▾ ▴

#25 Cursor CLI as a sixth agent runner, and the containment call from #8 revisited

closed
nobody
2026-09-23
2026-09-22
Anonymous
No

Originally created by: imshaikot

The side panel runs on whichever agent CLI you already have logged in — Claude Code, Codex, Antigravity, with Mistral Vibe and Grok Build in beta. Anysphere ships one too: Cursor CLI (cursor-agent), with a documented headless mode. Someone whose subscription is Cursor has no way into the side panel today.

[#8] passed on this, for a reason that no longer holds

Cursor was the other candidate. It configures MCP from files too, but has no sandbox flag, so it lands in the same host containment class as Antigravity — a sealed environment and nothing else. Grok is strictly the better first move.

That was right when it was written, and Grok was the better first move. It is not right now. Current builds ship all three of:

Lever Flag or file
OS sandbox --sandbox <enabled\|disabled>, plus --allow-paths, --readonly-paths, --blocked-patterns, --network. macOS Seatbelt, Linux Landlock + seccomp
Sandbox profile .cursor/sandbox.json — type one of workspace_readwrite / workspace_readonly / insecure_none, and a networkPolicy that can default to deny
Tool permissions .cursor/cli.json — permissions.allow / permissions.deny, matchers Shell(), Read(), Write(), WebFetch(), Mcp(server:tool), and deny takes precedence over allow

That is Grok Build's containment class, not Antigravity's, and it is worth revisiting on that basis alone.

What Cursor CLI offers against the Runner contract

What a runner needs Cursor CLI offers
Newline-delimited events -p --output-format stream-json --stream-partial-output — NDJSON, every event carrying session_id
Per-run MCP server No flag. mcpServers in a project .cursor/mcp.json, merged with the user's global ~/.cursor/mcp.json, project winning
Per-run tool permissions .cursor/cli.json, as above
Appended system prompt No flag. AGENTS.md at the project root, or .cursor/rules/*.mdc (the extension must be .mdc)
Sessions session_id, first seen in the system/init event; --resume <chatId>, --continue, agent ls
Auth cursor-agent login (NO_OPEN_BROWSER=1 prints the URL), or CURSOR_API_KEY
Skills for the / picker ~/.cursor/skills, ~/.agents/skills — and it also reads ~/.claude/skills and ~/.codex/skills for interop
Efforts none

So the shape is Antigravity's and Vibe's — files written into a per-conversation cwd through Plan.files — with Grok's containment.

Where it does not fit, and what each costs

  1. The assistant stream has four undocumented forms. Under --stream-partial-output, a genuine delta carries timestamp_ms and no model_call_id; a buffered flush right before a tool call adds model_call_id; a final flush carries neither and repeats the whole message. This is confirmed by Cursor staff on the forum, not by the reference docs. A reader that naively concatenates prints the answer several times. The terminal result event is the canonical text. This is the pitfall the Codex reader already solves — track what has been said, emit only the difference.

  2. There is no way to scope which MCP servers load. No per-invocation registration flag, and no flag to use a config exclusively rather than merging with the global one. The user's own external browsentic entry in ~/.cursor/mcp.json would load beside the run's, reaching the browser outside the run's gate — the Grok problem, but with no GROK_CLAUDE_MCPS_ENABLED=false equivalent to switch it off. The containable answer is to read the user's config at plan time (read only, never write it), enumerate every server that is not ours, and emit an Mcp(<name>:*) deny rule for each, since deny beats allow. --approve-mcps auto-approves all of them and must never be passed.

  3. No caller-supplied session id. Unlike Grok's --session-id, the id is server-assigned and only revealed in system/init. Back to the scrape-and-thread pattern that Claude Code and Codex already use.

  4. No documented error event. The only error signal found is is_error: true on the terminal result. A tool-denial or mid-turn failure shape is unverified.

  5. No sandbox backend on Windows. Seatbelt and Landlock only. Containment is materially weaker there, and the docs have to say so rather than imply parity.

  6. cursor-agent and agent are the same binary — the installer symlinks both into ~/.local/bin. agent is generic enough to be shadowed by something else on PATH, so the runner should always invoke cursor-agent explicitly.

The containment posture this should ship with

localTools: 'allowlist', with the .cursor/cli.json deny rules as the lever we actually depend on, and --sandbox enabled requested as belt-and-braces that is not relied upon.

The reason for that hedge: a forum report claims headless -p runs generated a sandbox helper policy ignoring the configured sandbox.json. A Cursor engineer disputed it on the same build and asked for a repro; the thread ends unresolved. Cursor's sandbox also wraps the shell subprocess rather than the cursor-agent process itself. Until that is settled against a real binary, the permission file carries the load and the note says as much out loud.

Proposed shape

cursor-agent -p <instruction> --output-format stream-json --stream-partial-output
             --sandbox enabled [--resume <sessionId>] [--model <model>]

Four files into the conversation's own cwd under ~/.browsentic/agents/cursor/run/<conversation>/:

// .cursor/mcp.json
{ "mcpServers": { "browsentic": {
  "type": "stdio", "command": "<node>", "args": ["<cli.js>", "mcp"],
  "env": { "BROWSENTIC_AGENT_RUN": "<runId>" } } } }

:::jsonc
// .cursor/cli.json — the lever
{ "permissions": {
  "allow": ["Mcp(browsentic:*)"],
  "deny": ["Shell(*)", "Write(**)", "Read(**)", "WebFetch(*)",
           "Mcp(<each of the user's other servers>:*)"] } }

:::jsonc
// .cursor/sandbox.json
{ "type": "workspace_readonly", "networkPolicy": { "default": "deny" } }

plus AGENTS.md carrying the system prompt.

A conversation gets one folder, rewritten every turn — the lesson from 6dc97df. task mode drops the MCP server entirely and denies everything, as the other runners' json plans do.

Catalog entry: label Cursor CLI, vendor Anysphere, bin cursor-agent, install curl https://cursor.com/install -fsS | bash, keepsEnv: ['CURSOR_'], and CURSOR_ added to SECRET_PREFIX so a Cursor key is sealed away from every other agent.

Two things this exposes in the guardrails

  1. vetPlan cannot see file content. Cursor would be the first runner whose primary lever lives in a file rather than in argv, and Requirements.files only asserts that a path is present. A .cursor/cli.json whose deny list had been emptied would pass. Requirements wants an optional fileContains map, checked beside the existing files loop — which retroactively strengthens Antigravity, Vibe and Grok, whose TOML and JSON containment is unasserted today.

  2. --trust cannot go into the global FORBIDDEN list. Cursor's --trust bypasses the workspace-trust prompt and should never be passed — but FORBIDDEN is checked against every runner, and Mistral Vibe requires --trust. It needs a per-kind forbidden list on Containment rather than a global entry. -f / --force (aliased --yolo) is safe to add globally; --yolo itself is already there.

Open questions, each about one command from an answer

  1. Does --sandbox enabled actually constrain a headless -p run on the installed build? This decides whether the sandbox is a second layer or just a hopeful flag.
  2. Do Mcp(server:*) matchers glob, or do they need exact tool names? If exact, StreamContext.mcpTools already supplies the live list, as it does for Vibe.
  3. Is --resume scoped to the working directory? One source says chats are listed across workspaces. It changes nothing if the conversation keeps one stable folder, but it is worth knowing.
  4. What does a denied tool or a mid-turn failure actually emit? Only is_error on result is documented.
  5. Does Cursor announce its toolset up front? If so, Grok's available_commands tripwire should be ported, failing AGENT_UNSAFE on anything unexpected.

Not to be confused with

Driving Browsentic from Cursor over MCP already works — that is MCP clients. This is about what runs the side panel.


Sources: CLI overview · Headless · Parameters · Output format · Permissions · Sandbox · MCP · ACP

Related: [#8], which passed on Cursor for a reason that no longer holds.

Related

Tickets: #26
Tickets: #27
Tickets: #56
Tickets: #57
Tickets: #8

Discussion

  • Anonymous

    Anonymous - 2026-09-22

    Originally posted by: imshaikot

    Probed against the real binary — cursor-agent 2026.09.18-9a7762b, darwin/arm64, installed from the official script. Several claims in the issue above came from docs and are wrong or incomplete. Corrections, all from --help and direct invocation:

    Better than described

    --mode <plan|ask> and --plan are top-level flags, not an ACP-only feature.

    --mode <mode>   Start in the given execution mode. plan: read-only/planning
                    (analyze, propose plans, no edits). ask: Q&A style for
                    explanations and questions (read-only).
    --plan          Start in plan mode (shorthand for --mode=plan).
    

    This matters for containment: it is a read-only lever that lives in argv, which vetPlan can assert with an ordinary pairs entry. The issue above assumed the permission file was the only per-run lever and that the guardrails would need a new fileContains check to see it. That check is still worth adding — the MCP deny rules do live in a file — but it is no longer the only thing standing between a run and the shell.

    create-chat mints a session id up front. cursor-agent create-chat is documented as "Create a new empty chat and return its ID", so friction 3 above ("no caller-supplied session id") is wrong — the Grok --session-id pattern is available after all, at the cost of one extra spawn. Caveat found the hard way: create-chat hangs indefinitely when unauthenticated rather than failing fast, so it needs a timeout guard.

    --model takes parameterised bracket overrides, e.g. claude-opus-4-8[context=1m,effort=high,fast=false]. So Cursor does expose reasoning effort, just not as its own flag — efforts need not be empty, though it would have to be encoded into the model string rather than passed separately.

    Confirmed as described

    • -p --output-format stream-json --stream-partial-output — all present, and -p's own help states it "Has access to all tools, including write and shell", which is the risk this runner has to close.
    • -f/--force, --yolo (an explicit alias for --force), --approve-mcps and --trust all exist and all must never be passed.
    • --sandbox <enabled|disabled> exists.
    • MCP registration is file-only. cursor-agent mcp offers login, list, list-tools, enable, disable — there is no add/remove, confirming there is no per-invocation flag and no way to scope which servers load.
    • Auth strings, captured against an empty scratch HOME:
    • cursor-agent -p "hello" → Error: Authentication required. Please run 'agent login' first, or set CURSOR_API_KEY environment variable. — exit 1
    • cursor-agent status → Not logged in — exit 0, so a readiness check must match stdout, not the exit code.

    Not found, contrary to the docs

    No --acp flag and no acp subcommand appear in the top-level help of this build. ACP is out of scope for this issue, but the claim that Cursor exposes it should not be relied on without a second look.

    New, and worth a decision

    • cursor-agent bedrock — "Configure AWS Bedrock usage for CLI". That makes Cursor the second runner after Claude Code with a federated backend, so CONTAINMENT.cursor may need a federated entry keeping AWS_* when Bedrock is configured, exactly as CLAUDE_CODE_USE_BEDROCK does today.
    • -e/--endpoint and CURSOR_API_ENDPOINT (default https://api2.cursor.sh) — a redirectable API endpoint.
    • install-shell-integration writes to ~/.zshrc. Never invoke it from the daemon.
    • --workspace <path-or-name> and --add-dir control workspace roots, alongside the sandbox path flags.

    Still unverified, and blocked

    Everything past the auth gate: the real stream event shapes (including the four assistant forms), whether --mode plan still permits MCP tool calls or blocks them along with edits, whether a project .cursor/cli.json is honoured in headless -p, and whether --sandbox enabled actually constrains the run. These need a signed-in account.

     
  • Anonymous

    Anonymous - 2026-09-22

    Originally posted by: imshaikot

    Now verified against a signed-in account, same build (2026.09.18-9a7762b). Two findings overturn what I wrote above.

    --trust is required, not forbidden

    I had it in the forbidden list, alongside --force and --approve-mcps. That would have made the runner unable to start at all:

    ⚠ Workspace Trust Required
      Cursor Agent can execute code and access files in this directory.
      To proceed, you can either:
        • Run 'agent' interactively to decide
        • Pass --trust, --yolo, or -f if you trust this directory
    

    A headless -p run in an untrusted folder refuses outright. There is no environment-variable escape — I searched the bundle's whole CURSOR_* namespace, and unlike Grok's GROK_FOLDER_TRUST there is nothing.

    --trust is also not what I assumed it was. It skips the workspace trust prompt; it grants no tool permission. --force (aliased --yolo) is the dangerous one — "force allow commands unless explicitly denied". The folder being trusted is one Browsentic created, inside its own state directory, containing only files Browsentic wrote, and vetPlan already asserts the cwd is inside stateDir. This is exactly Mistral Vibe's situation, which requires --trust for the same reason.

    So it moves to required, and the per-kind forbidden list keeps --approve-mcps and --auto-review. It still cannot go on the global list, since Vibe needs it too — the per-kind list earns its place either way.

    One cost worth recording: a --trust run adds an entry under ~/.cursor/projects/<sanitized-cwd>. Because a conversation keeps one folder, that is one entry per conversation rather than per run, and it ages out with the workspace sweep.

    The deny rules hold in headless mode — measured, not assumed

    The open question above was whether .cursor/cli.json is honoured by -p. It is. A run given the real deny list and asked to echo a marker:

    {"type":"tool_call","subtype":"completed","tool_call":{"shellToolCall":{"result":{"permissionDenied":{"command":"echo PWNED_MARKER", ...
    

    Denied, twice — the model retried and was refused again — and the turn ended with "The shell blocked echo PWNED_MARKER, so nothing was printed." That settles the containment posture: the permission file carries the load, and the OS sandbox is a second layer we do not depend on.

    The stream, as it really arrives

    Recorded rather than inferred, and it differs from the documented shapes in three ways that matter:

    • thinking events exist, as {"type":"thinking","subtype":"delta","text":...} with the text at the top level rather than inside message.content. A reader that treats any text field as prose will print the model's reasoning into the panel.
    • result carries token usage, in camelCase: {"inputTokens":18119,"outputTokens":51,"cacheReadTokens":512,"cacheWriteTokens":0}. None of the documentation mentioned it, and it means Cursor can feed the context card, unlike Antigravity and Vibe.
    • tool_call is not a single-entry object. It carries toolCallId, startedAtMs and hookAdditionalContexts as siblings of the tool itself, so taking the first key is luck. The tool is the key ending in ToolCall — shellToolCall, readToolCall. A top-level call_id does exist, so that part of the plan was right.

    The dedup rule from the forum thread is confirmed exactly: deltas carry timestamp_ms, and the closing flush repeats the entire message with neither timestamp_ms nor model_call_id.

    Where it stands

    Implemented and merged into the working tree behind beta. The guardrail sweep, ten Cursor tampering cases and thirty-five runner tests pass, with fixtures recorded from the real CLI. Still unproven: a whole side-panel conversation — the browser tools reaching a page through the run's own MCP server, and a follow-up turn resuming. That needs the panel rather than a scratch process, because a run only gates browser calls when the daemon itself started it.

     
  • Anonymous

    Anonymous - 2026-09-23

    Ticket changed by: imshaikot

    • status: open --> closed
     

Log in to post a comment.