Originally created by: imshaikot
Why Codex and Mistral Vibe navigate and reason so much worse than Claude in the side panel. Diagnosed from real runs on this machine — Codex rollouts in ~/.codex/sessions, a Vibe session log, and ~/.browsentic/daemon.log — and it is mostly integration, not the models.
A Claude run starts with no built-ins and every browser tool in its list. Neither of the other two was getting anything like that.
Codex keeps an MCP server's tools out of the model's tool list. With gpt-5.5 they sit behind tool_search; with a code-mode model (gpt-5.6-terra, tool_mode: "code_mode_only") they exist only inside exec. What the model can see is its memory and a web search, so that is what it used:
| Turn | What Codex did | Answer |
|---|---|---|
| "how many stars do the browsenric repo has ?" | web search, no browser call | 4 stars |
| "what is going on ?" | web search again | 8 stars |
| "can you go to the repo ?" | two tool_search calls, then the browser |
"Opened the Browsentic repo." |
Claude, in the same conversation minutes later, called the browser first and read 21 off the page.
Worse, the web search should not have been there at all: tools.web_search stopped existing in Codex 0.155, and an unknown config key is ignored in silence, so the switch had been doing nothing. The key is now top-level web_search="disabled", and "live" only while mapping a site.
Changed:
tool_search, how to call them from exec as tools.mcp__browsentic__* (and to pass an image on with image(result.content[0]), which text() drops), that the page is to be read rather than recalled, that a script run through exec has to start with // @exec: {"max_output_tokens": 25000}, and that a result over 25,000 tokens still comes back cut, so pages are read in pieces.web_search="disabled", or "live" for a mapping run.features.multi_agent=false is a containment requirement, not a preference: a sub-agent is spawned outside the run's gate and reports nothing to the panel.tool_output_token_limit=25000. Codex cuts a tool result at 10,000 tokens by default, and a dense page snapshot passed that in a real run at 14,847 tokens and lost its middle. 25,000 is what Claude Code allows an MCP result. An earlier draft of this PR said the key did not raise the cut. That was only true for code mode, where the script's // @exec line is needed as well.Vibe restores a resumed session from the folder it began in, not the folder it is started in now. With a folder per run, every turn after the first read the first turn's files — its BROWSENTIC_AGENT_RUN, which the daemon answers RUN_INACTIVE, and its AGENTS.md, so later turns' skill, A-Eye pick and attachments never arrived either.
In the recorded conversation that meant seven refused browser calls in a row, each shown to the model as ok: True with RUN_INACTIVE in the text, followed by an invented answer: "5,817 stars" for a repository that does not exist.
Changed: StreamContext carries the conversation a run belongs to, and Vibe names its folder after it, rewriting the config and prompt each turn. Grok already worked this way; both now say so through conversationDir().
main shows web_search present despite tools.web_search=false — the bug, reproduced. The same endpoint, with a stub MCP server returning 54,000 characters, shows the result arriving whole for gpt-5.5 and for gpt-5.6-terra with the // @exec line, and cut at 10,000 tokens without the key.config.toml and AGENTS.md between turns are picked up — which is what makes one folder per conversation the fix.yarn check: 1,307 tests pass, coverage floors hold. Vibe had no tests; it now has the ones this needed.invoked … for inactive run in daemon.log.tool_search costs Codex an extra turn often enough to be worth naming specific tools in the prompt.🤖 Generated with Claude Code
Ticket changed by: imshaikot