| Name | Modified | Size | Downloads / Week |
|---|---|---|---|
| Parent folder | |||
| e2e@0.17.0 source code.tar.gz | 2026-10-04 | 3.9 MB | |
| e2e@0.17.0 source code.zip | 2026-10-04 | 4.6 MB | |
| README.md | 2026-10-04 | 15.1 kB | |
| Totals: 3 Items | 8.5 MB | 0 | |
Minor Changes
-
#779
6639bfaThanks @fecolinhares! -copilot()reaches the models Copilot serves only over its Responses API, such asgpt-6-luna,gpt-5.3-codex, andgrok-4.5. On the first call it reads the plan's model listing and picks chat completions or the Responses API from the model'ssupported_endpoints; a model served over chat completions, or one the listing does not place, stays on chat completions. A listing that cannot be read sends that call over chat completions without remembering it, and a revoked login fails on the listing with the sign-in command. Responses models run through@ai-sdk/openai, whiche2e initnow adds for Copilot; without it the call names the package to install. Responses turns carry the same initiator and vision headers as chat turns.e2e models github-copilotmarks whatcopilot()cannot call withno chat or responsesin place ofno chat completions, and a disabled model asnot enabledonly. -
#794
c66176bThanks @neriousy! - Sign in with OpenCode Console to use OpenCode Zen and OpenCode Go models for agent steps.e2e login opencode-consoleruns Console's device flow: approve the code, pick the workspace, and the login is stored and refreshed like the other subscriptions.opencodeConsole('<id>')frome2e/oauth/opencode-consolecalls a Zen model for a bare id and a Go model for ago/id (go/deepseek-v4.1-flash); on its first call it reads the workspace config to pick the model's API (OpenAI chat completions, OpenAI Responses, Anthropic Messages, or Google) and loads only that AI SDK package, so@ai-sdk/anthropicand@ai-sdk/googlejoin the optional peers. Once the workspace config is read, a model it does not serve, or a Go model without the subscription, fails withMISCONFIGUREDnaming the fix. Each model sends onex-opencode-session, so a worker's calls stay on one upstream cache, and prompts cache the way e2e caches them for OpenAI and Anthropic directly: models served over Anthropic Messages get a breakpoint on the system prompt and, in agent steps, on the newest message, and Responses models get the system prompt's cache key. Reports name the modelopencode/<id>oropencode.go/<id>, and chat models of both plans read provider options underopencode.e2e models opencode-consolelists the enabled models tagged Zen or Go,e2e initoffers OpenCode Console among the subscriptions, andOPENCODE_API_KEY, the variable OpenCode reads, takes a Console service account key in place of the stored login. -
#741
bfa58c7Thanks @Marve10s! - Keep soft assertion failures in the report when a test skips itself, and show them in the CLI underSkipped After Failure. A failed attempt followed by a skip on retry also appears there with its original error.
Add failOnSkippedFailure: true to fail the run with exit code 1 for these
tests. The default is false, preserving the existing exit code. Tests retain
their skipped status and reason, teardown still runs, and a skip does not
trigger another retry.
- #713
7ab80bcThanks @okwasniewski! - Breaking: an interrupted test is no longer counted as failed.report.jsongainsrun.summary.interrupted, andrun.summary.failedcounts only failed and timed-out tests.run.summary.skippednow counts only selected tests, sopassed + failed + interrupted + flaky + skippedequalsselected; the tests a filter left out arediscovered - selected. Thelistreporter prints3 interruptedin its own column,summary.mdshows them with ⏹️ and gives them no failure block or page, andjunit.xmlwrites each one as a<skipped>whose message startsinterrupted:. A test that failed and whose retry an interrupt cut short stays failed, and its failure is the one reported. A--repeat-eachrun the interrupt stopped is listed as interrupted, not as a flake. The telemetry event gainstests_interrupted. In@e2e-dev/github, a test that was interrupted and then passed on a--last-failedrerun shows as passed, not flaky.
Patch Changes
-
#711
a3da00dThanks @okwasniewski! -device.installApp()no longer pins the build's file path as the app when agent-device reports no bundle id or package for it, which on Android madeapp.open()fail with "Android runtime hints require an installed package name". It fails at the install withENGINE_FAILUREnamingapp.bundleIdor installApp'sappoption; the mobile docs' Troubleshooting covers the Android cause (agent-device 0.21.18 reads an aapt2-built APK's package only through the SDK'saapt, which it does not find in the macOS default SDK location, or when the install added the package).app.open()on a device now keeps the engine's own reason it cannot launch, such as a build not installed yet withdevice.installApp(), instead of a generic "pin one with app.bundleId or app.appPath". -
#732
3bf295eThanks @okwasniewski! - The replay cache no longer passes a step whose effect did not happen. Entries are keyed by the agent and a hash of its redacted context, so one persona never replays another's steps and a changedagentContextrecords again. A route now includes the origin and the query (ids, tokens, and timestamps aside), and the end route must match exactly. A replay checks what the step made appear, with checked or selected state, and what it made disappear, needs at least one of those changes to happen during the replay, and stops on an alert the recording never saw. A count the step made appear is part of the effect, and an unnamed control with twins needs a named row or group to replay. Every existing entry is re-recorded on the first run after the upgrade. -
#713
7ab80bcThanks @okwasniewski! - A forced interrupt (a second Ctrl-C, or a CI cancel that sends SIGINT and then SIGTERM) still writesjunit.xmlandsummary.md. The runner used to abandon them along with every other reporter, which left the previous run's files beside the newreport.json. Custom reporters are still abandoned on force. -
#799
96ff11bThanks @okwasniewski! - A recorded step that removes a control while something else on screen keeps its text, such as a radio labelled "Express" replaced by a status reading "Express", now replays. Its recording used to fail its own end check on every replay (end-mismatch, orREPLAY_STALEunder--strict-cache). Re-record such a step once to pick up the fix. -
#726
f7c0756Thanks @okwasniewski! -locator.waitForadds Playwright'sattachedanddetachedstates and refuses any other state, or an option key besidesstateandtimeout, withINVALID_ARGUMENT. Before, every state butvisibleran ashidden, sowaitFor({ state: 'attached' })and a misspelled state passed at once on an absent node.toBeHidden,toBeVisible({ visible: false }),toBeAttached({ attached: false }), andwaitFordetachedandhiddenread a locator under a frame missing from the document as zero matches instead of timing out waiting for the frame.toHavePropertyreads a primitive on the path through its wrapper, as Jest does, sotoHaveProperty('label.length', 3)passes.expect(browser).toHaveURLtakesignoreCase, andurlMatchesine2e/enginetakes it as an optional fourth argument. -
#784
cc54c6bThanks @okwasniewski! - Every model call now identifies e2e, whichever provider serves it: theUser-Agentstarts withe2e/<version> (<platform>; <arch>)ahead of the AI SDK's own, andHTTP-Referer: https://tester.army/e2ewithX-Title: e2eattribute the traffic on the Vercel AI Gateway and OpenRouter. These replace anHTTP-RefererorX-Titleset on the provider instance. Before, only subscription logins sent the e2e user agent. -
#722
f0f9c8dThanks @okwasniewski! - Explicit navigation now uses an allowlist instead of a denylist:app.open,browser.goto, and the agent'snavigateverb admithttp:,https:, and the exactabout:blank, and every other scheme (chrome:,blob:,about:srcdoc, ...) isPOLICY_DENIED. A wrapped scheme such asview-source:file:///...no longer loads a local file; it isPOLICY_DENIEDlikefile:itself.device.openLinkanddevice.openAppalso refuseview-source:,blob:, andfilesystem:links.browser.setCookiesrefuses anabout:blankcookie URL withPOLICY_DENIED. -
#729
4ce01b2Thanks @okwasniewski! - A registered secret no longer leaks through the observed screen: a test id, an iframe name in a frame path, a selector, or an attribute holding one is masked in the model's text, an executor's tree, and cache entries, and a secret with a line break, tab, CRLF, no-break space, or repeated spaces that an engine collapsed and then cut at the name or text limit no longer leaves its leading part in observations or end anchors. -
#798
02ad48aThanks @okwasniewski! - Breaking:@e2e-dev/webdepends onplaywright-corepinned to an exact version instead of peering onplaywright. Projects no longer install Playwright themselves, and the engine always runs the Playwright it was tested against. Removeplaywrightfrom your dependencies unless your app uses it for its own tests:
bash
npm uninstall playwright
Install browsers in CI with the package's new command, which runs the engine's Playwright: npx @e2e-dev/web install chromium --with-deps (pnpm exec e2e-web install chromium --with-deps under pnpm) replaces npx playwright install chromium --with-deps. Open traces with npx playwright-core@1.63.0 show-trace <file> (pnpm dlx in a pnpm project). An app that also depends on @playwright/test keeps its own copy; the two share a browser cache only when their versions match. e2e init no longer adds playwright to a new project.
- #725
9b93575Thanks @okwasniewski! - Anexpect.pollthat the test body, anafterEachhook, or a fixture teardown returns without awaiting is now cancelled and fails that phase withSTEP_NOT_AWAITEDat the line of the call; in abeforeAllorafterAllhook it fails the hook withHOOK_FAILED. Before, the test passed and the poll's timeout failed whichever test ran next, or nothing at all.
The negation window of a negated expect(locator) or expect(browser) matcher now starts when the first read that saw the negation was issued, not when it returned, so a slow read counts toward it. With a timeout under a second, a negation that held for the whole budget now passes at the deadline; before, a slow first read made a true negation time out.
-
#777
79931d0Thanks @DeryFerd! - The replay cache abstracts short prefixed record ids as minted segments. A route segment of letters, a dash, and a digit tail of two or more (PROJ-016,INV-2041) now reduces to:id, so a step recorded on one record replays on the next instead of missing withwrong-contextand running live every time. A single trailing digit (page-2) stays a route word. -
#724
3a9667cThanks @okwasniewski! ---last-failedno longer goes green on tests it never ran. It reruns tests--max-failuresskipped, and a test or failedbeforeAll/afterAllanother filter leaves out stays owed in the report's newrun.carrieduntil a rerun runs it. A rerun also keeps the artifacts of the run it reruns and writes its own underartifacts/rerun-<n>/in the output directory (.e2eby default), so the folded pull request comment keeps its evidence and stays red while anything is owed. -
#721
94ddbfeThanks @okwasniewski! - Two credentials or two secrets whose names map to the same override variable (api-keyandapi_keyboth readE2E_SECRET_API_KEY) now fail the config load withINVALID_CONFIGinstead of both silently taking one value. -
#800
b5dc0f8Thanks @okwasniewski! ---strict-cachenow fails a step withREPLAY_STALEwhen the cache directory holds its recording under another key, for example after ane2eor engine upgrade or a change to the agent's context. Before, such a step ran live and spent model calls with no error. The message names the old entry file. A step with no recording still runs live.