Your tests stay Playwright, Selenium and pytest. We don't replace them. We make them enterprise-ready, and you own everything: code, evidence and history.

Every screen of the Studio, and what it does.

Tests are created, run, scheduled and judged from these screens. The command line stays for pipelines.

Standard code

What the files look like

Every test the Studio writes or imports is framework code your team already knows. The engine adds evidence, telemetry, verdicts and self-healing around it; it does not replace it.

  • Playwright, JavaScript and TypeScript: a Playwright Test file. The fixture import at the top is the Studio's; the body is page, locators and expect.
  • Playwright, Python: plain pytest-playwright, a test_ function that takes page. It runs with pytest alone.
  • Selenium, Python: pytest with a driver fixture. The workspace's conftest.py falls back to a plain WebDriver when the engine is not the one running it.
  • Selenium, JavaScript: selenium-webdriver code inside run({ driver, By, until }). The Studio's runner builds and quits the driver.
  • API tests use the Studio's own JavaScript format.

How to leave

Your tests, evidence and history are files in your workspace. To run the suite without Nexus: the Python tests run as they are, the Playwright JavaScript imports point back at @playwright/test (a visual checkpoint, if you added one, needs replacing), and the Selenium JavaScript run() functions get a few lines that start a driver. API tests would need rewriting for another runner.

Create

Patrol

Where tests are written, run and scheduled. The tree shows every test with a logo per browser, colored by its state, and the run strip over the editor says in words what each browser is doing.

  • Five tabs: Test cases, Run patrol, Nexus Recorder, AI writer and Test Adapter.
  • Targets: Desktop matrix (Chromium, Firefox and WebKit), Mobile emulated, Real device (a real Android phone), or Desktop + real device for the PC and an Android phone at once, with the evidence filed per identity.
  • One browser after another by default, or up to three at once and several tests at once when you set it. Selenium runs on Chromium and Firefox.
  • New tests rehearse in the sandbox, which writes no history, before their runs start to count.
  • Feature flags: each flag runs on, off or both, sent by cookie, header, local storage or query string, and the card names the combination a test fails under.

AI writer

Writes a test from what you describe, with the AI you choose: your own AI command-line tool (Claude Code, Codex or opencode) or an API, a local Ollama model, Anthropic or OpenAI. It writes Playwright in JavaScript, TypeScript or Python, and Selenium in JavaScript or Python.

  • Three ways to work: Chat with my AI, Step by step (plan the cases, write the ones you pick), or Bring my agent, a brief for the agent in your own terminal.
  • It can read up to four pages of your app, read-only and without signing in, so the test names real elements.
  • The CLI runs with no tools in an empty folder, or in its own read-only sandbox, and you consent to each destination first.
  • Every test passes the Vet, seven checks without AI: it has a test id, parses, uses the Studio's fixture, reaches nothing outside the page, has assertions that can fail, has no sleeps, forced clicks or mocks, and no fragile locators.
  • Try it runs the test twice in Chromium as a draft that writes nothing to evidence or history. Accept stays disabled until you have opened the code.

Your AI, through its own command line

  • Claude Code
  • Codex
  • opencode

The writer starts the CLI you are already signed into on this machine, so there is no API key to paste into the Studio. Your plan with that tool still applies. A provider API key or a local Ollama model is the other way in.

How the Studio starts Claude Code, argument for argument:

claude -p --output-format json --tools "" --disable-slash-commands

--strict-mcp-config --setting-sources "" --no-session-persistence

--permission-mode dontAsk --max-budget-usd 1

The prompt goes in on stdin. No tools, no MCP servers, none of your own settings or hooks, nothing saved as a session, and a spending cap on every call: one US dollar by default.

Nexus Recorder

Records what you do in the browser as test code. There is no AI in the loop: what you click is what gets written.

  • Record for Desktop, Mobile emulated or a Real device: an Android phone on USB, driving the phone's own Chrome.
  • Writes Playwright in JavaScript, TypeScript or Python, or Selenium in JavaScript or Python.
  • Or records the API calls the page makes as REST or GraphQL tests, without auth or cookie headers, and with secret-looking fields turned into environment variables.
  • Selector Maps scan a page without AI, grade each selector HIGH, MEDIUM or LOW, and keep the map on your disk.

Test Adapter

Brings the tests your team already has into the Studio, without rewriting them by hand.

  • Paste a test, drop a file or point at a folder: Playwright or Selenium in JavaScript, TypeScript or Python, under Playwright Test, mocha, jest, pytest or unittest.
  • Rewrites by rule, not by AI: the browser launch and setup move into the Studio's fixture, while assertions, conditions, loops and waits stay as they were.
  • Shows the diff side by side and the assertion count before and after. A lost assertion blocks the import unless you accept it, line by line.
  • An optional AI handles only what no rule covers, and its code faces the strict Vet. Kept tests rehearse in the sandbox until three clean runs.
  • Cypress, WebdriverIO, Robot Framework, Java and C# are not supported.

Repair line and self-healing

What happens when a locator breaks. Nothing is repaired behind your back.

  • Self-healing, set to Suggest by default: after a failure, a background browser replays the test up to the failed line and reads the page, so the candidates are ready. No AI, and the test stays red.
  • Repair line, from the gutter of that line: re-record just that step, or swap the selector, strongest first and XPath last. The change lands in the editor unsaved, and saving it is your approval.
  • Ask my AI which one can reorder the candidates, never add one. The previous version of the file stays in its timeline.
  • Locator health lists the open test's weak locators and repairs them the same way.

Run

Devices & Mobile

Real Android phones, connected by USB to the machine that runs the engine. No device farm, no minutes to buy.

  • Several phones at once, over adb and the Chrome DevTools Protocol, with a guide to connect the first one.
  • Watch screen mirrors a phone on your PC: controllable from this screen, and view-only when opened from Patrol.
  • Mobile emulation, from Playwright's device profiles: Chrome on Android, Safari on an emulated iPhone through WebKit, and a low-end profile. Real devices are Android phones only; an iPhone is always emulated.
  • Tests on a phone run one at a time, and a phone run always stays on the machine the phone is plugged into.

Watch

In Patrol, under Run patrol: re-runs a selection on a timer, unattended.

  • Repeat every 5 minutes to 24 hours, or Once, on a date. It ends when you stop it, after a number of runs, or on a date.
  • It runs as its own process, so closing the Studio does not stop it. It needs that machine on: while it sleeps, scheduled runs wait or are skipped.
  • Alerts on Telegram or Slack only when a state changes, as ALERT or RECOVERED.

CI/CD

Sets up the pipeline on screen instead of in a YAML editor.

  • Writes the GitHub Actions or GitLab CI file: browsers, visual, accessibility, chaos, loop and retry settings, and on GitHub a nightly regression at a UTC time with the tests you pick.
  • Runs on a cloud-hosted runner, your own machine, or your machine with Docker, with a runner script for bash or PowerShell.
  • Validates the pipeline locally first, then commits and pushes from the same screen. Remote runs can be started here and their evidence downloaded.
  • Also scriptable for CI: nexus ci exits 0 on a pass, 1 on a failure and 2 when nothing ran, and writes JUnit XML. Run your tests from CI/CD

Docker and SSH

The Studio is a Windows 11 desktop app, on your PC or, optionally, on a Windows 11 mini PC. From it, the engine runs a workspace's suites on this machine, in Docker on it, or over SSH on another machine. The Studio itself never runs in Docker or over SSH.

  • In Docker, the suite runs in the runner container on the same disk, so the evidence lands in your workspace as usual.
  • Over SSH, on a VPS with Ubuntu 24.04 LTS, Docker Engine and Compose v2, reached with key-based access: the engine runs the suite there in the same container image as Docker here, and the evidence stays there until you pull it.
  • Real Android phones, Selenium and Python suites, infrastructure chaos and results import always run on the Studio's own machine.
  • The top bar shows where the workspace runs and its RAM and CPU. Adaptive pacing, when you switch it on, sizes the browsers at once to the free RAM.

Prove

Evidence Locker

Every result, filed where you can read it and triage it.

  • Sections for test results, the run summary, bug reports, passed tests, history intelligence, risk ranking, Core Web Vitals, security audits, chaos resilience and API testing.
  • Each record opens on Report, Metadata, Telemetry, Raw Log and AI Analysis, the last one on by default and filled with your AI connected. A failure files its forensic report, metadata and one screenshot of the final frame, plus the video for a real failure (you choose which verdicts in Settings, under Keep failure video for); the raw log holds console errors, failed requests, security issues and page errors. A trace is kept too while tracing is on, which is the default.
  • Triage: Real bug or False alarm, with a required reason, whether it still counts, and what happens to the files: keep them, archive what matters, or delete them.
  • Share to Telegram, Slack or Jira, copy an image, or export a full HTML report. Nothing leaves the machine until you pick a channel and press send.

One real failure's folder, file by file →

API tests

REST and GraphQL tests with seven verdicts of their own. They run from Patrol like any test, and each endpoint files its own folder in the Evidence Locker's API testing section.

  • API_SECURE: every assertion held and the response shape matches the last one recorded. API_FLAKY: it passed on a retry. API_DRIFTED: it passed, but the shape changed, and the diff names each field added, removed or retyped. Drift counts as a pass until you make it strict.
  • API_BREACHED: an assertion failed. API_DEGRADED: slower than the latency budget, 2000 ms by default, and a verdict only when performance is strict. API_VULNERABLE: a request with no credentials was answered where a 401 or 403 was expected; a missing security header alone never makes it one. API_BLOCKED: a 429, or a 403 with a Cloudflare or Akamai signature, so the request never reached your product.
  • On screen: tiles for endpoints, passing, runs over the budget, shape changed, vulnerable ever and slowest observed; the filters all, failed, drifted and vulnerable; and a table of endpoint, verdict, min, median, p95, max, n and pass rate. An endpoint with under 20 observations shows a dash, not a percentile.
  • Click a run to open it on six tabs: Report, Request, Response, Performance, Schema and AI Analysis, the last filled with your AI connected. A failed step of a flow offers Re-run from step N and Re-run failed step.
  • Filed per run in api-evidence: request.json, response.json, performance.json, schema.json, report.md and metadata.json, with secrets redacted in headers and body. report.md ends with a curl command that replays the call. Passing runs keep the newest 20 per test; failures are always kept.
  • API Flow Recorder: in the Nexus Recorder, record the calls the page makes, then Review the recording before the suite is written. Calls lists each one with its method, path, status, category and the reason. Only business, auth, read and state-changing API calls start in the suite; telemetry beacons, third-party hosts, assets and preflights start out, and you can tick them back in. Flows chain a value from one response into a later request, such as a login token, and Redactions lists the secret-looking fields marked never written.
  • From the review: Generate test writes the suite, Generate OpenAPI writes a 3.1 draft with types and status codes and no recorded values, and Load test opens the calls as a Load plan.

A drift, as the Schema tab reads it

APIT311-1 GET /api/productsList API_DRIFTED

+6 fields, -2 fields

+ products[].brand + products[].id + products[].name + products[].price

+ products[].category.category + products[].category.usertype.usertype

- brands[].brand - brands[].id

APIT311 is a demo test written on purpose to drift: a live API does not change on demand, so it alternates between two catalog endpoints on successive runs, and its second run reads API_DRIFTED. A drift is the change a consumer meets in production while an assertion about status 200 still passes.

History

Each test on each identity, run after run.

  • Tabs for the summary, reliability, performance, API, failures, coverage, quality, and unit and integration results.
  • Trends by session, day or week: Regressed, Recovered, Still broken, Slower. The identity matrix sets every test against every browser and device.
  • Unit and integration: import JUnit XML from Jest, Vitest, pytest, Maven or gotestsum, kept apart from the end-to-end numbers.
  • Risk, in the Evidence Locker and in Patrol: a ranking by how often a test flips, breaks or costs time, where nothing is called flaky below five runs, and a predicted risk from a logistic regression over your history, with an EWMA fallback when data is thin.

What one row of History holds →

Visual Regression

Pixel comparison against the pictures you approved. The first run with --visual plants the baseline and every later run is compared with it. The screen opens on what needs a decision: Pixels moved.

  • Each regression is a card: the test, the checkpoint, the share of pixels that moved against its tolerance, "N of M px", the dimensions, how many areas changed, and a note when the capture never settled. Edge pixels re-blended by anti-aliasing are not counted.
  • Side by side, Swipe, Diff and Blink views, and the earlier runs of the same checkpoint: one regression is a fact, four in a row each a little worse is a UI drifting, and four identical ones is a test recorded against something unstable.
  • Baselines are kept per browser and platform. Approve this change replaces only that checkpoint's image; Approve all is for a change that hits many, such as a global header.
  • Settings on the same screen, with their defaults: Compare pixels on every run off, Full-page tolerance 1.0%, Element checkpoint tolerance 0.1%, A regression fails the test off (the regression is recorded and the test stays SECURE; on, it reads BREACHED), and Masked on every photo (CSS selectors) for a clock or an ad.
  • Filed per regression in visual-regression: baseline.png, current.png, diff.png (red changed, amber anti-aliasing ignored, grey unchanged), visual-report.md and result.json, the newest 10 per test and checkpoint at about 1 MB a set.

A checkpoint on one element

nexus patrol TC01 --visual

# inside the test: one element instead of the whole page

await visualCheck(page, 'cart-total', page.locator('#cart-total'));

An element checkpoint is held to a tenth of the tolerance a whole viewport needs: narrowing the scope makes the check strict without turning a clock or an ad into a failure. Use it when a spacing, color or layout change could break a page that every assertion still passes. Mask what changes on purpose instead of raising the tolerance.

Executive Report

The quality of a release on one PDF, for whoever signs it off.

  • Pick a period: all time, the last 7, 30 or 90 days, or a custom range.
  • A release call of go, caution or no-go, with confidence, coverage and pass rate.
  • Failures with their root cause and business impact, stability by test, security findings, API testing and the quality pyramid of end-to-end, integration and unit results.
  • Recommendations come from rules, not from an AI.

Stress and security

Chaos Lab

Resilience testing, from the browser to the container.

  • Browser faults: Network (≈3G) latency, an Offline blackout mid-test, a CPU throttle 4x on Chromium, or all at once, with packet loss and memory pressure on top.
  • Infrastructure: register a Docker container, or the app under test, then pause, stop, kill or restart it, with or without a test running against it.
  • Recovery is measured with a health probe every 500 ms: time to detect, time to recover, availability. The outcome reads RESILIENT, DEGRADED, NOT_RECOVERED or NOT_INJECTED.
  • Only registered targets, never the Studio's own containers. The fault is written down before it fires and always restored, and a pass where no fault fired is never called resilient.

Load

Concurrency, percentiles and a verdict, from the engine itself.

  • Worker threads in the engine, no external load tool, and a report in Apache JMeter's statistics.json schema.
  • Requests per second, p50, p90, p95 and p99, error rate and APDEX. The verdict is PASS, OVER_BUDGET or UNMEASURED.
  • Measure Web Vitals under load: a real Chromium page sampled before, every 15 seconds during, and after the load, for LCP, CLS, FCP, TTFB and TBT.
  • It arms only after you type AUTHORIZED, and a target that is not this machine asks again.

Security

Runs the security checks you configure and records what it found, with the evidence. Smart Monkey, the DAST scanner, lives on the Security screen under Run a scan.

  • Smart Monkey scans a target you are authorized to test: XSS, SQL injection, broken access control, IDOR, SSRF, command injection, path traversal, template injection, XXE, open redirects, CORS, security headers, cookie flags and weak JWT secrets, among others.
  • Pick a phase or run all of them: pre-auth, auth-transit, post-auth, http-probes (cookies, CORS, headers, JWT, redirects, files), ssrf, cmdi, xss-oob, sqli (blind, boolean and time-based), idor, or agent (orchestrated, budgeted discovery). A plain-language line explains the one you picked.
  • Out-of-band confirmation uses an HTTP listener on your own machine: each payload carries a random 32-character token, and SSRF, command injection and blind XSS count only when the target calls back with it. There is no DNS callback and no public service such as Interactsh or Burp Collaborator, so nothing about your target goes to a third party.
  • Authorization to test: a host with no recorded basis is refused before a packet is sent, and private and local addresses pass without one. Every scan records its authorization, operator, target, operation and a 60-minute window, sealed locally and never sent anywhere.
  • Each session files one folder per finding with audit_report.json, screenshot.png, replay.webm when the recording survived, and an executable replay.mjs that nexus replay runs against the finding's folder. The session root adds session_summary.json and results.sarif. The newest 10 sessions are kept.
  • SARIF 2.1.0, one result per finding: the rule is the finding's CWE, CRITICAL and HIGH read error, MEDIUM warning, the rest note, and the location is the probed URL. GitHub code scanning and GitLab ingest it directly, and nexus monkey writes it on a CI runner without the Studio.

One session on disk

monkey-audits/YYYY-MM-DD_HH-mm-ss_FAST_AUDIT/

session_summary.json results.sarif

BUG-NN_TYPE/ audit_report.json screenshot.png replay.webm replay.mjs

The folder is stamped in UTC and ends in FAST_AUDIT or DEEP_AUDIT. A finding is complete in its own folder, so it can go to whoever fixes it with its screenshot, its recording and the script that reproduces it; in the Evidence Locker, Archive… keeps those parts as well. Use it before a release, on a target you own or are authorized to test.

Settings

Quality checks inside a run

Switched on per run in Run patrol or in Settings. Each one is advisory until you make it strict, and then it becomes a verdict.

  • Accessibility with axe-core against WCAG 2.1 AA, plus touch-target size on phones.
  • Web Vitals against budgets (LCP 2.5 s, CLS 0.1, INP 200 ms, TTFB 800 ms), and a lazy-loading audit that grades from A to F how the page loads its images.
  • Block detection: a 403, a 429 or a challenge page reads BLOCKED. Automation signals are hidden from the site by default, and one switch turns that off.
  • Strict mode turns a break into A11Y_VIOLATION, PERFORMANCE_DEGRADED or SECURITY_VIOLATION.

Timeouts

How long a Playwright step may wait before the run reads TIMEOUT. Four dials in Settings, under Execution & Timeouts. The values below are the defaults.

Settings · Execution & Timeouts

  • Action timeout (click, fill…)30000 msNEXUS_ACTION_TIMEOUT · every click, fill or wait on a locator
  • Assertion timeout10000 msNEXUS_EXPECT_TIMEOUT · how long an expect() keeps retrying
  • Page navigation timeout60000 msNEXUS_NAV_TIMEOUT · page.goto() and its redirects
  • Per-test hard timeout360 sNEXUS_TC_TIMEOUT · the whole test, six minutes

One step can take its own limit in the test's code, as standard Playwright:

await page.getByRole('link', { name: 'More information...' }).click({ timeout: 5000 });

await expect(page.getByRole('heading', { level: 1 })).toHaveText('Example Domain', { timeout: 15000 });

TC91 is a demo test written on purpose against example.com's old link text, More information...; the page now says Learn more. Its click waited the default 30000 ms for a link that was no longer there, then read TIMEOUT. A longer limit does not hide a broken element: it still fails, later.

AI Providers

The AI behind the forensic analysis of failures and the triage of Smart Monkey findings.

  • Local only is the default: a model on this machine through Ollama, with no fallback to the cloud.
  • Auto, Anthropic or OpenAI when you choose, and the screen says plainly when prompts are leaving your network. Keys are write-only.

Settings

The switches behind the other screens, in one place.

  • Integrations: Telegram, Jira and Slack. Off until you switch it on; then Jira files a failure (BREACHED, BREACHED_FLAKY or CHAOS_BREACHED) or a TIMEOUT, never a block, a crash, an unreachable site or a configuration error.
  • Evidence and retention: how many bug reports, passing screenshots and scan sessions each test keeps.
  • Selectors and self-healing, adaptive pacing by RAM pressure, and the automation signals the browser shows.

The full inventory

19 more capabilities, at a glance

Nexus Recorder

Record a flow into Playwright (JavaScript, TypeScript, Python) or Selenium (JavaScript, Python), on the desktop, emulated mobile or a real Android phone, or into the API calls the app makes. No AI in the loop.

Test Adapter

Paste a test or point at a folder: Playwright or Selenium in JavaScript, TypeScript or Python. It shows the diff, keeps every assertion, and rehearses the test before it counts.

Repair line

From the line a test failed on: re-record just that step, or swap the selector, strongest first and XPath last. The old version stays in the file's timeline.

Self-healing that asks first

After a failure, the Studio replays the test to that line in the background and has candidate selectors ready. The run stays red until you choose one.

Locator health and selector maps

Lists the open test's weak locators and repairs them in place. A selector map scans a page without AI and grades every selector it finds.

Watch

Re-runs a selection every 5 minutes to 24 hours, or once on a date, and alerts only when a state changes. It keeps running after the Studio closes.

CI/CD, set up on screen

Writes the GitHub Actions or GitLab CI pipeline, with a nightly regression, and validates it locally first. For the pipeline itself: a strict exit code and JUnit XML.

Feature flags

Runs a test once per on and off variant of your flags and names the combination it fails under.

Quality pyramid

Import JUnit XML from Jest, Vitest, pytest, Maven or gotestsum and see unit and integration results beside your end-to-end runs.

Risk ranking and prediction

Ranks each test on each identity by how often it flips, breaks or costs time. A logistic regression over your own history adds a risk score, with an EWMA fallback when data is thin.

Visual regression

Compare side by side, swipe, diff or blink against baselines kept per browser and platform, and approve a change in one click.

Accessibility

WCAG 2.1 AA checks with axe-core, plus touch-target size on phones. Advisory by default; in strict mode a break becomes an A11Y_VIOLATION.

Web Vitals and lazy loading

LCP, CLS, INP, FCP, TTFB and TBT, on by default and judged against budgets, plus a lazy-loading audit that grades how the page loads its images.

Block detection

A 403, a 429 or a bot-challenge page is read as BLOCKED, not as a failure of your app, and never files a Jira ticket.

API suites

REST and GraphQL: status, fields, types and latency budgets, with verdicts of their own and the request, response and schema filed as evidence.

Executive report

A PDF with a release call (go, caution or no-go), coverage, stability, failures by root cause, security findings and rule-based recommendations.

Jira, Slack and Telegram

Jira files a bug on BREACHED, BREACHED_FLAKY, CHAOS_BREACHED or TIMEOUT with 24-hour dedup, Slack gets Block Kit alerts, and Telegram gets alerts from a bot of your own, paired by QR.

Docker and SSH

The engine runs the same suite in the runner container on the Studio's machine, or over SSH on a VPS in the same container image, with the same verdicts. Each workspace picks its runtime. Over SSH the evidence stays on that host until you pull it.

Desktop app

Nexus Studio for Windows 11 x64, built on Tauri, on your PC or a mini PC with 16 GB of RAM or more: runs, evidence, devices and history in one window.

Try every screen on your own app.

Free for 30 days, with every feature, on your machines. No card needed.