Your tests stay Playwright, Selenium and pytest. We don't replace them. We make them enterprise-ready, and you own everything: code, evidence and history.

One failed test. Everything it left on your disk.

A real run, not a mock-up: TC92 on 2026-09-30, headless, in Chromium, Firefox and WebKit. It opens example.com, follows Learn more to IANA and expects the heading to read Example Domain. The page says Example Domains. These are the Chromium run's files, their real sizes and what is inside each one.

evidence/bug-reports/chromium/TC92 - Iana heading/2026-09-30_14-34 [BREACHED]

5 files
in the run's folder, plus one row in History
531.2 kB
the whole folder, the video included
32 fields
in metadata.json, for machines
9 sections
in report.md, the forensic report

What one failure leaves behind

The folder a failure files, opened.

Line 23 of TC92_Iana_heading.js expects the heading to read Example Domain. IANA's page says Example Domains: in 10000 ms the locator resolved 23 times and read the plural every time. BREACHED, on all three engines. Here is the frame it failed on, and every file it wrote.

https://www.iana.org/help/example-domains
The IANA page at the moment TC92 failed: its heading reads Example Domains, where the test expected Example Domain
screenshot.png · 89.8 kB · one viewport screenshot of the final frame, 1280×720: what a user would have seen when the check ran
Error: expect(locator).toHaveText(expected) failed

Locator:  getByRole('heading', { level: 1 })
Expected: "Example Domain"
Received: "Example Domains"
Timeout:  10000ms

  23 × locator resolved to <h1 id="example-domains">Example Domains</h1>
     - unexpected value "Example Domains"

In the Evidence Locker the same record opens on five tabs:

ReportMetadataTelemetryRaw LogAI Analysis

report.md5.3 kB

The forensic report, in Markdown: it reads without the Studio.

  1. Incident summary
  2. Root cause analysis
  3. Network telemetry
  4. Failure timeline
  5. Historical stability
  6. Self-healing log
  7. Evidence artifacts
  8. Performance budget
  9. Lazy-loading audit

Use it to file a ticket: the Incident summary names the test, browser, page and business impact, and the Root cause analysis the failed check, selector, expected and received values and the source line.

metadata.json2.2 kB

32 fields, machine-readable. A few of them, as written:

"status": "BREACHED",
"pageUrl": "https://www.iana.org/help/example-domains",
"selector": "getByRole('heading', { level: 1 })",
"expectedVal": "\"Example Domain\"",
"receivedVal": "\"Example Domains\"",
"timeoutMs": "10000ms",
"securityIssueCount": 3,
"attempt": { "number": 2, "of": 2 },
"lazy": { "score": 85, "grade": "B", … }

Use it to filter failures with a script instead of a Markdown parser. It is also the file the chain of custody seals: a folder with a metadata.json is a record.

raw-data.json5.1 kB

The raw log: every console line, network error, security issue and page error of the run, and the performance samples.

"consoleLogs": [],
"networkErrors": [],
"securityIssues": [
  "Missing CSP: https://example.com/",
  "Missing CSP: https://example.com/s.js",
  "Missing CSP: https://iana.org/help/example-domains"
],
"jsErrors": [],
"performance_metrics": { "lcp": 992, … }

Use it to see what the page did around the failure. A FLAKY or AUTO_RECOVER pass keeps the failed attempt's raw-data.json, so a retry that hid the failure still leaves its log.

Telemetrymeasured

Measured at the end of the test, against your budgets. The bar is the measurement over its budget.

  • LCP992 / 2500 ms
  • FCP992 / 1800 ms
  • CLS0.0295 / 0.1
  • TTFB727 / 800 ms
  • TBT0 / 200 ms
  • DOM114 / 1500 nodes
  • Heap5 / 50 MB
  • INPnot measured

All within budget, and the report flags TTFB as close to it. INP reads not measured, never a pass: nothing was clicked or typed on the page it measured. The lazy-loading audit scored the page 85 of 100, grade B: one image without a width and height.

Use it to tell a slow page from a wrong assertion.

video.webm428.8 kB
A frame of video.webm, eight seconds in: the IANA page with its Example Domains heading, while the check retried
0:08 of 0:16

The whole session, 16.6 seconds: example.com, the click on Learn more, then IANA's page while the check retried. Kept because the run was BREACHED, and it plays in the Evidence Locker. A real failure keeps its video: BREACHED, BREACHED_FLAKY and CHAOS_BREACHED by default, TIMEOUT and BLOCKED when you turn them on in Settings. CRASHED and UNREACHABLE are never offered, because that recording comes out broken or blank. Only Playwright web and emulated mobile runs record: a real Android phone, a Selenium run and an API test do not.

Use it to see how the page got to the failed check, step by step; the screenshot above is the last frame.

AI analysisyour AI

On by default, with your AI connected. With a local Ollama model, the default, or another provider you pick, the engine adds four headings to the report: Diagnosis, Probable cause, Impact and Recommended action. When your own AI command-line tool is on in the AI writer, it reads report.md after the run instead, for up to five failures per run, and adds a fifth heading:

  • Diagnosis
  • Classification (command-line tool only)
  • Probable cause
  • Impact
  • Recommended action

The classification is exactly one of APP_BUG, APP_CHANGE, TEST_AUTHORSHIP, TEST_SETUP or INFRA_FLAKY. No AI connected: no analysis, and the five files stay the same.

trace.zip is a sixth file. With Record a forensic trace on failure on, which is the default, a BREACHED folder also files a Playwright trace (DOM snapshots, screenshots and sources), opened with npx playwright show-trace trace.zip. It is not recorded on a real Android phone. TC92 ran before tracing became the default, so its folder holds the five files above.

On the home page: TC91's TIMEOUT folder, from a demo test staged against example.com's old link text →

Verdicts

The 14 verdicts, and what each one tells you to do next.

The verdict is in the folder's name: TC92's ends in [BREACHED]. The classifier reads a failure in a fixed order: the harness first, then a crash, a block, an unreachable host and a timeout, and only then BREACHED_FLAKY or BREACHED. So a BREACHED is a failure no environment pattern explains. Each card says whose turn it is and what that verdict files.

  • SECURENobody

    Passed on its own: no retry, no repaired selector. Nothing to do. Filed in evidence-passed-test, the newest 5 per test and browser.

    screenshot.pngmetadata.json

  • CHAOS_RESILIENTNobody

    Passed while a fault was injected, so the app held. Filed in chaos-runs, apart from the bug reports, so chaos never skews the stability numbers.

    screenshot.pngchaos-meta.jsonchaos-report.mdmetadata.jsonvideo.webm

  • AUTO_RECOVERTest

    Green, but not unassisted: the self-healer swapped a broken selector. The locator in the test is dead, and the next DOM change breaks the suite for real. Fix the locator. raw-data.json appears when a retry was needed, and it is the failed attempt's.

    screenshot.pngmetadata.jsonraw-data.json

  • FLAKYTest or app

    Failed, then a retry passed it, and nothing was repaired. That is unexplained instability in the app or the test, not a fix. Read the failed attempt's raw-data.json, then the History trend.

    screenshot.pngmetadata.jsonraw-data.json

  • PERFORMANCE_DEGRADEDApp

    Assertions passed and the page was too slow: a Web Vitals or lazy-loading budget was exceeded. Advisory until you make the check strict.

    screenshot.pngmetadata.json

  • A11Y_VIOLATIONApp

    Assertions passed, accessibility did not. metadata.json lists each axe-core rule that fired, with its impact and node count. Advisory until you make the check strict.

    screenshot.pngmetadata.json

  • SECURITY_VIOLATIONApp

    Assertions passed, the security headers did not (a missing CSP, for example). Fix the response headers. Advisory until you make the check strict.

    screenshot.pngmetadata.json

  • BREACHEDApp

    An assertion failed and no environment or harness pattern explains the error. Read Expected and Received in report.md, watch the video, then mark it Real bug or False alarm, with a reason.

    report.mdmetadata.jsonraw-data.jsonscreenshot.pngvideo.webmtrace.zip

  • BREACHED_FLAKYApp or test

    Failed on the last attempt after an earlier retry: the failure itself is inconsistent. Compare the attempts before you file a bug. Its video is kept by default.

    report.mdmetadata.jsonraw-data.jsonscreenshot.pngvideo.webm

  • CHAOS_BREACHEDApp

    Broke under an injected fault. A resilience finding, kept apart from the functional bug reports; the report records the fault applied and the impact on the user. Retention never removes a run that broke. Its video is kept by default.

    screenshot.pngvideo.webmchaos-meta.jsonchaos-report.mdmetadata.json

  • TIMEOUTEnvironment

    Never answered in time, which is not proof the app is broken. Check how fast the page was in Telemetry before you touch a dial: a longer limit does not hide a broken element, it fails later. Its video is off by default; turn it on in Settings, under Keep failure video for.

    report.mdmetadata.jsonraw-data.jsonscreenshot.pngvideo.webm

  • UNREACHABLEEnvironment

    The host could not be reached at all: DNS or the connection failed before any page was observed. Check the URL, the network and the target's availability. Not a finding about the app. No page loaded, so its video would be blank and none is offered.

    report.mdmetadata.jsonraw-data.jsonscreenshot.pngvideo.webm

  • BLOCKEDEnvironment

    The site defended itself: a 403, a 429, a block page or a bot challenge. metadata.json names the defense when one was on screen. Not a finding about the app, and Jira never files a block. Its video is off by default; turn it on in Settings, under Keep failure video for.

    report.mdmetadata.jsonraw-data.jsonscreenshot.pngvideo.webm

  • CRASHEDHarness

    The browser or the process died: a closed page, a DevTools protocol error, a dead WebDriver session. Infrastructure, not the app. Look at the machine and the driver. The recording of a dead browser comes out broken, so none is offered.

    report.mdmetadata.jsonraw-data.jsonscreenshot.pngvideo.webm

Not among the 14: CONFIG_ERROR (the harness could not start the test: a broken import, a missing .env value), SKIPPED and INTERRUPTED. None is a result about your app, so none files an evidence folder or a History row. A CONFIG_ERROR is still written in full, with a hint, to a log under .run-cache/config-errors/, so the broken test gets fixed. API tests have seven verdicts of their own: API tests.

History

Every run writes one row. The rows judge the test.

Beside the folder, each run appends a row to history/<browser>/<test>/runs.json: the last 50 runs of each test on each browser, by default. These are TC92's three rows, one per browser, as written. The same error_signature on all three says it is one error, not three.

  1. webkitBREACHED

    status
    breached
    duration_ms
    13570
    error_category
    UI_UNSTABLE_ELEMENT
    error_severity
    MEDIUM
    error_signature
    70f7532d
    perf.lcp
    638 ms
    perf.ttfb
    null
    evidence_ref: bug-reports/webkit/TC92 - Iana heading/2026-09-30_14-33 [BREACHED]
  2. firefoxBREACHED

    status
    breached
    duration_ms
    20013
    error_category
    UI_UNSTABLE_ELEMENT
    error_severity
    MEDIUM
    error_signature
    70f7532d
    perf.lcp
    4353 ms
    perf.ttfb
    3877 ms
    evidence_ref: bug-reports/firefox/TC92 - Iana heading/2026-09-30_14-34 [BREACHED]
  3. chromiumBREACHED

    status
    breached
    duration_ms
    12291
    error_category
    UI_UNSTABLE_ELEMENT
    error_severity
    MEDIUM
    error_signature
    70f7532d
    perf.lcp
    992 ms
    perf.ttfb
    727 ms
    evidence_ref: bug-reports/chromium/TC92 - Iana heading/2026-09-30_14-34 [BREACHED]

A value the browser cannot measure is written as null, never as zero: WebKit reports no TTFB here. Row times are UTC, as the row writes them; folder names use the machine's local time, UTC−5 on this one.

Every field of a row

schema_versiontimestampstatusdetail_statusretry_attemptattemptsrun_idruntimeplatformduration_msbrowserchaosvarianterror_summaryerror_categoryerror_severityerror_signatureperfraw_log_refevidence_ref

perf holds LCP, FCP, CLS, TTFB, TBT, INP, long tasks, DOM size, heap, every page visited and the lazy-loading score. raw_log_ref and evidence_ref point back at the folder above.

What the engine prints from it

TC92 chromium BREACHED 13.7s

expect(locator).toHaveText(expected) failed

at TC92_Iana_heading.js:23 waiting for getByRole('heading', { level: 1 })

stability CRITICAL 0% · 0/1 runs · 1 failed in a row

# TC91, the staged demo on the home page: a BREACHED, then a TIMEOUT

stability CRITICAL 0% · 0/2 runs · 2 failed in a row

The row's duration_ms, 12291, is the test itself, as Playwright timed it, on the second of its two attempts.

The same rows feed the report's Historical stability section, the History screen's trends (Regressed, Recovered, Still broken, Slower) and the risk ranking, where nothing is called flaky below five runs.

The History screen →

Retain evidence

A pass keeps little. A failure keeps everything. Old files go.

A suite that runs every night fills a disk. The engine caps every folder it writes, per test and browser, so the disk stops growing at a size you can predict. A failure still keeps what it takes to reproduce it: the error with its selector and values, the page URL, the frame, the console, network and page-error lines, the runtime and platform, and the run's context, what it was measured against. The caps are set in Settings, under Evidence & Retention.

  • A BREACHED failure

    report.mdmetadata.jsonraw-data.jsonscreenshot.pngvideo.webmtrace.zip

    The full folder above: 531.2 kB for TC92 on Chromium, 428.8 kB of it the video. trace.zip comes on top while tracing is on.

  • A TIMEOUT

    report.mdmetadata.jsonraw-data.jsonscreenshot.pngvideo.webm

    The same files without the recording, because TIMEOUT is off by default for video: 23.9 kB for TC91's TIMEOUT, the staged demo on the home page.

  • A pass

    screenshot.pngmetadata.json

    Filed apart from the bug reports, the newest 5 per test and browser. The screenshot is what visual regression compares. A FLAKY or AUTO_RECOVER pass keeps its folder too, with the failed attempt's raw-data.json, while Keep evidence of a FLAKY / AUTO_RECOVER pass is on.

Settings · Evidence & Retention · defaults

  • Bug reports kept per test20NEXUS_BUGREPORT_RETAIN
  • Passing screenshots kept5NEXUS_PASSED_RETAIN
  • Keep evidence of a FLAKY / AUTO_RECOVER passonNEXUS_RECOVERED_RETAIN
  • Raw run history window50 runsNEXUS_HISTORY_RETAIN
  • Smart Monkey sessions kept10NEXUS_MONKEY_RETAIN
  • Run contexts kept (what each run tested against)100NEXUS_RUN_CONTEXT_RETAIN
  • Record failure videoonNEXUS_VIDEO_ENABLED
  • Keep failure video for3 verdictsNEXUS_VIDEO_VERDICTS=BREACHED,BREACHED_FLAKY,CHAOS_BREACHED
  • Record a forensic trace on failureonNEXUS_TRACE_ENABLED

After every run, printed by the engine

[CACHE RETENTION] 3 files archived → .run-cache/telemetry/2026-09/run_30-1434.zip

[CACHE RETENTION] .cache purged (freed 8.9 KB)

[CACHE RETENTION] Archive health: 1/2 files · 9.3 KB / 500 MB used

These three lines followed TC92's run. The run's telemetry scratch files are never evidence: they go into a zip, the loose copies are deleted, and at most 2 zips of up to 7 days stay, within 500 MB. Those three limits are NEXUS_CACHE_ZIPS, NEXUS_CACHE_RETENTION_DAYS and NEXUS_CACHE_MAX_MB in .env, not fields in Settings.

In Settings, Clear cache drops the workspace's test-results and playwright-report folders and closes any browser a crashed or killed run left behind. It is refused while a patrol is running.

Workspace · Disk · the limit each folder has today

area limit today

Bug reports newest 20 runs per test and browser

Passed tests newest 5 runs per test and browser

API evidence newest 20 passing runs per test; failures kept

Chaos runs newest 20 held runs; runs that broke kept

Smart Monkey newest 10 sessions

Visual regressions newest 10 per test and checkpoint

Run context newest 100 runs

Run history newest 50 runs per test and browser

Chain of custody append-only, never trimmed

Archive no automatic limit

The Disk panel lists every folder a run writes to, biggest first, with its size, file count and oldest date, the total and the free space on the drive. A folder with no automatic limit is flagged. Clean is offered only where it is safe: the run cache, test file versions older than 14 days and Executive reports older than 30 days. It never touches evidence or the archive.

Evidence Locker · keep what matters, free the rest

Archive… copies exactly the parts you tick, Screenshot and Report ticked to start, into View → Archive. The copy sits beside the evidence folder, not inside it, so no retention sweep reaches it, and the run in the Locker is not touched.

Restore, on the Archive screen, puts a record back where it was and back into the count. If its folder was freed, the archived parts are written back into the Evidence Locker. A False alarm verdict still keeps it out of the count.

Free disk sends a folder to the recycle bin. Where the record is sealed, the engine writes the disposition first, so nexus verify reports PURGED, not TAMPERED. Retention does the same when it removes the oldest folder, and the ledger keeps proving what the folder hashed to, who removed it and when. Sealing is opt-in: until the first Real bug or False alarm verdict, or nexus verify --seal, the Locker reads UNVERIFIED.

See what your own failures leave behind.

Free for 30 days, with every feature, on your machines. No card needed.