Pass or fail is not enough: how 14 verdicts are decided
A red test should mean the application is broken. When red can also mean the DNS blinked, the site showed a CAPTCHA or the browser died, the team learns to rerun reds instead of reading them, and the real defect hides in the noise. Nexus Studio gives every run one of 14 verdicts instead of two. This page explains how the engine picks one, in the order it decides, and where the method has limits.
By the Nexus Studio team, Verdict System LLC. October 7, 2026.
Three questions, in order
- Did the test produce a result at all?
- If it passed, did it pass cleanly?
- If it failed, whose fault is it: the harness, the environment, or the application?
Only one answer is a bug report: a failure that belongs to the application. Every other answer is information about something else, and it gets a name that says so.
No result is not a result
Three states are not among the 14, because none of them says anything about your application: SKIPPED (the test never ran), INTERRUPTED (Ctrl+C, a --max-failures cutoff or a global timeout stopped it) and CONFIG_ERROR (the harness could not start it). None of them files an evidence folder or a History row.
INTERRUPTED exists for a concrete reason. Without it, a stopped test fell through to the pass branch and was recorded as SECURE: a green verdict for a test that never finished.
How a pass is graded
- SECURE: passed on the first attempt, and nothing was repaired.
- AUTO_RECOVER: passed, but self-healing had to replace a selector in this test. The locator the author wrote is dead, and the next change to the page breaks the suite for real, so a healed pass is never reported as clean. Fix the original selector.
- FLAKY: failed, then a retry passed, and nothing was healed. Something is unstable, in the application or in the test, and nothing was fixed.
One exception: when the first attempt failed only because it ran out of time and the retry passed, the verdict is SECURE. That pass shows latency, not instability.
How a failure is assigned
When a test fails, the engine reads the error message and the network log of the run and tests them against groups of patterns in a fixed order. The first group that matches decides the verdict, so the order is part of the design.
- CONFIG_ERROR first:
Cannot find module, aReferenceErrororSyntaxError, a missing environment variable, a fixture that failed in setup. If the harness never started, any later symptom is a consequence of that. The patterns are deliberately narrow, limited to errors that only test code or its environment can produce: one CONFIG_ERROR too many would blame the harness for a real bug in your application. - CRASHED:
Target closed,Page crashed, a DevTools protocol error,SessionNotCreatedError, an invalid session id. The browser or the driver died. It is infrastructure: look at the machine and the driver. - BLOCKED: a 403 or a 429 in the error or in the network log, a CAPTCHA, an access-denied page, or a bot challenge the engine recognized on screen. The site defended itself, and the request never reached your product.
- UNREACHABLE: the network said no before any page was observed: the name did not resolve (
net::ERR_NAME_NOT_RESOLVEDin Chromium,NS_ERROR_UNKNOWN_HOSTin Firefox,Could not resolve hostnamein WebKit), the connection was refused, or the machine was offline. A certificate error is not here: that one is a finding. - TIMEOUT: Playwright's
Timeout 30000ms exceeded, Selenium'sWait timed out after, aTimeoutError, or a test Playwright marked as timed out. A navigation that ran out of time is a TIMEOUT, not UNREACHABLE: it started, and the page was slow. - BREACHED_FLAKY: nothing above explains it, and the test had already been retried. It failed, and the failure itself is inconsistent. Compare the attempts before you file a bug.
- BREACHED: nothing above explains it, on a single attempt. An assertion failed, and no environment or harness pattern accounts for the error. This is the bug report.
Every group covers both Playwright and Selenium WebDriver wording. Before the Selenium patterns existed, a WebDriver TimeoutError and an outdated chromedriver's SessionNotCreatedError were both reported as BREACHED, as if the application had a defect.
A real Android phone that cannot take the run, because it fell asleep or is locked, is BLOCKED too: an unmet prerequisite, not a product failure. The rule came from a measured case. A phone fell asleep halfway through a campaign, and a login test that had passed four times in a row came out as a critical red.
Chaos runs
When the run injects faults from Chaos Lab, the verdict answers a different question: did the application hold up? A pass, SECURE or AUTO_RECOVER, becomes CHAOS_RESILIENT. Anything else becomes CHAOS_BREACHED, including a FLAKY pass and a TIMEOUT, because under an injected fault those are the impact. SKIPPED, INTERRUPTED and CONFIG_ERROR stay as they are: a test that did not run says nothing about resilience.
Strict checks that can change a pass
Three checks run next to your assertions: security headers, accessibility with axe-core, and performance budgets for Web Vitals and lazy loading. They record what they find on every run. They change the verdict only when you make them strict in Settings. On an existing application, any of them would turn the whole suite red on the first day, and a suite that is always red is a suite nobody reads.
When they are strict, they apply only to a test that passed, in this order: SECURITY_VIOLATION, A11Y_VIOLATION, PERFORMANCE_DEGRADED. Security goes first because, of the three, it is the one that can end in an incident rather than a complaint. A failed test is never escalated: BREACHED already says something more serious, and an advisory on top of it would hide the finding.
What CI sees
The seven passing verdicts exit 0: SECURE, AUTO_RECOVER, FLAKY, CHAOS_RESILIENT, PERFORMANCE_DEGRADED, A11Y_VIOLATION and SECURITY_VIOLATION. A reservation shows in the report and in History, not as a broken build. The other seven exit 1. The pytest and Selenium runners exit 2 when the run could not start at all, for example with no Python or with a browser the engine has no driver for.
Limits of the method
- The classification reads error text. A message you wrote yourself that contains
403,ForbiddenorAccess Deniedreads as BLOCKED, and an error that containsis not a functionreads as CONFIG_ERROR. Keep custom assertion messages specific, and open metadata.json when a verdict surprises you. - A TIMEOUT is not proof that the application works. It says the run cannot tell. A test that times out on every run deserves a look.
- Retries are what separate FLAKY from SECURE and BREACHED_FLAKY from BREACHED. With retries off, an intermittent failure is a BREACHED.
- API tests have seven verdicts of their own, decided by other rules: see API tests.
See it on a real failure
The Evidence page follows one failed test and lists every file each verdict leaves on your disk: report.md, metadata.json, raw-data.json, a screenshot and, for a real failure, the video. The 30-day trial runs the same engine on your own tests.