Is this test flaky, or is it a real bug?
A flaky test failed and then passed on a retry without anyone fixing anything, and a real bug fails for a reason no retry changes. Nexus Studio, by Verdict System LLC, tells them apart by recording every retry: FLAKY for a failure that a retry passed, BREACHED for a real finding, and BREACHED_FLAKY for a failure that is itself inconsistent. A retry never turns a defect into a green run.
By the Nexus Studio team, Verdict System LLC. October 7, 2026.
What is a flaky test?
A flaky test is one whose attempts disagree: it fails once and passes on another try, with the code and the test unchanged. In Nexus Studio the word has one meaning. FLAKY is a test that failed, then passed on a retry, with nothing healed. Something is unstable, in the application or in the test, and nothing was fixed.
A hard failure that happens on every attempt is not flaky, even when the self-healing step tried and failed to repair a selector. An earlier version of the engine filed that case as FLAKY. Teams read FLAKY as "the test is bad" and move on, so a broken application shipped as noise. The engine no longer does that.
Do retries hide real bugs?
They can, and the reason is the number of states. With pass and fail only, a retry that passes turns a red run green, and an intermittent defect in the application becomes invisible. That is one of the most expensive kinds of bug to find. When red can also mean a blinked network or a dead browser, the team learns to rerun reds instead of reading them, and the suite loses trust.
Nexus Studio keeps the retry visible. A test that failed and then passed is not SECURE. It exits 0 in CI, so it does not break the build, but it shows in the report and in History as FLAKY, a pass with a reservation. See how the 14 verdicts are decided for the full order of checks.
What is the difference between FLAKY, BREACHED and BREACHED_FLAKY?
- FLAKY: failed, then a retry passed it, and nothing was repaired. Read the failed attempt's raw-data.json, then the History trend. It exits 0.
- BREACHED: the test failed on a single attempt, and no environment or harness pattern explains the error. An assertion failed. This is the bug report, and it exits 1.
- BREACHED_FLAKY: the test failed on its last attempt after an earlier retry, and no environment or harness pattern explains it. The failure itself is inconsistent. Compare the attempts before you file a bug. It exits 1.
Two neighbors keep the picture honest. AUTO_RECOVER is a pass where self-healing replaced a selector, so the author's locator is dead and the test is not clean. A first attempt that failed only on time and then passed on retry is SECURE, because that shows latency, not instability.
Before a failure can be BREACHED or BREACHED_FLAKY, the classifier rules out the other explanations in a fixed order: a harness error, a crash, a block, an unreachable host and a timeout. Only a failure that none of those explains is a finding. That is why a timeout or a CAPTCHA does not arrive as a bug report.
How are retries counted?
The engine reads the attempt number of the run. A passing test with a retry number above zero is FLAKY, unless a selector was healed, in which case it is AUTO_RECOVER. A failing or timed-out test with a retry number above zero, and no environment explanation, is BREACHED_FLAKY. A clean pass needs retry zero and no healed selector.
FLAKY means the attempts disagreed, and nothing else. A failure on a single attempt, with zero retries, is never labeled flaky. Each History row also records retry_attempt and attempts, so the count stays with the run.
What evidence does each verdict leave?
- FLAKY: screenshot.png, metadata.json and raw-data.json. The folder keeps the failed attempt's raw-data.json, so a retry that hid the failure still leaves its log. This follows the setting Keep evidence of a FLAKY / AUTO_RECOVER pass, which is on by default.
- BREACHED_FLAKY: report.md, metadata.json, raw-data.json and screenshot.png, and its failure video is kept by default.
- BREACHED: the same files, with the video kept by default.
The verdict is in the evidence folder's name, for example a name that ends in [BREACHED]. The Evidence page follows one failed test through every file.
How does History spot a test that flips?
History keeps one row per run for each test and browser or device, with the verdict, the retry attempt and the error signature. From those rows it shows trends by session, day or week: Regressed, Recovered, Still broken and Slower. A test that goes green, red and green again shows up there instead of disappearing into reruns.
The risk ranking, in the Evidence Locker and in Patrol, orders tests by how often they flip, break or cost time. It also shows a predicted risk from a logistic regression over your history, with an EWMA fallback when data is thin. Nothing is called flaky below five runs, because one or two runs cannot show a pattern.
Limits
- Retries are what separate FLAKY from SECURE and BREACHED_FLAKY from BREACHED. With retries off, an intermittent failure is a BREACHED.
- FLAKY says the attempts disagreed. It does not say why. Finding the cause in the application or in the test is still work for a person.
- The classification reads error text, so a custom assertion message that contains words such as
403orForbiddencan read as an environment problem. Open metadata.json when a verdict surprises you. - Below five runs, History and the risk ranking call nothing flaky.
Where to go next
The Evidence page shows every file a failed test leaves, and How Nexus Studio decides 14 test verdicts explains the order of the checks. The 30-day trial runs the same engine on your own tests.