
Capability
Stress Testing
Pressure-test answer fragility by removing evidence
Stress Testing is the systematic process of probing answer resilience by selectively removing evidence sources and measuring how the answer changes. It provides the quantitative foundation for fragility scores in the Trust Gradient and helps identify answers that might appear well-supported but are actually dependent on a single critical source.
The stress testing process works through controlled evidence removal experiments. For each claim in an answer, the system identifies all supporting evidence sources and then systematically removes them one at a time, re-running the synthesis to observe the impact. If removing a single source causes the claim to change or disappear, the claim is highly fragile. If removing any combination of sources up to a threshold still leaves the claim intact, it is resilient.
The minimum evidence set is the smallest subset of sources that is sufficient to sustain a given claim. Identifying this set is valuable because it reveals the true depth of evidence support behind each assertion. A claim might appear well-supported because it is cited by ten sources, but if nine of those sources are all derived from a single original study, the minimum evidence set is actually just one, and the claim is far more fragile than it appears.
Fragility scores are computed from the stress testing results and represent the inverse of evidence resilience. A score of zero means the claim survives the removal of all but one source. A high fragility score means the claim changes with the removal of a single source. These scores are surfaced through the Trust Gradient in the Proof Surface, giving users a quantitative measure of how robust each claim is against evidence changes.
Stress Testing feeds directly into the Coverage Labelling system. Claims with high fragility scores contribute to partial or thin coverage labels, alerting users that the answer rests on a narrow evidence base. The Coverage-Driven Sourcing system uses fragility data to prioritise ingestion requests, targeting domains and topics where additional sources would most improve answer resilience.
The stress testing process is resource-intensive, involving multiple re-runs of the retrieval and synthesis pipeline for each claim in each answer. This is why comprehensive stress testing is reserved for audit-mode assurance, where the additional cost is justified by the need for maximum evidence rigour. Verified-mode answers receive a lighter-weight fragility assessment based on source diversity metrics.
For the CueCrux platform, Stress Testing embodies the principle that transparency requires honest assessment of limitations. Rather than presenting all answers as equally reliable, the platform actively probes for weaknesses and communicates them clearly to users.





