Acceptance Criteria for a Scoped Adversarial Campaign Against an LLM Application: What Must Be Proven
By Mario Alexandre · July 18, 2026 · 10 min read
For a scoped adversarial campaign against an LLM application, an acceptance criteria decision begins with authorized application access, scope documentation, prohibited actions, and test data, together with a separate assignment of incident contacts. This acceptance criteria guide connects a scoped adversarial campaign against an LLM application to the workflow, evidence, named owners, failure handling, and catalog limits without promising a buyer-specific result.
The direct answer
Write a test for “scope and prohibited actions are signed off” before execution and keep “testing without written authorization” as a release-blocking counterexample.
For a scoped adversarial campaign against an LLM application, the relevant audience is teams that need attack evidence, not only a design checklist. The decision should cover authorization, threat modeling, attack-case selection, safe execution, per-attack evidence, control mapping, prioritization, and retest planning. The supplied boundary starts with authorized application access, scope documentation, prohibited actions, and test data, together with a separate assignment of incident contacts and ends with a threat model, per-attack evidence record, and prioritized control list, presented in reviewable form.
A scoped campaign cannot certify the system, prove the absence of unknown vulnerabilities, or replace broader application and infrastructure security testing.
Turn each requirement into a proof obligation
The expected deliverable is a threat model, per-attack evidence record, and prioritized control list.
| Acceptance statement | Observable evidence | Criterion-specific negative fixture | Evidence supplier | Acceptance adjudicator |
|---|---|---|---|---|
| “scope and prohibited actions are signed off” | a versioned normal, alternate, and failure-flow receipt with raw observed output for the statement “scope and prohibited actions are signed off” | For “scope and prohibited actions are signed off”, place a destructive synthetic action outside the authorized test matrix and omit reviewer sign-off, then require the sandboxed campaign preflight to block execution. | system owner | security observer |
| “attacks map to named threat hypotheses” | a versioned normal, alternate, and failure-flow receipt with raw observed output for the statement “attacks map to named threat hypotheses” | For “attacks map to named threat hypotheses”, label a synthetic test as data-exfiltration risk while its payload only changes response style, then require the threat mapping check to expose the mismatch. | red-team lead | security observer |
| “each result has reproducible evidence” | a source-to-claim trace with quoted support and a separate readback for the statement “each result has reproducible evidence” | For “each result has reproducible evidence”, create a synthetic finding that omits its payload, environment identity, and replay steps, then require evidence validation to identify every missing reproduction input. | data owner | security observer |
| “findings separate exploitability from impact” | a versioned normal, alternate, and failure-flow receipt with raw observed output for the statement “findings separate exploitability from impact” | For “findings separate exploitability from impact”, give a synthetic high-impact scenario an unreachable attack path but mark exploitability from severity alone, then require the rubric to surface the conflation. | data owner | security observer |
| “control fixes have retest cases” | a versioned normal, alternate, and failure-flow receipt with raw observed output for the statement “control fixes have retest cases” | For “control fixes have retest cases”, add a synthetic input-filter repair while omitting direct and encoded attack variants from the retest manifest, then require coverage comparison to reject the incomplete control check. | remediation owner | security observer |
The security observer adjudicates every pass, hold, or fail verdict against these registered statements.
Cover more than the happy path
The normal flow should establish whether “scope and prohibited actions are signed off” holds. An alternate flow should vary a permitted input while testing whether “attacks map to named threat hypotheses” holds. The failure flow should use a fixture demonstrating “successful prompts recorded without downstream impact evidence” and verify containment.
Add a recovery flow for “unsafe data or tools left in scope”.
Judge evidence quality and freshness
For a scoped adversarial campaign against an LLM application, a result from another environment cannot prove that “each result has reproducible evidence” holds in the buyer's environment.
Define pass, hold, and fail before execution
| Disposition | Meaning for this product | Required action |
|---|---|---|
| Pass | Current evidence establishes the applicable conditions, including “findings separate exploitability from impact” | The system owner may authorize the next bounded step |
| Hold | Evidence is missing, stale, mixed, or unable to rule on “testing without written authorization” | Name the absent proof and keep the current state |
| Fail | The observed result contradicts a required condition or exposes “controls recommended without a retest condition” | The remediation owner stops or rolls back the affected slice and requests an acceptance hold |
Keep sign-off independent
The implementer may produce artifacts, but the security observer should judge whether “control fixes have retest cases” holds against criteria written before the result was seen.
Record the business decision of the system owner, the technical evidence reviewed by the data owner, the acceptance verdict recorded by the security observer, and residual risk accepted by the remediation owner.
A screenshot or self-score cannot prove that “control fixes have retest cases” holds under the failure condition “controls recommended without a retest condition”.
Reopen criteria when the system changes
Changes to authorization, threat modeling, attack-case selection, safe execution, per-attack evidence, control mapping, prioritization, and retest planning can invalidate a test even when the requirement text stays the same.
How the sources bound the acceptance criteria decision
For a scoped adversarial campaign against an LLM application, the live catalog limits the offer to two elements. The supplied boundary is authorized application access, scope documentation, prohibited actions, and test data, together with a separate assignment of incident contacts. The catalog names the deliverable as a threat model, per-attack evidence record, and prioritized control list. It cannot establish whether “scope and prohibited actions are signed off” holds in the buyer's environment.
Connect those narrow roles to a local fixture for “attack cases copied from a checklist without system context” rather than treating citation status as a pass.
For a scoped adversarial campaign against an LLM application, limit the conclusion to the documented workflow and let the red-team lead retain the current source-to-claim map. Keep the source decision provisional while the failure case “unsafe data or tools left in scope” remains unresolved.
Product-specific acceptance criteria review drills
These drills connect a scoped adversarial campaign against an LLM application to concrete inputs, failures, acceptance statements, and owners. For a scoped adversarial campaign against an LLM application, the drills map each criterion to a reviewable verdict.
Acceptance for a scoped adversarial campaign against an LLM application is judged against a boundary record covering authorized application access, scope documentation, prohibited actions, and test data, never live protected material. Incident contacts are assigned separately from material custody. The system owner requires synthetic, non-secret cases; messages, writes, state changes, and all other external effects stay inside the fixture throughout and after each case.
Requirement trace
Frame the requirement trace review around “testing without written authorization”. Before testing a response, the system owner captures the input, decision boundary, and residual state.
Pair a scope record covering authorized application access, scope documentation, prohibited actions, and test data, together with a separate assignment of incident contacts with a direct observation of whether “scope and prohibited actions are signed off” holds. The red-team lead retains the source and result together.
The security observer bases the outcome for the requirement trace review on “scope and prohibited actions are signed off” and keeps a threat model, per-attack evidence record, and prioritized control list bounded to that finding. In the requirement trace review, evidence for “scope and prohibited actions are signed off” maps support to pass, contradiction to fail, and unresolved to hold.
Repeat the judgment when the workflow boundary for authorization, threat modeling, attack-case selection, safe execution, per-attack evidence, control mapping, prioritization, and retest planning adds a new handoff or removes the rollback state used in the test.
Normal-flow result
Use the normal-flow result review to examine what follows from the failure case “attack cases copied from a checklist without system context”. Before intervention, the red-team lead retains the observable handoff.
Give the data owner an authorized, read-only boundary record covering authorized application access, scope documentation, prohibited actions, and test data, together with a separate assignment of incident contacts plus the criterion “each result has reproducible evidence”. Their receipt identifies any missing proof.
The security observer accepts, rejects, or returns the evidence for “each result has reproducible evidence”. Completion of another condition cannot substitute for it. In the normal-flow result review, evidence for “each result has reproducible evidence” maps support to pass, contradiction to fail, and unresolved to hold.
Expire the disposition if the red-team lead cannot reproduce the case for “attack cases copied from a checklist without system context” under the recorded authority.
Alternate-flow result
Represent the failure case “successful prompts recorded without downstream impact evidence” explicitly in the alternate-flow result review. The data owner captures the relevant input, action, and residual condition.
Use a scope record covering authorized application access, scope documentation, prohibited actions, and test data, together with a separate assignment of incident contacts to reproduce the case and inspect whether “control fixes have retest cases” holds. Store the comparison under the alternate-flow result review, not in operator memory.
The security observer closes the alternate-flow result review with a bounded ruling on “control fixes have retest cases”. The ruling does not certify untested behavior in a threat model, per-attack evidence record, and prioritized control list. In the alternate-flow result review, evidence for “control fixes have retest cases” maps support to pass, contradiction to fail, and unresolved to hold.
A new dependency, owner, or instance of “successful prompts recorded without downstream impact evidence” expires the evidence for the alternate-flow result review and requires a focused rerun.
Failure-flow result
Model the failure-flow result review with a safe fixture involving “unsafe data or tools left in scope”. The data owner names the affected action and its permitted consequence.
Use “attacks map to named threat hypotheses” as the explicit criterion for a case drawn from the boundary covering authorized application access, scope documentation, prohibited actions, and test data, together with a separate assignment of incident contacts. The resulting receipt belongs to the remediation owner.
The security observer moves forward only after the record supports the finding “attacks map to named threat hypotheses”. Conflicting evidence makes the security observer record fail and preserve the prior state. In the failure-flow result review, evidence for “attacks map to named threat hypotheses” maps support to pass, contradiction to fail, and unresolved to hold.
Recheck the drill when the operating path no longer matches authorization, threat modeling, attack-case selection, safe execution, per-attack evidence, control mapping, prioritization, and retest planning or when the rollback evidence expires.
Independent verdict
The independent verdict review starts with the failure case “controls recommended without a retest condition”. Its first owner is the remediation owner, who captures the current workflow state without changing it.
Create a versioned boundary record covering authorized application access, scope documentation, prohibited actions, and test data, together with a separate assignment of incident contacts, then test whether “findings separate exploitability from impact” holds; keep the case result with its exact input identity.
The security observer links the finding “findings separate exploitability from impact” to go, revise, or stop in the decision record. It does not treat completion of a threat model, per-attack evidence record, and prioritized control list as proof of every outcome. In the independent verdict review, evidence for “findings separate exploitability from impact” maps support to pass, contradiction to fail, and unresolved to hold.
Reopen the case if the operating response to “controls recommended without a retest condition” changes, even when the title and stated requirement remain the same.
Evidence expiry
Create a safe fixture for “testing without written authorization” and attach it to the evidence expiry review. The system owner observes the relevant part of authorization, threat modeling, attack-case selection, safe execution, per-attack evidence, control mapping, prioritization, and retest planning.
Compare the candidate result with a frozen scope record covering authorized application access, scope documentation, prohibited actions, and test data, together with a separate assignment of incident contacts for “scope and prohibited actions are signed off”. Preserve both sides of the comparison.
The security observer closes the evidence expiry review only when the record resolves “scope and prohibited actions are signed off”; otherwise the listed deliverable remains provisional. In the evidence expiry review, evidence for “scope and prohibited actions are signed off” maps support to pass, contradiction to fail, and unresolved to hold.
Return the record to hold when the fixture, dependency, or permission used to judge whether “scope and prohibited actions are signed off” holds changes materially.
Frequently asked question
What acceptance criteria should I use for LLM Security Red-Team?
Require observable evidence that scope and prohibited actions are signed off and include “testing without written authorization” as a negative case. The security observer should record pass, hold, or fail before expansion.
A product bridge, with a boundary
The LLM Security Red-Team is the relevant sincLLM offer for this narrow problem. The frozen live catalog describes its required boundary as authorized application access, scope documentation, prohibited actions, and test data, together with a separate assignment of incident contacts and its deliverable as a threat model, per-attack evidence record, and prioritized control list. The buyer must judge fit and results in its own environment; the catalog does not certify compliance, safety, or technical sufficiency.
Sources and claim boundaries
- sincLLM product catalog: The bounded product description, required inputs, stated deliverable, and product bridge.
- MITRE ATLAS: A living knowledge base of adversary tactics and techniques involving AI-enabled systems.
- NIST AI Risk Management Framework: A voluntary, use-case-agnostic framework for governing, mapping, measuring, and managing AI risk.
The source list constrains what the article may claim and cannot substitute for tests, readbacks, or accountable review in the target environment.