Skip to content

Dose-response: verified-leakage rate vs attack pressure (defensive curve) #36

Description

@bamdadd

Context

We report a single leakage rate per case, but a defensive question is how leakage responds to dose along an axis that already exists in the published patterns the suite ships — e.g. how many times the injected instruction is repeated, or the encoding depth in the encoded family. Plotting verified-leakage rate against that pressure is a robustness dose-response curve. This is explicitly not about making attacks more effective: it only varies existing published-pattern parameters and measures defensive response; no new attack potency, no tuning-to-fire.

Acceptance criteria

  • Add a harness that sweeps one existing, published-pattern axis (injection-repetition count and/or encoding depth) over a fixed set of levels, using cases already in the suite.
  • Emit verified-leakage rate (and hijack-ASR) per dose level with bootstrap 95% CIs, 3+ seeds, mean ± std; seeds/hardware/wall-clock recorded.
  • Produce a dose-response plot (leakage rate vs pressure), framed as a defensive robustness curve in the caption.
  • Body/docs state the guardrail: no new attack strings are introduced and nothing is tuned against a model's behaviour; the scorer stays the deterministic ground truth.

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions