Autonomous Offensive Agent Safety Scorecard
Score any offensive security vendor — including Sprocket — against the seven required properties of a safe autonomous offensive agent. Each property is assessed independently across three maturity levels. This is a measurement framework, not a certification: a high score means a vendor has assessed its controls against these properties, not that it has passed a threshold.
Score each property on evidence, not on assurances.
For each of the seven properties, pick the highest level the vendor can demonstrate. If they describe a control but cannot show you the mechanism, the artifact, or the log, score the level below. Each property carries an “evidence to ask for” prompt — use it as the question in the room.
The critical transition is Level 1 to Level 2: the shift from behavioral enforcement (an instruction in a prompt, which the agent can be argued out of) to structural enforcement (a control in code, outside the model’s decision loop, which it cannot bypass).
A behavioral control on data exfiltration means the agent probably will not exfiltrate data. A structural control means it cannot. When a vendor answers a safety question by quoting their system prompt, that is Level 1 — regardless of how well written the prompt is.
Seven properties.
Properties scored.
Per-property profile
A vendor's weakest property matters more than its average. One property at Level 1 caps what the whole system can be trusted with, because the properties are not independent — Governable is what keeps the other six enforceable when something unanticipated happens.