DeepSweep
Initial benchmark · October 5, 2026

Permission checks need context

An AI-agent settings check should account for where a setting lives, which rules override it, and what the agent version supports. We tested DeepSweep against those questions and kept the failures in the results.

Brad McEvilly, founder of DeepSweep · Run completed 2026-10-05T16:25:35Z · Report version 1

What we measured: configuration diagnosis on 24 authored cases, using a source-level build of DeepSweep's 2.0.24 review engine. This evaluates DeepSweep's interpretation against a frozen documentation-based answer key. It does not measure native-agent behavior, attack prevention or the installed extension's end-to-end behavior.

The initial results

Each case had one stated criterion. All 24 remain in the accounting, including failures and unresolved judgments. Restrictive controls are negative only for their stated question; they are not declarations that a whole setup is safe.

Exposure cases: 1 met the criterion, 4 did not, 3 unresolved, out of 8; Restrictive / ignored-setting controls: 7 met the criterion, 2 did not, 2 unresolved, out of 11; Malformed configuration: 2 met the criterion, 0 did not, 0 unresolved, out of 2; Unsupported setting values: 0 met the criterion, 3 did not, 0 unresolved, out of 3
Counts on this finite authored set. The bars do not estimate performance across developers or real-world attacks.
Case groupCasesMet criterionDid not meetUnresolved
Exposure cases8143
Restrictive / ignored-setting controls11722
Malformed configuration2200
Unsupported setting values3030

A failure on an exposure case is a missed or incorrect diagnosis. A failure on a restrictive control is an incorrect assertion about the specific exposure. Invalid-input cases test whether the tool explains that it could not assess the configuration. These are different questions, so we publish them separately.

What the results change

Corrections will be checked against the same frozen cases and reported as a separate result with a new source revision. Once a case is known, passing it is a regression check; fresh cases are needed to test whether the improvement generalizes.

How we ran it

We authored 24 configuration bundles and their expected interpretations from the official documentation before collecting DeepSweep's outputs. The set contains 8 exposure cases, 11 restrictive or ignored-setting controls, 2 malformed inputs and 3 unsupported-value controls. A separate review role challenged the judgments before publication.

The engine ran locally with authored files in isolated folders. The runner did not call a model or execute Claude Code or Codex. The evaluation used the source revision associated with DeepSweep 2.0.24, compiled for a constrained harness; it did not run a downloaded VSIX or the editor UI. Historical benchmark suites are excluded from these counts.

Source revision: f2f46b0da7d1a4c8a82d73ab4e46691a8ff2b692. Case/oracle freeze: October 5, 16:18:52 UTC; source-target freeze: 16:20:14 UTC. Documentation retrieved October 5, 2026. Claude Code scope follows the documented change from version 2.1.257. Codex follows the retrieved documentation; no native binary version was tested.

Before publication, an applicability review found that four cases assumed project trust that the test setup had not supplied or established. Those judgments are now unresolved. The frozen inputs and raw outputs were not changed, and all 24 cases remain in the denominator.

Limits worth keeping visible

This is a vendor-conducted self-evaluation on a small, deliberately selected set. The reviewers worked within the same evaluation team. The cases are related and are not a random sample. We make no population-accuracy, competitive-superiority, certification, customer-growth or risk-reduction claim. The private corpus and implementation are not published, which limits outside reproduction; the public release contains methodology and aggregates only.

Check your own configuration

Start with your agent's current permission documentation and inspect which settings actually loaded. Use DeepSweep's local review as an additional check, with the limitations above in mind. Review suggested changes before applying them.

Choose your editor

Sources and downloads