Permission checks need context
An AI-agent settings check should account for where a setting lives, which rules override it, and what the agent version supports. We tested DeepSweep against those questions and kept the failures in the results.
The initial results
Each case had one stated criterion. All 24 remain in the accounting, including failures and unresolved judgments. Restrictive controls are negative only for their stated question; they are not declarations that a whole setup is safe.

| Case group | Cases | Met criterion | Did not meet | Unresolved |
|---|---|---|---|---|
| Exposure cases | 8 | 1 | 4 | 3 |
| Restrictive / ignored-setting controls | 11 | 7 | 2 | 2 |
| Malformed configuration | 2 | 2 | 0 | 0 |
| Unsupported setting values | 3 | 0 | 3 | 0 |
A failure on an exposure case is a missed or incorrect diagnosis. A failure on a restrictive control is an incorrect assertion about the specific exposure. Invalid-input cases test whether the tool explains that it could not assess the configuration. These are different questions, so we publish them separately.
What the results change
- Three exposure cases lacked the required diagnosis. A fourth detected the broad exposure but overstated its reach, so it also failed the predeclared criterion. Correct explanations matter alongside detection.
- Two restrictive or ignored-setting controls produced incorrect effective-permission claims. The corrective work is to resolve setting scope and precedence before describing what an agent can do.
- Four cases are unresolved because our test setup assumed a trusted project without supplying or establishing that trust state. We cannot count those as successful diagnoses or engine failures. This is a limitation of our evaluation.
- One additional restrictive-control judgment remains unresolved: the output avoided the disputed warning but did not acknowledge a narrow permission named in the criterion. We retained that ambiguity rather than count it as a pass.
- Both malformed inputs were identified in the captured diagnostics. None of the three unsupported values received the required explanation that the setting could not be assessed.
- In two confined-automation controls, the skipped-prompt warning did not assert that sandbox boundaries were removed. That absence of a false assertion met the stated criterion; it does not establish a full model of effective permissions.
Corrections will be checked against the same frozen cases and reported as a separate result with a new source revision. Once a case is known, passing it is a regression check; fresh cases are needed to test whether the improvement generalizes.
How we ran it
We authored 24 configuration bundles and their expected interpretations from the official documentation before collecting DeepSweep's outputs. The set contains 8 exposure cases, 11 restrictive or ignored-setting controls, 2 malformed inputs and 3 unsupported-value controls. A separate review role challenged the judgments before publication.
The engine ran locally with authored files in isolated folders. The runner did not call a model or execute Claude Code or Codex. The evaluation used the source revision associated with DeepSweep 2.0.24, compiled for a constrained harness; it did not run a downloaded VSIX or the editor UI. Historical benchmark suites are excluded from these counts.
Source revision: f2f46b0da7d1a4c8a82d73ab4e46691a8ff2b692. Case/oracle freeze: October 5, 16:18:52 UTC; source-target freeze: 16:20:14 UTC. Documentation retrieved October 5, 2026. Claude Code scope follows the documented change from version 2.1.257. Codex follows the retrieved documentation; no native binary version was tested.
Before publication, an applicability review found that four cases assumed project trust that the test setup had not supplied or established. Those judgments are now unresolved. The frozen inputs and raw outputs were not changed, and all 24 cases remain in the denominator.
Limits worth keeping visible
This is a vendor-conducted self-evaluation on a small, deliberately selected set. The reviewers worked within the same evaluation team. The cases are related and are not a random sample. We make no population-accuracy, competitive-superiority, certification, customer-growth or risk-reduction claim. The private corpus and implementation are not published, which limits outside reproduction; the public release contains methodology and aggregates only.
Check your own configuration
Start with your agent's current permission documentation and inspect which settings actually loaded. Use DeepSweep's local review as an additional check, with the limitations above in mind. Review suggested changes before applying them.
Choose your editor