Jules skills · Jules adversarial review

Jules adversarial review

Use when an implementer agent (Codex, Claude, Grok, Gemini or a human) has pushed a commit and you want an independent, read-only second opinion from Google Jules before merge: correctness, quality or defensive security review lanes via the Jules REST API, with plan vetting, refusal handling, findings normalisation and read-only compliance checks. Not for having Jules write code or open PRs.

Human twin this page · AI twin raw SKILL.md (same text, unrendered, for agents)

Folder in the private repo Kuglys/skills-jules: skills/jules-adversarial-review/ → SKILL.md, scripts/jules_review.py, scripts/tests/test_jules_review.py. Scripts and test fixtures are not published on this site.

Diagram

Open the pipeline and session-lifecycle diagram →

Diagram snapshot: review pipeline and Jules session lifecycle

Use Google Jules as an independent, read-only reviewer of code that another agent wrote. You get a different model's view, delivered as structured findings that a gate (evaluator, tests, human) decides on. Jules never changes the repository in this workflow.

CLI: scripts/jules_review.py (relative to this skill folder), absolute path /home/box/agent-data/workflows/jules-adversarial-review/scripts/jules_review.py. Python 3 stdlib only. Unit tests: scripts/tests/ (run python3 -m unittest discover -s scripts/tests from the skill folder).

1. Prerequisites

2. Lanes: which to run

Lane Ask Use when Notes
correctness (primary) Defects, wrong assumptions, error handling, state, races, leaks, dead code, API misuse, behaviour that contradicts the docs Always, for any non-trivial change Strongest lane in testing. Took 6 to 16 min to report
quality README gaps, regression risks, architecture, tests, browser/a11y, edge cases Features, UI, refactors Fast. Jules may skip planning entirely
security Defensive security review with an ownership + authorisation + defensive-purpose preamble Anything exposed: HTTP handlers, auth, file serving, user input Without the preamble, "adversarial security review" prompts were refused. Keep the preamble verbatim
scanner (not Jules) Semgrep / CodeQL / dependency audit Always for security-sensitive repos, and whenever the security lane is refused Deterministic second source

All prompts are vetted templates inside the CLI. They include a strict read-only clause (no file edits, no scratch files, no commit, push, publish or PR), an optional --context focus (diff summary, feature description), the target branch/commit, and a parseable output format ending in Confidence: x.

3. Run it

S=scripts/jules_review.py          # or the absolute path above
# one lane, end to end (start, watch, report):
python3 $S run --repo OWNER/REPO --branch BRANCH --commit SHA --lane correctness \
  --context "path/to/diff-summary.md or a sentence about what changed" --out ./jules-runs/$(date +%Y%m%d-%H%M) --gh
# all lanes in parallel + summary.md with a verdict suggestion:
python3 $S matrix --repo OWNER/REPO --branch BRANCH --commit SHA --lanes correctness,quality,security --out ./jules-runs/X --gh
# piecewise:
python3 $S start --repo OWNER/REPO --branch BRANCH --lane quality         # prints session_id + URL
python3 $S watch SESSION_ID --timeout-min 30                              # exit 0 ok, 2 REFUSED, 3 NEEDS_HUMAN, 4 FAILED, 5 TIMEOUT
python3 $S report SESSION_ID --out DIR --gh                               # session.json, activities.json, review.md, findings.json

Run long jobs in the background and check on them. A lane takes about 2 to 20 min. Sessions sit in COMPLETED about 7 min after the report arrives; --settle-min 3 lets watch return REPORTED early once the report has been idle.

What watch does on each 30 s poll: 1. Plan generated: keyword check for write steps (modify/edit/create files/commit/push/PR/submit/"pre-commit"), ignoring negated mentions ("No modifications ... will be made"). If the plan is read-only it approves it. Otherwise it sends the standard revise message, at most 2 times, then stops with NEEDS_HUMAN. The check is deliberately conservative: Jules' boilerplate "Complete pre-commit steps" triggers one revision. 2. Jules waiting for feedback: one short read-only nudge (max 2), then NEEDS_HUMAN. It never widens scope. 3. Last agentMessaged matches refusal phrasing and has no findings: REFUSED. 4. State COMPLETED/FAILED or timeout: done.

findings.json holds {lane, session_id, url, status, refused, confidence, findings:[{category, severity, file, line, title, reasoning, fix}], clear_categories, compliance:{patch_present, patch_files, pr_present, github}, raw_report}. Parsing is best-effort across Jules' varying layouts, and the raw text is always kept.

4. Verify before you trust (mandatory)

Jules reports read well but are not ground truth. - Check every high/critical finding against a local clone at the same SHA. Open the cited file and lines. Line numbers are sometimes off by one or two. A correct conclusion can come with a wrong explanation. - Drop speculative findings ("if X were None..." when the code guarantees it never is). - Treat the self-reported confidence (often 0.95) as inflated. Confidence is decided by verification, not by Jules. - Look for what it missed: different lanes catch different bugs, so the union of lanes plus a scanner beats any single lane. - The matrix verdict (FAIL if any high/critical, INCOMPLETE if a lane is missing, else PASS) is a suggestion for the evaluator or human gate. It never merges anything.

5. Never

6. Refusals

If a lane is refused (often security-flavoured wording without ownership context): 1. Do not disguise or reword the request to slip past the refusal. Record it verbatim (review.md, status REFUSED). 2. Route that review axis to a deterministic scanner (Semgrep, CodeQL, dependency audit), a different reviewer, or a human. 3. Only the vetted security template, with its honest ownership/defensive preamble, is used up front. That framing is legitimate context, not an evasion. It is not something to escalate after a refusal.

7. Jules API quirks (v1alpha, observed)

Base https://jules.googleapis.com/v1alpha, header X-Goog-Api-Key. - requirePlanApproval is not echoed by create or GET, and Jules may skip the plan and answer directly. - Create response lacks state/createTime. GET .../activities returns 404 on a brand-new session and {} when there are none; treat both as an empty list. - Refusal is a normal agentMessaged; the state goes IN_PROGRESS then COMPLETED (not FAILED). - sessionCompleted activity is often absent. Use session.state. The final report = last agentMessaged. - :approvePlan / :sendMessage return {}. One approval creates two planApproved activities with the same planId (deduplicate). A revise via :sendMessage yields a new planGenerated with a new plan id. The first plan step omits index. - Every progressUpdated repeats the cumulative changeSet.gitPatch. A gitPatch with only baseCommitId = no change. session.outputs appears only for a non-empty patch (adds suggestedCommitMessage) or a PR. - Planning time ranges from 0 to about 10 min. Poll every 30 s with a budget of 20 to 35 min. - The API is alpha. Keep all calls inside the CLI/adapter so a breaking change touches one file.

8. Examples