---
name: Jules adversarial review
description: >-
  Use when an implementer agent (Codex, Claude, Grok, Gemini or a human) has
  pushed a commit and you want an independent, read-only second opinion from
  Google Jules before merge: correctness, quality or defensive security review
  lanes via the Jules REST API, with plan vetting, refusal handling, findings
  normalisation and read-only compliance checks. Not for having Jules write code
  or open PRs.
---
# Jules adversarial review

Use Google Jules as an **independent, read-only reviewer** of code that another agent wrote. You get a different model's view, delivered as structured findings that a gate (evaluator, tests, human) decides on. Jules never changes the repository in this workflow.

CLI: `scripts/jules_review.py` (relative to this skill folder), absolute path `/home/box/agent-data/workflows/jules-adversarial-review/scripts/jules_review.py`. Python 3 stdlib only. Unit tests: `scripts/tests/` (run `python3 -m unittest discover -s scripts/tests` from the skill folder).

## 1. Prerequisites

- `JULES_API_KEY` in the environment (create it at jules.google.com, Settings, API keys). Never echo, print or commit it. The CLI reads it and never logs it.
- The target GitHub repo is **connected as a Jules source**. Check with `python3 scripts/jules_review.py sources` (lists `sources/github/<owner>/<repo>` and each default branch).
- The code to review is **pushed** to a branch Jules can see. Pin the commit SHA you mean (`--commit`).
- Optional: `gh` authenticated for the GitHub-side compliance check (`--gh`).
- Keep extra Jules MCP integrations off for review work, apart from a docs MCP such as Context7. Every connected integration gives the reviewer more reach.

## 2. Lanes: which to run

| Lane | Ask | Use when | Notes |
|---|---|---|---|
| `correctness` (primary) | Defects, wrong assumptions, error handling, state, races, leaks, dead code, API misuse, behaviour that contradicts the docs | Always, for any non-trivial change | Strongest lane in testing. Took 6 to 16 min to report |
| `quality` | README gaps, regression risks, architecture, tests, browser/a11y, edge cases | Features, UI, refactors | Fast. Jules may skip planning entirely |
| `security` | Defensive security review **with an ownership + authorisation + defensive-purpose preamble** | Anything exposed: HTTP handlers, auth, file serving, user input | Without the preamble, "adversarial security review" prompts were refused. Keep the preamble verbatim |
| scanner (not Jules) | Semgrep / CodeQL / dependency audit | Always for security-sensitive repos, and **whenever the security lane is refused** | Deterministic second source |

All prompts are vetted templates inside the CLI. They include a strict read-only clause (no file edits, no scratch files, no commit, push, publish or PR), an optional `--context` focus (diff summary, feature description), the target branch/commit, and a parseable output format ending in `Confidence: x`.

## 3. Run it

```bash
S=scripts/jules_review.py          # or the absolute path above
# one lane, end to end (start, watch, report):
python3 $S run --repo OWNER/REPO --branch BRANCH --commit SHA --lane correctness \
  --context "path/to/diff-summary.md or a sentence about what changed" --out ./jules-runs/$(date +%Y%m%d-%H%M) --gh
# all lanes in parallel + summary.md with a verdict suggestion:
python3 $S matrix --repo OWNER/REPO --branch BRANCH --commit SHA --lanes correctness,quality,security --out ./jules-runs/X --gh
# piecewise:
python3 $S start --repo OWNER/REPO --branch BRANCH --lane quality         # prints session_id + URL
python3 $S watch SESSION_ID --timeout-min 30                              # exit 0 ok, 2 REFUSED, 3 NEEDS_HUMAN, 4 FAILED, 5 TIMEOUT
python3 $S report SESSION_ID --out DIR --gh                               # session.json, activities.json, review.md, findings.json
```

Run long jobs in the background and check on them. A lane takes about 2 to 20 min. Sessions sit in COMPLETED about 7 min after the report arrives; `--settle-min 3` lets `watch` return `REPORTED` early once the report has been idle.

What `watch` does on each 30 s poll:
1. Plan generated: keyword check for write steps (modify/edit/create files/commit/push/PR/submit/"pre-commit"), ignoring negated mentions ("No modifications ... will be made"). If the plan is read-only it approves it. Otherwise it sends the standard revise message, at most 2 times, then stops with **NEEDS_HUMAN**. The check is deliberately conservative: Jules' boilerplate "Complete pre-commit steps" triggers one revision.
2. Jules waiting for feedback: one short read-only nudge (max 2), then NEEDS_HUMAN. It never widens scope.
3. Last `agentMessaged` matches refusal phrasing and has no findings: **REFUSED**.
4. State COMPLETED/FAILED or timeout: done.

`findings.json` holds `{lane, session_id, url, status, refused, confidence, findings:[{category, severity, file, line, title, reasoning, fix}], clear_categories, compliance:{patch_present, patch_files, pr_present, github}, raw_report}`. Parsing is best-effort across Jules' varying layouts, and the raw text is always kept.

## 4. Verify before you trust (mandatory)

Jules reports read well but are not ground truth.
- **Check every high/critical finding against a local clone at the same SHA.** Open the cited file and lines. Line numbers are sometimes **off by one or two**. A correct conclusion can come with a wrong explanation.
- Drop speculative findings ("if X were None..." when the code guarantees it never is).
- Treat the self-reported **confidence (often 0.95) as inflated**. Confidence is decided by verification, not by Jules.
- Look for what it **missed**: different lanes catch different bugs, so the union of lanes plus a scanner beats any single lane.
- The matrix verdict (`FAIL` if any high/critical, `INCOMPLETE` if a lane is missing, else `PASS`) is a **suggestion** for the evaluator or human gate. It never merges anything.

## 5. Never

- Never click **Publish branch / Create PR** in the Jules UI for a review session, and never set `automationMode`. Even read-only sessions can carry a patch: Jules may write scratch probe scripts in its VM, and they appear in `session.outputs` with a suggested commit message. Record it under compliance and leave it unpublished.
- Never auto-merge on Jules' say-so. Fixes go back to the implementer, who pushes a new commit, and the review reruns.
- Never paste the API key into prompts, logs, chats or repos.

## 6. Refusals

If a lane is refused (often security-flavoured wording without ownership context):
1. **Do not disguise or reword the request to slip past the refusal.** Record it verbatim (`review.md`, status REFUSED).
2. Route that review axis to a deterministic scanner (Semgrep, CodeQL, dependency audit), a different reviewer, or a human.
3. Only the vetted security template, with its honest ownership/defensive preamble, is used up front. That framing is legitimate context, not an evasion. It is not something to escalate after a refusal.

## 7. Jules API quirks (v1alpha, observed)

Base `https://jules.googleapis.com/v1alpha`, header `X-Goog-Api-Key`.
- `requirePlanApproval` is not echoed by create or GET, and Jules **may skip the plan** and answer directly.
- Create response lacks `state`/`createTime`. `GET .../activities` returns **404** on a brand-new session and `{}` when there are none; treat both as an empty list.
- Refusal is a normal `agentMessaged`; the state goes IN_PROGRESS then COMPLETED (not FAILED).
- `sessionCompleted` activity is often absent. Use `session.state`. The final report = **last `agentMessaged`**.
- `:approvePlan` / `:sendMessage` return `{}`. One approval creates **two `planApproved` activities** with the same `planId` (deduplicate). A revise via `:sendMessage` yields a new `planGenerated` with a new plan id. The first plan step omits `index`.
- Every `progressUpdated` repeats the **cumulative** `changeSet.gitPatch`. A gitPatch with only `baseCommitId` = no change. `session.outputs` appears only for a non-empty patch (adds `suggestedCommitMessage`) or a PR.
- Planning time ranges from 0 to about 10 min. Poll every 30 s with a budget of 20 to 35 min.
- The API is alpha. Keep all calls inside the CLI/adapter so a breaking change touches one file.

## 8. Examples

- After an implementer pushes commit `abc123` to `feature/x` of `OWNER/REPO`: run `matrix` with all three lanes, read `summary.md`, verify each high finding in a clone, then send confirmed findings back to the implementer.
- Small fix: `run --lane correctness` only, plus the repo's own tests.
- Earlier boundary tests (example repo `OWNER/test-repo`): the bare "adversarial security review" prompt was refused, while quality, correctness and the ownership-stated security prompt were accepted. An early correctness prompt left scratch probe scripts in the session outputs. After the stricter "no scratch files" clause was added, the CLI's own correctness template produced a clean session: no patch, exact line references, and a report in the requested format that parsed with no errors. Nothing reached GitHub in any test.
