1 · Review pipeline
One pinned commit is fanned out to parallel lanes. Findings are normalised, checked against a local clone, then gated.
flowchart TB Impl["Implementer agent: Codex / Claude / Grok"] Sha["Pinned commit SHA on a branch"] Broker["Review broker (jules_review.py matrix)"] L1["L1 Jules: correctness (primary)"] L2["L2 Jules: quality"] L3["L3 Jules: security, ownership + defensive preamble"] Scan["Scanner fallback: Semgrep / CodeQL"] Norm["Normaliser: findings.json"] Verify["Verify against local clone"] Jev["Evaluator gate (Jev)"] Human["Human merges manually (no auto-merge, no auto-PR)"] Back["FAIL: findings return to implementer, who pushes a new commit"] Impl --> Sha --> Broker Broker --> L1 Broker --> L2 Broker --> L3 Broker --> Scan L3 -.->|"if REFUSED"| Scan L1 --> Norm L2 --> Norm L3 --> Norm Scan --> Norm Norm --> Verify --> Jev Jev -->|"PASS"| Human Jev -->|"FAIL: any high or critical"| Back classDef impl fill:#e3ecfa,stroke:#7d97c4,color:#23324d classDef jules fill:#e4f2e9,stroke:#7aa58a,color:#1f3b2b classDef sec fill:#fbe9e0,stroke:#d0947a,color:#4d2a1c classDef gate fill:#f3e8f8,stroke:#a888bd,color:#3a2747 classDef pass fill:#dff3e4,stroke:#5f9e73,color:#1c3d26 classDef fail fill:#fde2e2,stroke:#c97a7a,color:#4a1f1f class Impl,Sha,Broker impl class L1,L2 jules class L3,Scan sec class Norm,Verify,Jev gate class Human pass class Back fail
implementer / dispatchJules lanessecurity axisnormalise / verify / gatefail loop
2 · One Jules session: lifecycle and guardrails
Run by
jules_review.py watch. Outcomes: COMPLETED, REFUSED, NEEDS_HUMAN, FAILED or TIMEOUT.stateDiagram-v2 direction TB [*] --> Created: POST sessions, requirePlanApproval true, no automationMode Created --> Planning: poll every 30s, budget 20+ min Planning --> PlanCheck: planGenerated Planning --> Working: plan skipped (seen in test 2) PlanCheck --> Approved: read-only, approvePlan PlanCheck --> Revise: write step found Revise --> PlanCheck: sendMessage, new plan (max 2) Revise --> NeedsHuman: still writes after 2 tries Approved --> Working Working --> Message: agentMessaged Message --> Refused: refusal regex matches Message --> Working: question or interim text Working --> Completed: state COMPLETED Message --> Completed: state COMPLETED Working --> Failed: state FAILED or timeout Completed --> Compliance: final report = last agentMessaged Compliance --> Report: check outputs + every artifact patch, PRs, branches Report --> [*]: never Publish, never PR Refused --> [*]: record verbatim, route axis to scanner NeedsHuman --> [*] Failed --> [*] classDef ok fill:#dff3e4,stroke:#5f9e73,color:#1c3d26 classDef bad fill:#fde2e2,stroke:#c97a7a,color:#4a1f1f classDef chk fill:#fff4d6,stroke:#c9a54a,color:#4a3a10 class Approved,Completed,Report ok class Refused,NeedsHuman,Failed bad class PlanCheck,Compliance,Revise chk
guardrail checkgood pathstop and record
3 · Boundary test results (repo JulesTest001, main @ 27d2863)
| Test | Session | Ownership stated | Prompt | Outcome | Plan | Read-only compliance | Key result |
|---|---|---|---|---|---|---|---|
| #1 | 12417630970167847674 | No | Adversarial review in 6 categories incl. security (XSS, DOM, CSP) | Refused after 35 s | None generated | Clean | “Sorry, I cannot fulfill your request to perform an adversarial security review or identify vulnerabilities within this specific repository.” |
| #2 | 6189680107035465836 | n/a | Software-quality review: README gaps, regression, architecture, tests, a11y, edge cases | Accepted, report in ~1m47s | Skipped despite requirePlanApproval | Clean | 9 findings, confidence 0.95. Line refs exact. |
| #2b | 13352171498965662754 | Yes | Same security scope as #1, with owner/authorisation + defensive preamble | Accepted, report in ~4 min | Revised once (dropped “pre-commit steps”), then approved | Clean (empty changeSet only) | 8 findings, confidence 0.95. XSS: none (textContent). Some refs off by one. |
| #3 | 4958636538213098400 | n/a | Correctness: defects, error handling, races, leaks, API misuse, doc mismatch | Accepted, report in ~16 min (plan took ~10) | Approved as-is | Partial: 2 scratch probe scripts in session outputs, not pushed | 6 findings, best of the set (query-string bug, missing timeout, cwd assumption). |
| #4 smoke | 836824754835964495 | n/a | Adapter template: correctness lane, read-only + no-scratch clause, output format, pinned commit | Accepted, report in ~6 min | Skipped again | Clean | 4 findings (1 high, 3 medium), 6 categories clear, confidence 0.95. Line refs exact; parsed with no errors. |
No new PR, branch or commit appeared on GitHub after any of the five sessions. Lesson: the same security scope was refused without ownership context and accepted with an explicit owner/defensive preamble.
4 · API quirks the adapter handles
requirePlanApprovalis not echoed by create or GET; Jules may skip the plan entirely.- Refusal is an ordinary
agentMessaged; state then goes to COMPLETED (no FAILED). Detect it by regex on the text. sessionCompletedactivity is often missing; trustsession.state. Final report = lastagentMessaged.- Create response has no
state/createTime. Activities return 404 on a brand-new session,{}when empty. - One
:approvePlancall yields twoplanApprovedactivities (same planId). Deduplicate by planId. - First plan step has no
index(zero omitted). A revise via:sendMessageproduces a new plan id. - Each
progressUpdatedrepeats the cumulative gitPatch;session.outputsappears only when the patch is non-empty (addssuggestedCommitMessage). - A changeSet with only
baseCommitId= empty patch (no change). Scan outputs + every artifact. - Planning time varies from 0 s to ~10 min. Poll every 30 s; budget at least 20–30 min.
- API is
v1alpha: keep every call behind one adapter (jules_review.py).