mirror of
https://github.com/obra/superpowers
synced 2026-08-05 01:06:06 +00:00
Compare commits
19 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 36f3883f4e | |||
| 05c2393b82 | |||
| 78cc189244 | |||
| be76350536 | |||
| 419dec7755 | |||
| 2b195749df | |||
| 8acf8e5f24 | |||
| 50a924b0c4 | |||
| 2a977c7095 | |||
| 50f787ca5c | |||
| 61f669ebc9 | |||
| e7a4285985 | |||
| e9686d5c09 | |||
| d8189d1587 | |||
| db4538fcb8 | |||
| 9b8b14fe12 | |||
| 7c560e048b | |||
| 75756d2900 | |||
| 2e7d681591 |
@@ -21,7 +21,7 @@
|
|||||||
"workflow"
|
"workflow"
|
||||||
],
|
],
|
||||||
"skills": "./skills/",
|
"skills": "./skills/",
|
||||||
"hooks": {},
|
"hooks": "./hooks/hooks-codex.json",
|
||||||
"interface": {
|
"interface": {
|
||||||
"displayName": "Superpowers",
|
"displayName": "Superpowers",
|
||||||
"shortDescription": "Planning, TDD, debugging, and delivery workflows for coding agents",
|
"shortDescription": "Planning, TDD, debugging, and delivery workflows for coding agents",
|
||||||
|
|||||||
@@ -92,6 +92,22 @@ Superpowers is available via the [official Codex plugin marketplace](https://git
|
|||||||
|
|
||||||
- Select `Install Plugin`.
|
- Select `Install Plugin`.
|
||||||
|
|
||||||
|
#### Codex: compaction re-injection hook
|
||||||
|
|
||||||
|
Codex compacts long sessions, replacing the transcript with a summary that
|
||||||
|
drops Superpowers' skill instructions mid-run — long autonomous workflows
|
||||||
|
(like subagent-driven-development) then drift back to harness defaults.
|
||||||
|
Claude Code re-injects the bootstrap after every compaction; the plugin ships
|
||||||
|
a SessionStart hook (`hooks/hooks-codex.json`) that restores the same
|
||||||
|
behavior on Codex (0.145+). It fires only on post-compaction re-starts
|
||||||
|
(`source: "compact"`) and is silent at normal session start.
|
||||||
|
|
||||||
|
The hook installs with the plugin — no configuration needed. Codex asks you
|
||||||
|
to review and trust it once, the first time it loads after install or update.
|
||||||
|
Headless automation (CI, eval harnesses) must pass
|
||||||
|
`--dangerously-bypass-hook-trust` instead, because untrusted hooks are
|
||||||
|
skipped silently.
|
||||||
|
|
||||||
### Cursor
|
### Cursor
|
||||||
|
|
||||||
- In Cursor Agent chat, install from marketplace:
|
- In Cursor Agent chat, install from marketplace:
|
||||||
|
|||||||
@@ -237,10 +237,12 @@ nesting differ per harness**.
|
|||||||
- Manifests: `.cursor-plugin/plugin.json` is the Shape A manifest example that
|
- Manifests: `.cursor-plugin/plugin.json` is the Shape A manifest example that
|
||||||
points the harness at `./skills/` and the right `hooks-*.json`. Claude Code's
|
points the harness at `./skills/` and the right `hooks-*.json`. Claude Code's
|
||||||
`.claude-plugin/plugin.json` sets neither field — it auto-discovers `skills/`
|
`.claude-plugin/plugin.json` sets neither field — it auto-discovers `skills/`
|
||||||
and `hooks/hooks.json` by convention. Do **not** copy Codex's
|
and `hooks/hooks.json` by convention. Codex's `.codex-plugin/plugin.json`
|
||||||
`.codex-plugin/plugin.json` for Shape A: it declares an empty `hooks` object
|
points `hooks` at `./hooks/hooks-codex.json` — a compaction-only hook, not a
|
||||||
specifically to suppress Codex's `hooks/hooks.json` auto-discovery, because
|
bootstrap injector: Codex surfaces skills natively at session start, so its
|
||||||
Codex surfaces skills natively and runs no session-start hook.
|
hook fires only on post-compaction re-starts. The explicit pointer also
|
||||||
|
suppresses Codex's `hooks/hooks.json` auto-discovery fallback, which would
|
||||||
|
otherwise run the Claude Code hook.
|
||||||
|
|
||||||
> **A hook *system* is not a session-start *event*.** A harness can have a
|
> **A hook *system* is not a session-start *event*.** A harness can have a
|
||||||
> `hooks.json` mechanism — and even contain the literal string `SessionStart` in
|
> `hooks.json` mechanism — and even contain the literal string `SessionStart` in
|
||||||
@@ -785,7 +787,7 @@ Use this as the live index; when in doubt, read the files, not this table.
|
|||||||
| Harness | Entry point | Bootstrap mechanism | Tool mapping | Tests | Distribution |
|
| Harness | Entry point | Bootstrap mechanism | Tool mapping | Tests | Distribution |
|
||||||
|---|---|---|---|---|---|
|
|---|---|---|---|---|---|
|
||||||
| Claude Code | `.claude-plugin/plugin.json` + `hooks/hooks.json` | shell hook → `hooks/session-start` (`hookSpecificOutput.additionalContext`) | native `Skill` tool; no adapter file needed | `tests/hooks/` | marketplace |
|
| Claude Code | `.claude-plugin/plugin.json` + `hooks/hooks.json` | shell hook → `hooks/session-start` (`hookSpecificOutput.additionalContext`) | native `Skill` tool; no adapter file needed | `tests/hooks/` | marketplace |
|
||||||
| Codex | `.codex-plugin/plugin.json` (declares empty `hooks`) | native skill discovery (no session-start hook) | `references/codex-tools.md` | `tests/codex/`, `tests/codex-plugin-sync/` | fork sync (`scripts/sync-to-codex-plugin.sh`) |
|
| Codex | `.codex-plugin/plugin.json` + `hooks/hooks-codex.json` | native skill discovery at startup; shell hook → `hooks/session-start-codex` re-injects after compaction only | `references/codex-tools.md` | `tests/codex/`, `tests/codex-plugin-sync/` | fork sync (`scripts/sync-to-codex-plugin.sh`) |
|
||||||
| Cursor | `.cursor-plugin/plugin.json` + `hooks/hooks-cursor.json` | shell hook → `hooks/session-start` (`additional_context`) | none needed (Claude Code–compatible tool surface) | `tests/hooks/` | hand-authored |
|
| Cursor | `.cursor-plugin/plugin.json` + `hooks/hooks-cursor.json` | shell hook → `hooks/session-start` (`additional_context`) | none needed (Claude Code–compatible tool surface) | `tests/hooks/` | hand-authored |
|
||||||
| Copilot CLI | (shares Claude Code hook path; `COPILOT_CLI` env) | shell hook → `hooks/session-start` (`additionalContext`) | none needed (Claude Code–compatible tool surface) | `tests/hooks/` | — |
|
| Copilot CLI | (shares Claude Code hook path; `COPILOT_CLI` env) | shell hook → `hooks/session-start` (`additionalContext`) | none needed (Claude Code–compatible tool surface) | `tests/hooks/` | — |
|
||||||
| Gemini CLI | `gemini-extension.json` + `GEMINI.md` | instructions file `@`-includes bootstrap + mapping | `references/gemini-tools.md` | — | `gemini extensions install` |
|
| Gemini CLI | `gemini-extension.json` + `GEMINI.md` | instructions file `@`-includes bootstrap + mapping | `references/gemini-tools.md` | — | `gemini extensions install` |
|
||||||
|
|||||||
@@ -0,0 +1,17 @@
|
|||||||
|
{
|
||||||
|
"hooks": {
|
||||||
|
"SessionStart": [
|
||||||
|
{
|
||||||
|
"matcher": "compact",
|
||||||
|
"hooks": [
|
||||||
|
{
|
||||||
|
"type": "command",
|
||||||
|
"command": "\"${PLUGIN_ROOT}/hooks/run-hook.cmd\" session-start-codex",
|
||||||
|
"async": false,
|
||||||
|
"timeout": 30
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
}
|
||||||
Executable
+56
@@ -0,0 +1,56 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# Codex SessionStart hook for the superpowers plugin.
|
||||||
|
#
|
||||||
|
# Codex re-fires SessionStart with source:"compact" after every context
|
||||||
|
# compaction (verified on codex-cli 0.145.0). Compaction replaces the live
|
||||||
|
# context with a summary, which sheds the using-superpowers bootstrap and any
|
||||||
|
# active skill's instructions — the measured cause of mid-session dispatch
|
||||||
|
# drift in long multi-agent runs. This hook re-injects the bootstrap at
|
||||||
|
# exactly that moment, restoring the same re-injection Claude Code performs
|
||||||
|
# via its "startup|clear|compact" SessionStart matcher.
|
||||||
|
#
|
||||||
|
# On source:"startup" it emits nothing: the native Codex plugin path owns
|
||||||
|
# session-start injection, and duplicating it here would recreate the
|
||||||
|
# redundancy that led to the original session-start-codex hook's removal.
|
||||||
|
#
|
||||||
|
# Codex injects raw hook stdout into the model's context (verified with
|
||||||
|
# sentinel probes), so output is plain text — not the JSON envelopes other
|
||||||
|
# harnesses require of hooks/session-start.
|
||||||
|
#
|
||||||
|
# A hook failure must never break a session: every path fails open to empty
|
||||||
|
# output and exit 0.
|
||||||
|
|
||||||
|
set -u
|
||||||
|
|
||||||
|
payload="$(cat 2>/dev/null || true)"
|
||||||
|
|
||||||
|
# Act only on post-compaction re-fires. Tolerate arbitrary whitespace around
|
||||||
|
# the JSON colon; anything unparseable falls through to a silent no-op.
|
||||||
|
if ! printf '%s' "$payload" | grep -qE '"source"[[:space:]]*:[[:space:]]*"compact"'; then
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
|
|
||||||
|
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||||
|
PLUGIN_ROOT="$(cd "${SCRIPT_DIR}/.." && pwd)"
|
||||||
|
|
||||||
|
using_superpowers_content="$(cat "${PLUGIN_ROOT}/skills/using-superpowers/SKILL.md" 2>/dev/null)" || using_superpowers_content=""
|
||||||
|
if [ -z "$using_superpowers_content" ]; then
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
|
|
||||||
|
# printf instead of heredocs throughout: heredocs hang on bash 5.3+.
|
||||||
|
# See: https://github.com/obra/superpowers/issues/571
|
||||||
|
printf '%s\n' "<EXTREMELY_IMPORTANT>"
|
||||||
|
printf '%s\n\n' "You have superpowers."
|
||||||
|
printf '%s\n\n' "**Below is the full content of your 'superpowers:using-superpowers' skill - your introduction to using skills. For all other skills, use the 'Skill' tool:**"
|
||||||
|
printf '%s\n' "$using_superpowers_content"
|
||||||
|
printf '%s\n\n' "</EXTREMELY_IMPORTANT>"
|
||||||
|
printf '%s\n' "<CONTEXT_RESTORED>"
|
||||||
|
printf '%s\n' "Your context was just summarized (compacted). The summary preserves your progress but not your working instructions — the files are authoritative."
|
||||||
|
printf '%s\n' ""
|
||||||
|
printf '%s\n' "Before your next tool call:"
|
||||||
|
printf '%s\n' "- Re-read the SKILL.md of any skill you are mid-way through executing. If you are executing subagent-driven-development, re-read skills/subagent-driven-development/SKILL.md."
|
||||||
|
printf '%s\n' "- On Codex, also re-read skills/using-superpowers/references/codex-tools.md and follow its dispatch rules on every spawn_agent call."
|
||||||
|
printf '%s\n' "</CONTEXT_RESTORED>"
|
||||||
|
|
||||||
|
exit 0
|
||||||
@@ -40,8 +40,9 @@ Options:
|
|||||||
-h, --help Show this help.
|
-h, --help Show this help.
|
||||||
|
|
||||||
The archive is rootless: .codex-plugin/, assets/, skills/, README.md, LICENSE,
|
The archive is rootless: .codex-plugin/, assets/, skills/, README.md, LICENSE,
|
||||||
and CODE_OF_CONDUCT.md sit at the archive root. Source-only repo files, hooks, tests,
|
CODE_OF_CONDUCT.md, and the Codex SessionStart hook (hooks/hooks-codex.json plus
|
||||||
docs, and other harness manifests are intentionally not shipped.
|
its two scripts) sit at the archive root. Source-only repo files, other-harness
|
||||||
|
hooks, tests, docs, and other harness manifests are intentionally not shipped.
|
||||||
EOF
|
EOF
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -238,6 +239,9 @@ git -C "$REPO_ROOT" -c tar.umask=0022 archive --format=tar "$REF" -- \
|
|||||||
LICENSE \
|
LICENSE \
|
||||||
README.md \
|
README.md \
|
||||||
assets \
|
assets \
|
||||||
|
hooks/hooks-codex.json \
|
||||||
|
hooks/run-hook.cmd \
|
||||||
|
hooks/session-start-codex \
|
||||||
skills \
|
skills \
|
||||||
| tar -xpf - -C "$STAGE"
|
| tar -xpf - -C "$STAGE"
|
||||||
|
|
||||||
@@ -333,7 +337,7 @@ esac
|
|||||||
|
|
||||||
unexpected_paths="$(
|
unexpected_paths="$(
|
||||||
printf '%s\n' "$archive_paths" |
|
printf '%s\n' "$archive_paths" |
|
||||||
grep -E '(^superpowers/|^\.agents/|^hooks/|package\.json$|^\.git|^\.pytest_cache|^\.ruff_cache|^scripts/|^tests/|^docs/|^evals/|^lib/|^\.claude|^\.cursor|^\.kimi|^\.opencode|^\.pi|^AGENTS\.md$|^CLAUDE\.md$|^GEMINI\.md$|^RELEASE-NOTES\.md$|^CHANGELOG\.md$)' || true
|
grep -E '(^superpowers/|^\.agents/|^hooks/hooks\.json$|^hooks/hooks-cursor\.json$|^hooks/session-start$|package\.json$|^\.git|^\.pytest_cache|^\.ruff_cache|^scripts/|^tests/|^docs/|^evals/|^lib/|^\.claude|^\.cursor|^\.kimi|^\.opencode|^\.pi|^AGENTS\.md$|^CLAUDE\.md$|^GEMINI\.md$|^RELEASE-NOTES\.md$|^CHANGELOG\.md$)' || true
|
||||||
)"
|
)"
|
||||||
if [[ -n "$unexpected_paths" ]]; then
|
if [[ -n "$unexpected_paths" ]]; then
|
||||||
printf '%s\n' "$unexpected_paths" | sed 's/^/ /' >&2
|
printf '%s\n' "$unexpected_paths" | sed 's/^/ /' >&2
|
||||||
|
|||||||
@@ -34,6 +34,15 @@ Subagent (general-purpose):
|
|||||||
|
|
||||||
Your review is read-only on this checkout. Do not mutate the working tree, the index, HEAD, or branch state in any way. Use tools like `git show`, `git diff`, and `git log` to inspect history. If you need a working copy of a different revision, check it out into a separate temporary directory (e.g. `git worktree add /tmp/review-[SHA] [SHA]`) — never move HEAD on this checkout.
|
Your review is read-only on this checkout. Do not mutate the working tree, the index, HEAD, or branch state in any way. Use tools like `git show`, `git diff`, and `git log` to inspect history. If you need a working copy of a different revision, check it out into a separate temporary directory (e.g. `git worktree add /tmp/review-[SHA] [SHA]`) — never move HEAD on this checkout.
|
||||||
|
|
||||||
|
## You Do Not Dispatch Subagents
|
||||||
|
|
||||||
|
Do all of this review yourself. Never spawn a subagent to review part
|
||||||
|
of the diff, and never spawn another reviewer for a second opinion.
|
||||||
|
This process already provides every review seat the work gets; a
|
||||||
|
reviewer you spawn duplicates one of them at full cost, and its
|
||||||
|
verdict counts for nothing. If the diff feels too large for one
|
||||||
|
pass, review it in passes yourself and say so in your report.
|
||||||
|
|
||||||
## What to Check
|
## What to Check
|
||||||
|
|
||||||
**Plan alignment:**
|
**Plan alignment:**
|
||||||
|
|||||||
@@ -156,16 +156,27 @@ a ledger file, not only in todos.
|
|||||||
Read the plan once, note its context and Global Constraints, and create a
|
Read the plan once, note its context and Global Constraints, and create a
|
||||||
todo per task.
|
todo per task.
|
||||||
|
|
||||||
Before dispatching Task 1, scan the plan once for conflicts:
|
Before dispatching Task 1, scan the plan once for conflicts, writing down
|
||||||
|
what you checked as you check it:
|
||||||
|
|
||||||
- tasks that contradict each other or the plan's Global Constraints
|
- tasks that contradict each other or the plan's Global Constraints
|
||||||
- anything the plan explicitly mandates that the review rubric treats as a
|
- anything the plan explicitly mandates that the review rubric treats as a
|
||||||
defect (a test that asserts nothing, verbatim duplication of a logic block)
|
defect (a test that asserts nothing, verbatim duplication of a logic block)
|
||||||
|
|
||||||
Rule on everything you find before execution begins — each finding against
|
The scan's output is a table, not a verdict. One row for every pair of tasks
|
||||||
the plan text that mandates it — and record each ruling in the ledger. If the
|
that share a file or an interface: the two tasks, what one produces against
|
||||||
scan is clean, proceed without comment. The review loop remains the net for
|
what the other consumes, and what you found. One row for every task: whether
|
||||||
conflicts that only emerge from implementation.
|
its own text agrees with itself — the tests it specifies against the code it
|
||||||
|
specifies, the files it creates against the files it later touches. "The scan
|
||||||
|
is clean" without those rows is not a scan you ran.
|
||||||
|
|
||||||
|
Write the table to the ledger. Rule on everything you find before execution
|
||||||
|
begins — each finding against the plan text that mandates it — and record
|
||||||
|
each ruling in the ledger. If the scan is clean, proceed without comment.
|
||||||
|
Rule on each conflict it surfaces — the spec is the binding authority, the
|
||||||
|
plan is its argument — record the ruling beside its row, and dispatch
|
||||||
|
Task 1. The review loop remains the net for conflicts that only emerge from
|
||||||
|
implementation.
|
||||||
|
|
||||||
## Model Selection
|
## Model Selection
|
||||||
|
|
||||||
@@ -206,10 +217,29 @@ that implementer. Single-file mechanical fixes also take the cheapest tier.
|
|||||||
|
|
||||||
## The Task Loop
|
## The Task Loop
|
||||||
|
|
||||||
|
**Batch small same-shape work.** When the plan lists several tasks that are
|
||||||
|
each a small, independent edit of the same kind — the same one-line fix,
|
||||||
|
constant change, or field addition repeated across files — do not dispatch
|
||||||
|
one subagent per task. Compose ONE dispatch brief listing every file and
|
||||||
|
its change, send the whole batch to a single subagent, and review its diff
|
||||||
|
as one unit. Reserve one-dispatch-per-task for work that needs its own
|
||||||
|
judgment, its own tests, or its own review surface.
|
||||||
|
|
||||||
Everything you paste into a dispatch prompt — and everything a subagent
|
Everything you paste into a dispatch prompt — and everything a subagent
|
||||||
prints back — stays resident in your context for the rest of the session
|
prints back — stays resident in your context for the rest of the session
|
||||||
and is re-read on every later turn. Hand artifacts over as files.
|
and is re-read on every later turn. Hand artifacts over as files.
|
||||||
|
|
||||||
|
**Waiting on dispatched subagents:** never poll a wait interface with
|
||||||
|
short timeouts, and never sit in one silent, open-ended wait either.
|
||||||
|
While you have local work — ledger updates, packaging the next review,
|
||||||
|
reading reports — keep working; child results arrive on their own.
|
||||||
|
When you are genuinely idle, wait in bounded stretches (five to ten
|
||||||
|
minutes, where your platform allows), and between stretches post one
|
||||||
|
line of status and reconcile your live children: list them, and chase
|
||||||
|
any that finished without reporting. A bounded stretch keeps nearly
|
||||||
|
all of a long wait's efficiency while guaranteeing a stuck or lost
|
||||||
|
child is noticed within minutes, not at the end of the session.
|
||||||
|
|
||||||
### 1. Dispatch the implementer
|
### 1. Dispatch the implementer
|
||||||
|
|
||||||
Record BASE (`git rev-parse HEAD`) before dispatching — the review package
|
Record BASE (`git rev-parse HEAD`) before dispatching — the review package
|
||||||
@@ -236,6 +266,12 @@ and fix-round diffs need it.
|
|||||||
later dispatches — a real session's dispatch hit 42k chars of which 99%
|
later dispatches — a real session's dispatch hit 42k chars of which 99%
|
||||||
was pasted history. A fresh subagent needs its task, the interfaces it
|
was pasted history. A fresh subagent needs its task, the interfaces it
|
||||||
touches, and the global constraints. Nothing else.
|
touches, and the global constraints. Nothing else.
|
||||||
|
- The dispatch carries the no-subagents contract (it is in the
|
||||||
|
implementer template): the implementer never dispatches subagents —
|
||||||
|
not helpers, and never a reviewer. Review arrives from you, after the
|
||||||
|
report. In real sessions, every reviewer a worker spawned duplicated
|
||||||
|
the task review the controller dispatched anyway — a full extra
|
||||||
|
review seat per task.
|
||||||
- If an earlier task parked a finding in the area this task touches, carry
|
- If an earlier task parked a finding in the area this task touches, carry
|
||||||
a pointer to that ledger entry in the dispatch.
|
a pointer to that ledger entry in the dispatch.
|
||||||
- Record the implementer's agent identity from the dispatch result —
|
- Record the implementer's agent identity from the dispatch result —
|
||||||
@@ -459,6 +495,7 @@ Use superpowers:finishing-a-development-branch.
|
|||||||
| "The fix was small, skip the re-review" | Unreviewed fixes are how regressions land. Every round ends with a scoped re-review. |
|
| "The fix was small, skip the re-review" | Unreviewed fixes are how regressions land. Every round ends with a scoped re-review. |
|
||||||
| "Reviews slow the loop down" | The loop without reviews is just unverified churn. Reviews are the loop's brakes and steering. |
|
| "Reviews slow the loop down" | The loop without reviews is just unverified churn. Reviews are the loop's brakes and steering. |
|
||||||
| "Ledger bookkeeping is overhead" | The ledger is what survives compaction. Controllers without one have re-dispatched entire completed task sequences. |
|
| "Ledger bookkeeping is overhead" | The ledger is what survives compaction. Controllers without one have re-dispatched entire completed task sequences. |
|
||||||
|
| "The implementer spawned its own reviewer — free extra assurance" | It's a duplicate seat reviewing the same diff; the task review is the gate. A worker-spawned reviewer is a defect to flag, not rigor. |
|
||||||
|
|
||||||
## Example Workflow
|
## Example Workflow
|
||||||
|
|
||||||
|
|||||||
@@ -47,6 +47,18 @@ Subagent (general-purpose):
|
|||||||
While iterating, run the focused test for what you're changing; run the
|
While iterating, run the focused test for what you're changing; run the
|
||||||
full suite once before committing, not after every edit.
|
full suite once before committing, not after every edit.
|
||||||
|
|
||||||
|
## You Do Not Dispatch Subagents
|
||||||
|
|
||||||
|
Do all of this task's work yourself. Never spawn a subagent to
|
||||||
|
implement part of the task, and above all never spawn a reviewer to
|
||||||
|
check your work. Self-review (below) means reading your own diff.
|
||||||
|
Review is the controller's job: after you report, it dispatches a
|
||||||
|
fresh reviewer against your diff. A reviewer you spawn duplicates
|
||||||
|
that review at full cost, and its approval counts for nothing in
|
||||||
|
the process. If you catch yourself thinking "an independent review
|
||||||
|
would strengthen my report" — that review is already scheduled.
|
||||||
|
Report instead.
|
||||||
|
|
||||||
## Code Organization
|
## Code Organization
|
||||||
|
|
||||||
You reason best about code you can hold in context at once, and your edits are more
|
You reason best about code you can hold in context at once, and your edits are more
|
||||||
|
|||||||
@@ -43,6 +43,15 @@ Subagent (general-purpose):
|
|||||||
Your review is read-only on this checkout. Do not mutate the working
|
Your review is read-only on this checkout. Do not mutate the working
|
||||||
tree, the index, HEAD, or branch state in any way.
|
tree, the index, HEAD, or branch state in any way.
|
||||||
|
|
||||||
|
## You Do Not Dispatch Subagents
|
||||||
|
|
||||||
|
Do all of this review yourself. Never spawn a subagent to review part
|
||||||
|
of the diff, and never spawn another reviewer for a second opinion.
|
||||||
|
This process already provides every review seat the work gets; a
|
||||||
|
reviewer you spawn duplicates one of them at full cost, and its
|
||||||
|
verdict counts for nothing. If the diff feels too large for one
|
||||||
|
pass, review it in passes yourself and say so in your report.
|
||||||
|
|
||||||
## Scope
|
## Scope
|
||||||
|
|
||||||
Your scope is the findings list and the fix diff. Verdict every finding.
|
Your scope is the findings list and the fix diff. Verdict every finding.
|
||||||
|
|||||||
@@ -52,6 +52,15 @@ Subagent (general-purpose):
|
|||||||
Your review is read-only on this checkout. Do not mutate the working
|
Your review is read-only on this checkout. Do not mutate the working
|
||||||
tree, the index, HEAD, or branch state in any way.
|
tree, the index, HEAD, or branch state in any way.
|
||||||
|
|
||||||
|
## You Do Not Dispatch Subagents
|
||||||
|
|
||||||
|
Do all of this review yourself. Never spawn a subagent to review part
|
||||||
|
of the diff, and never spawn another reviewer for a second opinion.
|
||||||
|
This process already provides every review seat the work gets; a
|
||||||
|
reviewer you spawn duplicates one of them at full cost, and its
|
||||||
|
verdict counts for nothing. If the diff feels too large for one
|
||||||
|
pass, review it in passes yourself and say so in your report.
|
||||||
|
|
||||||
## Do Not Trust the Report
|
## Do Not Trust the Report
|
||||||
|
|
||||||
Treat the implementer's report as unverified claims about the code. It
|
Treat the implementer's report as unverified claims about the code. It
|
||||||
@@ -86,6 +95,12 @@ Subagent (general-purpose):
|
|||||||
- **Misunderstood:** right feature built the wrong way, wrong problem
|
- **Misunderstood:** right feature built the wrong way, wrong problem
|
||||||
solved
|
solved
|
||||||
|
|
||||||
|
If the brief lists several files each with its own change (a batched
|
||||||
|
dispatch), check the diff against that list file by file: every listed
|
||||||
|
file must have its corresponding hunk. A listed file the diff never
|
||||||
|
touches is a Missing finding, no matter how clean the rest of the
|
||||||
|
batch looks.
|
||||||
|
|
||||||
If a requirement cannot be verified from this diff alone (it lives in
|
If a requirement cannot be verified from this diff alone (it lives in
|
||||||
unchanged code or spans tasks), report it as a ⚠️ item instead of
|
unchanged code or spans tasks), report it as a ⚠️ item instead of
|
||||||
broadening your search.
|
broadening your search.
|
||||||
|
|||||||
@@ -7,7 +7,92 @@ Add to your Codex config (`~/.codex/config.toml`):
|
|||||||
multi_agent = true
|
multi_agent = true
|
||||||
```
|
```
|
||||||
|
|
||||||
This enables `spawn_agent`, `wait_agent`, and `close_agent` for skills like `dispatching-parallel-agents` and `subagent-driven-development`. When using subagent-driven-development, close reviewer subagents when their review returns. Keep each implementer subagent open until its task's review passes — the fix loop resumes the implementer — then close it. If your harness cannot send another message to a spawned agent, dispatch each fix round as a fresh implementer carrying the brief, the report file, and the findings.
|
This enables the multi-agent tools that skills like
|
||||||
|
`dispatching-parallel-agents` and `subagent-driven-development` use.
|
||||||
|
Which tools you get depends on the multi-agent version your model
|
||||||
|
preset selects (current presets run V2; older ones run V1). Trust your
|
||||||
|
actual tool list over any table — including this one — when they
|
||||||
|
disagree.
|
||||||
|
|
||||||
|
- **Spawning:** give children a clean context with
|
||||||
|
`spawn_agent {fork_turns: "none"}`; the default `"all"` copies your
|
||||||
|
entire transcript into the child. On Codex 0.145+, role files under
|
||||||
|
`~/.codex/agents/` attach to isolated forks via `agent_type`.
|
||||||
|
Full-history forks accept `model` and `reasoning_effort` overrides
|
||||||
|
(only `agent_type` is refused there) — isolated forks are the SDD
|
||||||
|
default for context hygiene, not because overrides require them.
|
||||||
|
- **Fix rounds:** resume the implementer with `followup_task` — it
|
||||||
|
delivers your message, triggers a turn, and transparently reloads a
|
||||||
|
child the harness evicted. Never dispatch a fresh implementer on the
|
||||||
|
theory that a spawned agent cannot be messaged again; on V2 it
|
||||||
|
always can.
|
||||||
|
- **Lifecycle:** V2 has no `close_agent`. Finished children are
|
||||||
|
evicted automatically when slots are needed; leaving them unclosed
|
||||||
|
costs nothing. Only V1 sessions have `close_agent` — there, close
|
||||||
|
reviewers when their review returns, and close each implementer
|
||||||
|
after its task's review passes.
|
||||||
|
- **Model names:** never copy a model name from a skill, table, or old
|
||||||
|
session into `spawn_agent` without checking it against your current
|
||||||
|
spawn allowlist — V2 accepts only V2-capable presets and hard-errors
|
||||||
|
on the rest.
|
||||||
|
|
||||||
|
## Waiting on children
|
||||||
|
|
||||||
|
`wait_agent` is an event subscription, not a poll: a long wait wakes
|
||||||
|
the moment a child produces mailbox activity, with the same latency as
|
||||||
|
a short one. Short-timeout polling buys nothing and costs a tool call —
|
||||||
|
and a context rebill — per poll. In measured sessions, roughly
|
||||||
|
two-thirds of all wait calls were short polls that timed out.
|
||||||
|
|
||||||
|
- While you still have local work, do not wait at all. A completed
|
||||||
|
child's final answer is pushed into your mailbox and arrives with
|
||||||
|
your next turn.
|
||||||
|
- When you are genuinely idle with children outstanding, wait in
|
||||||
|
bounded stretches: `wait_agent` with `timeout_ms` 300000-600000
|
||||||
|
(5-10 minutes). After each stretch — wake or timeout — post one
|
||||||
|
status line, run `list_agents`, and chase any child that finished
|
||||||
|
without reporting. Never stack polls shorter than five minutes; the
|
||||||
|
event subscription wakes a bounded stretch just as fast as a short
|
||||||
|
one.
|
||||||
|
- Completion mail cannot wake an idle controller (it is delivered
|
||||||
|
without triggering a turn); covering that idle window is
|
||||||
|
`wait_agent`'s only job. A stretch that times out with no activity
|
||||||
|
is your cue to reconcile, not to shorten the next stretch.
|
||||||
|
|
||||||
|
## Model routing on spawns
|
||||||
|
|
||||||
|
Every `spawn_agent` you issue — including when you are yourself a
|
||||||
|
spawned child running a fan-out — sets `model` AND `reasoning_effort`
|
||||||
|
explicitly, per the Model Selection rules of the skill you are
|
||||||
|
executing. Setting `model` alone is a trap: the child's effort
|
||||||
|
silently resets to that model's default, not to yours.
|
||||||
|
|
||||||
|
Ask your human partner to add a machine-level backstop to
|
||||||
|
`~/.codex/config.toml` so any spawn that slips through still routes to
|
||||||
|
a deliberate tier instead of silently inheriting the session's most
|
||||||
|
expensive model:
|
||||||
|
|
||||||
|
```toml
|
||||||
|
[agents]
|
||||||
|
default_subagent_model = "<a mid-tier model from your spawn allowlist>"
|
||||||
|
default_subagent_reasoning_effort = "medium"
|
||||||
|
```
|
||||||
|
|
||||||
|
## Compaction sheds these instructions
|
||||||
|
|
||||||
|
Context compaction replaces your transcript with a summary that keeps
|
||||||
|
your progress but not your working instructions — the first
|
||||||
|
post-compaction dispatch is where routing drift starts, and once one
|
||||||
|
bare spawn lands, the broken pattern becomes its own precedent. The
|
||||||
|
plugin ships a compaction re-injection hook (`hooks/hooks-codex.json`,
|
||||||
|
Codex 0.145+) that restores the bootstrap after every compaction; it
|
||||||
|
needs one-time trust approval, so if you never see a
|
||||||
|
`<CONTEXT_RESTORED>` block after a compaction, tell your human partner
|
||||||
|
the hook may be untrusted or unsupported on this version. Without it,
|
||||||
|
re-ground yourself: when a summary appears in your context, re-read
|
||||||
|
this file and the SKILL.md of the skill you are mid-way through
|
||||||
|
executing before your next dispatch, and trust the ledger over your
|
||||||
|
summarized memory of what happened.
|
||||||
|
|
||||||
## Environment Detection
|
## Environment Detection
|
||||||
|
|
||||||
|
|||||||
@@ -52,25 +52,37 @@ if not plugin_manifest.exists():
|
|||||||
manifest = json.loads(plugin_manifest.read_text(encoding="utf-8"))
|
manifest = json.loads(plugin_manifest.read_text(encoding="utf-8"))
|
||||||
assert_equal(manifest.get("name"), plugin.get("name"), "plugin manifest name")
|
assert_equal(manifest.get("name"), plugin.get("name"), "plugin manifest name")
|
||||||
|
|
||||||
# Codex auto-discovers a plugin's hooks/hooks.json whenever the Codex manifest
|
# The Codex manifest must declare its hooks explicitly. An absent field makes
|
||||||
# has no `hooks` field: load_plugin_hooks falls back to a hardcoded
|
# load_plugin_hooks fall back to a hardcoded DEFAULT_HOOKS_CONFIG_FILE =
|
||||||
# DEFAULT_HOOKS_CONFIG_FILE = "hooks/hooks.json" and registers it. That file is
|
# "hooks/hooks.json" — the Claude Code SessionStart hook, which injects the
|
||||||
# the Claude Code SessionStart hook, it is tracked in this repo, and this
|
# bootstrap at startup and must not run on Codex. The explicit pointer both
|
||||||
# marketplace installs the whole repo root (source url "./"), so on Codex the
|
# registers the Codex compaction re-injection hook and overrides that fallback.
|
||||||
# fallback re-registers the SessionStart hook and its install-time trust prompt.
|
|
||||||
# Declaring an empty inline hooks object ({}) parses as an empty inline hook set
|
|
||||||
# and suppresses the auto-discovery. An absent field, an empty array ([]), and
|
|
||||||
# an empty inline list all collapse back to the fallback, so the value must be
|
|
||||||
# exactly an empty object.
|
|
||||||
hooks_config = repo_root / "hooks" / "hooks.json"
|
hooks_config = repo_root / "hooks" / "hooks.json"
|
||||||
if not hooks_config.exists():
|
if not hooks_config.exists():
|
||||||
raise AssertionError("hooks/hooks.json must exist (Claude Code SessionStart hook)")
|
raise AssertionError("hooks/hooks.json must exist (Claude Code SessionStart hook)")
|
||||||
|
|
||||||
assert_equal(
|
assert_equal(
|
||||||
manifest.get("hooks"),
|
manifest.get("hooks"),
|
||||||
{},
|
"./hooks/hooks-codex.json",
|
||||||
"Codex manifest must declare empty hooks {} to suppress hooks/hooks.json auto-discovery",
|
"Codex manifest must point hooks at the Codex hook config (an absent field "
|
||||||
|
"falls back to auto-discovering the Claude Code hooks/hooks.json)",
|
||||||
)
|
)
|
||||||
|
|
||||||
|
codex_hooks_path = repo_root / "hooks" / "hooks-codex.json"
|
||||||
|
if not codex_hooks_path.exists():
|
||||||
|
raise AssertionError("hooks/hooks-codex.json must exist (Codex manifest points at it)")
|
||||||
|
|
||||||
|
codex_hooks = json.loads(codex_hooks_path.read_text(encoding="utf-8"))
|
||||||
|
session_start = codex_hooks["hooks"]["SessionStart"]
|
||||||
|
assert_equal(len(session_start), 1, "Codex SessionStart hook group count")
|
||||||
|
assert_equal(session_start[0].get("matcher"), "compact", "Codex hook matcher")
|
||||||
|
entry = session_start[0]["hooks"][0]
|
||||||
|
assert_equal(entry.get("type"), "command", "Codex hook type")
|
||||||
|
command = entry.get("command", "")
|
||||||
|
if "${PLUGIN_ROOT}" not in command or not command.endswith("session-start-codex"):
|
||||||
|
raise AssertionError(
|
||||||
|
f"Codex hook command must run session-start-codex via ${{PLUGIN_ROOT}}: {command!r}"
|
||||||
|
)
|
||||||
|
|
||||||
print("Codex marketplace manifest looks good")
|
print("Codex marketplace manifest looks good")
|
||||||
PY
|
PY
|
||||||
|
|||||||
@@ -141,7 +141,7 @@ tar_extracted="$TEST_ROOT/tar-extracted"
|
|||||||
write_metadata_fixture "$metadata_source"
|
write_metadata_fixture "$metadata_source"
|
||||||
|
|
||||||
source_hooks="$(python3 -c 'import json; print(json.load(open("'"$REPO_ROOT"'/.codex-plugin/plugin.json")).get("hooks"))')"
|
source_hooks="$(python3 -c 'import json; print(json.load(open("'"$REPO_ROOT"'/.codex-plugin/plugin.json")).get("hooks"))')"
|
||||||
assert_equals "$source_hooks" "{}" "source Codex manifest suppresses local hook auto-discovery"
|
assert_equals "$source_hooks" "./hooks/hooks-codex.json" "source Codex manifest declares the Codex hook config"
|
||||||
|
|
||||||
if output="$("$SCRIPT_UNDER_TEST" --allow-dirty --metadata-source "$metadata_source" --output "$archive" 2>&1)"; then
|
if output="$("$SCRIPT_UNDER_TEST" --allow-dirty --metadata-source "$metadata_source" --output "$archive" 2>&1)"; then
|
||||||
pass "package script exits successfully"
|
pass "package script exits successfully"
|
||||||
@@ -163,10 +163,13 @@ assert_contains "$output" "SHA-256:" "reports archive checksum"
|
|||||||
extract_archive "$archive" "$extracted"
|
extract_archive "$archive" "$extracted"
|
||||||
|
|
||||||
archive_paths="$(list_archive "$archive" | normalize_archive_paths)"
|
archive_paths="$(list_archive "$archive" | normalize_archive_paths)"
|
||||||
unexpected_pattern='(^superpowers/|^\.agents/|^hooks/|package\.json$|^\.git|^\.pytest_cache|^\.ruff_cache|^scripts/|^tests/|^docs/|^evals/|^lib/|^\.claude|^\.cursor|^\.kimi|^\.opencode|^\.pi|^AGENTS\.md$|^CLAUDE\.md$|^GEMINI\.md$|^RELEASE-NOTES\.md$|^CHANGELOG\.md$)'
|
unexpected_pattern='(^superpowers/|^\.agents/|^hooks/hooks\.json$|^hooks/hooks-cursor\.json$|^hooks/session-start$|package\.json$|^\.git|^\.pytest_cache|^\.ruff_cache|^scripts/|^tests/|^docs/|^evals/|^lib/|^\.claude|^\.cursor|^\.kimi|^\.opencode|^\.pi|^AGENTS\.md$|^CLAUDE\.md$|^GEMINI\.md$|^RELEASE-NOTES\.md$|^CHANGELOG\.md$)'
|
||||||
assert_not_matches "$archive_paths" "$unexpected_pattern" "archive excludes source-only paths"
|
assert_not_matches "$archive_paths" "$unexpected_pattern" "archive excludes source-only paths"
|
||||||
assert_contains "$archive_paths" ".codex-plugin/plugin.json" "archive includes Codex manifest"
|
assert_contains "$archive_paths" ".codex-plugin/plugin.json" "archive includes Codex manifest"
|
||||||
assert_contains "$archive_paths" "skills/brainstorming/SKILL.md" "archive includes skills"
|
assert_contains "$archive_paths" "skills/brainstorming/SKILL.md" "archive includes skills"
|
||||||
|
assert_contains "$archive_paths" "hooks/hooks-codex.json" "archive includes Codex hook config"
|
||||||
|
assert_contains "$archive_paths" "hooks/session-start-codex" "archive includes Codex hook script"
|
||||||
|
assert_contains "$archive_paths" "hooks/run-hook.cmd" "archive includes hook runner"
|
||||||
assert_contains "$archive_paths" "skills/brainstorming/agents/openai.yaml" "archive includes OpenAI skill metadata"
|
assert_contains "$archive_paths" "skills/brainstorming/agents/openai.yaml" "archive includes OpenAI skill metadata"
|
||||||
assert_contains "$archive_paths" "assets/app-icon.png" "archive includes app icon"
|
assert_contains "$archive_paths" "assets/app-icon.png" "archive includes app icon"
|
||||||
assert_contains "$archive_paths" "assets/superpowers-small.svg" "archive includes composer icon"
|
assert_contains "$archive_paths" "assets/superpowers-small.svg" "archive includes composer icon"
|
||||||
|
|||||||
Executable
+115
@@ -0,0 +1,115 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||||
|
REPO_ROOT="$(cd "$SCRIPT_DIR/../.." && pwd)"
|
||||||
|
HOOK_UNDER_TEST="$REPO_ROOT/hooks/session-start-codex"
|
||||||
|
CONFIG_UNDER_TEST="$REPO_ROOT/hooks/hooks-codex.json"
|
||||||
|
|
||||||
|
FAILURES=0
|
||||||
|
|
||||||
|
pass() {
|
||||||
|
echo " [PASS] $1"
|
||||||
|
}
|
||||||
|
|
||||||
|
fail() {
|
||||||
|
echo " [FAIL] $1"
|
||||||
|
FAILURES=$((FAILURES + 1))
|
||||||
|
}
|
||||||
|
|
||||||
|
# run_hook <stdin-payload> — echoes hook stdout; fails the calling test on
|
||||||
|
# non-zero exit. env -i mirrors the codex hook executor's clean environment.
|
||||||
|
run_hook() {
|
||||||
|
printf '%s' "$1" | env -i PATH="${PATH:-}" bash "$HOOK_UNDER_TEST"
|
||||||
|
}
|
||||||
|
|
||||||
|
echo "Codex SessionStart hook tests"
|
||||||
|
|
||||||
|
startup_payload='{"session_id":"s","hook_event_name":"SessionStart","model":"gpt-5.6-terra","source":"startup"}'
|
||||||
|
if output="$(run_hook "$startup_payload")" && [ -z "$output" ]; then
|
||||||
|
pass "source=startup emits nothing and exits 0"
|
||||||
|
else
|
||||||
|
fail "source=startup emits nothing and exits 0"
|
||||||
|
printf '%s\n' "$output" | head -3 | sed 's/^/ /'
|
||||||
|
fi
|
||||||
|
|
||||||
|
compact_payload='{"session_id":"s","hook_event_name":"SessionStart","model":"gpt-5.6-terra","source":"compact"}'
|
||||||
|
if output="$(run_hook "$compact_payload")"; then
|
||||||
|
ok=1
|
||||||
|
for needle in \
|
||||||
|
"<EXTREMELY_IMPORTANT>" \
|
||||||
|
"You have superpowers." \
|
||||||
|
"name: using-superpowers" \
|
||||||
|
"<CONTEXT_RESTORED>" \
|
||||||
|
"subagent-driven-development/SKILL.md" \
|
||||||
|
"references/codex-tools.md"; do
|
||||||
|
if [[ "$output" != *"$needle"* ]]; then
|
||||||
|
ok=0
|
||||||
|
echo " missing: $needle"
|
||||||
|
fi
|
||||||
|
done
|
||||||
|
if [ "$ok" -eq 1 ]; then
|
||||||
|
pass "source=compact emits bootstrap plus re-read addendum"
|
||||||
|
else
|
||||||
|
fail "source=compact emits bootstrap plus re-read addendum"
|
||||||
|
fi
|
||||||
|
else
|
||||||
|
fail "source=compact emits bootstrap plus re-read addendum (hook exited non-zero)"
|
||||||
|
fi
|
||||||
|
|
||||||
|
# Whitespace-tolerant source matching (serializers vary).
|
||||||
|
spaced_payload='{"hook_event_name":"SessionStart", "source" : "compact"}'
|
||||||
|
if output="$(run_hook "$spaced_payload")" && [[ "$output" == *"<CONTEXT_RESTORED>"* ]]; then
|
||||||
|
pass "whitespace around the source key still triggers injection"
|
||||||
|
else
|
||||||
|
fail "whitespace around the source key still triggers injection"
|
||||||
|
fi
|
||||||
|
|
||||||
|
if output="$(printf '' | env -i PATH="${PATH:-}" bash "$HOOK_UNDER_TEST")" && [ -z "$output" ]; then
|
||||||
|
pass "empty stdin fails open to no output, exit 0"
|
||||||
|
else
|
||||||
|
fail "empty stdin fails open to no output, exit 0"
|
||||||
|
fi
|
||||||
|
|
||||||
|
if output="$(run_hook 'not json at all {{{')" && [ -z "$output" ]; then
|
||||||
|
pass "garbage stdin fails open to no output, exit 0"
|
||||||
|
else
|
||||||
|
fail "garbage stdin fails open to no output, exit 0"
|
||||||
|
fi
|
||||||
|
|
||||||
|
# A compact mention inside some other field must not trigger injection.
|
||||||
|
decoy_payload='{"hook_event_name":"SessionStart","source":"startup","cwd":"/tmp/compact"}'
|
||||||
|
if output="$(run_hook "$decoy_payload")" && [ -z "$output" ]; then
|
||||||
|
pass "compact appearing outside the source field does not trigger"
|
||||||
|
else
|
||||||
|
fail "compact appearing outside the source field does not trigger"
|
||||||
|
fi
|
||||||
|
|
||||||
|
if node -e '
|
||||||
|
const config = JSON.parse(require("fs").readFileSync(process.argv[1], "utf8"));
|
||||||
|
const group = config.hooks.SessionStart[0];
|
||||||
|
if (group.matcher !== "compact") {
|
||||||
|
console.error(`hook matcher is ${JSON.stringify(group.matcher)}, expected "compact"`);
|
||||||
|
process.exit(1);
|
||||||
|
}
|
||||||
|
const entry = group.hooks[0];
|
||||||
|
if (entry.type !== "command") {
|
||||||
|
console.error(`hook type is ${JSON.stringify(entry.type)}, expected "command"`);
|
||||||
|
process.exit(1);
|
||||||
|
}
|
||||||
|
if (!entry.command.includes("${PLUGIN_ROOT}") || !/run-hook\.cmd" session-start-codex$/.test(entry.command)) {
|
||||||
|
console.error(`unexpected command shape: ${entry.command}`);
|
||||||
|
process.exit(1);
|
||||||
|
}
|
||||||
|
' "$CONFIG_UNDER_TEST"; then
|
||||||
|
pass "hooks-codex.json runs session-start-codex via \${PLUGIN_ROOT} on compact"
|
||||||
|
else
|
||||||
|
fail "hooks-codex.json runs session-start-codex via \${PLUGIN_ROOT} on compact"
|
||||||
|
fi
|
||||||
|
|
||||||
|
if [[ "$FAILURES" -gt 0 ]]; then
|
||||||
|
echo "STATUS: FAILED ($FAILURES failure(s))"
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
echo "STATUS: PASSED"
|
||||||
Reference in New Issue
Block a user