* fix(creator): raise max_history_messages to 80 for polling workflows
Creator Hand's async video_generate path polls video_status every 15-20s
until completion (1-3 min typical), consuming ~5-15 turns per video
request. Combined workflows (video + TTS + music) plus normal back-and-
forth cross the kernel default of 40 messages quickly, which surfaced
in user logs as:
WARN run_agent_loop: Trimming old messages at safe turn boundary
agent=creator:creator-hand total_messages=41 trimming=2
INFO run_agent_loop: prompt cache metrics for turn
hit_ratio=0.0 creation=0 read=0
Every turn was hitting the trim cap and invalidating the prompt-cache
prefix. 80 covers ~30 polling iterations plus a comfortable pre-context
window without runaway memory growth. Other hands keep the default 40.
* ci(refresh-cache): open PR instead of pushing directly to main
Branch protection on `main` started rejecting the workflow's auto-commit
with GH006 "Changes must be made through a pull request" — see run
25632824585 on 2026-05-10 against commit 6785807 (the first push that
hit the tightened protection). Direct push is precisely what the file's
own security comment (#1) warns against ("Compromised maintainer pushes
a malicious plugins-index.json directly to main. Mitigation: GitHub
branch protection on main requires PR review"), so the fix preserves
that gate rather than working around it.
The workflow now creates a short-lived `automation/refresh-indexes-<sha>`
branch, commits the regen there, pushes, and opens a PR back to main
via `gh pr create`. Maintainers see a one-click squash-merge.
Permissions: add `pull-requests: write` to the existing `contents: write`
so `gh pr create` can be authorised through the default GITHUB_TOKEN.
The post-merge run on the index PR is a no-op (no diff under
`hands/**`, `plugins/**`, etc. between consecutive states), so no
`[skip ci]` marker is needed and no loop is possible.
Without this fix, every content PR landing on main leaves
plugins-index.json + registry-index.json stale, blocking new agents and
hands from reaching daemons until a maintainer manually regenerates.
* fix(hands): raise max_history_messages on long-workflow coordinators
Three hand coordinators have workflows that routinely exceed the kernel
default history cap on a single user turn:
- researcher (max_iterations=80) — deep web_search → web_fetch →
summarize loops with multi-source synthesis. 80 iterations × ~4
messages each → 200+ messages per user turn. Set to 120.
- devops (max_iterations=60) — incident response and CI/CD fan out
into long shell_exec chains (logs, retries, post-mortems). Set to 80.
- predictor (max_iterations=60) — long reasoning chains accumulating
signals across many web/knowledge queries, with scheduled re-checks
referring back. Set to 80.
Creator's existing override is rephrased "raise above the kernel
default" so the comment stays correct regardless of the order this PR
and the upstream kernel-default bump (librefang side) land in.
Other hands (lead/linkedin/reddit/clip/analytics/apitester/browser/
collector/strategist) stay on the kernel default; the upstream bump
covers them.
Round-3 PR re-review follow-ups:
LOW — CODEOWNERS missed /scripts/build-plugins-index.mjs and
/wrangler.toml. Both can change the bytes that get signed without
touching the sign step. Replaced the per-file enumeration with /scripts/
catch-all and added wrangler.toml.
HIGH — workflow comment said the post-sign verify step "catches a buggy
or tampered sign-script run". The "tampered" claim was wrong: any
attacker who can edit sign-plugins-index.mjs in a PR can edit the
verify step (and the embedded pubkey) in the same diff. Re-stated as
"catches accidental regressions only — CODEOWNERS is what stops
adversarial edits".
Branch protection on main has been enabled separately via gh API
(force-push + deletion blocked, PR review required, CODEOWNERS
enforced; bypass for repo owner and github-actions[bot] so the
auto-publish workflow keeps working).
The first iteration moved signing from the Cloudflare worker into this
repo's CI to fix the worker-as-sign-oracle defect. The re-review pointed
out that this just relocated the trust problem: anyone who lands a
commit on main gets the resulting bytes signed automatically. Mitigations:
* CODEOWNERS — sign-script, build scripts, the workflow itself, and
the auto-generated index/sig files are owned by the registry owner.
Combined with branch protection requiring CODEOWNERS review, a PR
touching the signing infrastructure or the artefacts it produces
cannot land without explicit owner sign-off. Plugin contributions
under plugins/<name>/ are covered by the standard PR-review rules
but don't trip CODEOWNERS unless they touch the signing path.
* SHA-pinned actions — actions/checkout@v4 and actions/setup-node@v4
replaced with full-SHA refs (v4.2.2 / v4.1.0). Blocks
action-supply-chain swaps (a malicious mutable tag move on a
popular action would otherwise execute in the same job that has
REGISTRY_PRIVATE_KEY in scope).
* Post-sign self-verify — a new step verifies plugins-index.json.sig
against the committed pubkey (not a secret) BEFORE the .sig hits
main. Catches a buggy sign-script run, an env-leak that produces
zero-bytes output, or a half-applied edit.
* REGISTRY_PRIVATE_KEY env scope — explicitly noted in workflow
comment that the secret is on the sign step ONLY (where it was
already), not the job. Prevents future contributors lifting it
to job-level out of convenience.
* In-workflow security model docstring — enumerates the residual
threat model (compromised maintainer, malicious PR, sign-step
bug, leaked refresh token) and what each defense addresses.
Trust root is now: GitHub branch protection on main + CODEOWNERS on
sign infrastructure + SHA-pinned actions + post-sign verification.
A maintainer with push rights can still ship malicious bytes through
a code-review bypass — that residual is the same as for any signed
package registry and falls outside what CI alone can mitigate.
Branch protection on `main` MUST be configured by an org admin to
match the assumptions in CODEOWNERS:
- require pull request reviews (at least 1)
- require review from CODEOWNERS
- dismiss stale approvals on new commits
- restrict who can push directly to main (org admins only)
Pair with the worker-side simplification on the librefang PR — the
worker is now a pure transport (no key material, no signing) and this
repo's CI takes over signature production.
scripts/sign-plugins-index.mjs reads REGISTRY_PRIVATE_KEY from a GitHub
Actions secret, signs plugins-index.json with Ed25519, and writes
plugins-index.json.sig alongside it. Aborts loudly when the secret is
missing so a misconfigured CI can't silently ship an unsigned payload.
The workflow now runs build → sign → commit (.json + .sig) → push →
poke worker /refresh. The worker fetches the committed .json + .sig
verbatim and stores both — the daemon then verifies against the
embedded pubkey it ships with.
Closes PR review CRITICAL #1: the worker is no longer a sign-anything
oracle reachable via REGISTRY_REFRESH_TOKEN. Trust root is now this
repo's branch protection + Actions secret scope, not a token any CI
job that can talk to stats.librefang.ai can use to mint signatures.
Note: the keypair was rotated as part of this change (PR not yet
merged so no daemon TOFU pins exist). New pubkey:
ClGa0Ucap8NdrKAy1rw9Tt6A9I8eg4zJ53+xIuKMuq0=
The plugins-index.json.sig committed here is signed with the matching
new private key, in lockstep with the daemon EMBEDDED_REGISTRY_PUBKEY
constant and all three worker [vars] entries.
Pair with the existing plugins-index.json path (which feeds the
daemon's signed install lane). registry-index.json mirrors the
dict-shaped payload the registry-worker's cron currently builds via
40+ GitHub Contents API calls — but built locally from the checked-out
tree by scripts/build-registry-index.mjs, so the worker only fetches
ONE file per category type (2 total: plugins + registry) on refresh.
Workflow now picks up content changes across all 8 category dirs
(was: plugins/ only) so dashboard updates land within seconds of a
push instead of waiting for the 02:00 UTC cron tick.
The dashboard's /api/registry endpoint reads kv_store('registry_data');
the worker's forced-refresh now writes that key with these bytes and
purges the Cache-API entry, so the next dashboard hit sees fresh
data instead of the 1h-cached previous payload.
Generated counts on first build: 11p 17h 32a 60s 44c 57pr 22w 33mcp.
Walking ~40+ plugin TOMLs via the GitHub Contents API from the worker
exceeded the Workers Free 50-subrequest-per-invocation limit, leaving
the daemon's signed plugins index either empty or partial after every
forced refresh.
Move the walk into the repo: scripts/build-plugins-index.mjs reads each
plugins/<name>/plugin.toml directly from the checked-out tree and emits
a sorted flat array (name, version?, description?, needs?) at
plugins-index.json. The CI workflow regenerates and commits this file
on every push under plugins/, then pokes the worker's
/api/registry/refresh — which now fetches the single committed
plugins-index.json (1 subrequest), validates the JSON shape, and
re-signs it with Ed25519. Refresh cost is now constant in registry
size, not linear.
The dashboard's dict-shaped /api/registry payload is unchanged — that
still rebuilds via the daily 02:00 UTC cron.
Without this, dashboard / daemon see content changes only after the
next 02:00 UTC cron tick (up to ~24h delay). The action POSTs to
stats.librefang.ai/api/registry/refresh with a bearer token shared
with the worker secret of the same name; until both secrets exist on
their respective sides, the endpoint returns 503 and the action fails
loud — no silent half-deploy.
Filters on directories the worker actually reads (plugins/, agents/,
skills/, hands/, channels/, providers/, workflows/, mcp/) so README
edits don't burn worker invocations.
Known limitation: the worker rebuild walks ~40+ GitHub Contents API
subrequests, which on Workers Free truncates partway through and
leaves some categories empty. This action fires the right path; the
underlying budget fix (Workers Paid, or pre-building the index in
this repo) is tracked separately.