hive-agent-forge-token and hive-agent-queue-credential both run with
UMask=0377. Their scripts captured bao's stderr in `err="$(mktemp)"`,
which under that umask is created 0400; the very next `2>"$err"` on the
`bao login` line cannot reopen it for writing, so bash fails the
redirect with "Permission denied" before bao ever runs. The `if !`
around the login then took the only error branch it had and printed
"this agent's certificate was refused by the swarm secret store" — the
store was never contacted. No agent has fetched either credential.
The stderr file now lives in each unit's own 0700 RuntimeDirectory and
is removed before every redirect into it, so the redirect creates it —
the idiom forge-token.nix already used for its staging file.
The login's error branch now says which of these happened, then quotes
bao's output:
- `$err` could not be created, so bao never ran;
- the store answered with HTTP 4xx (refusal) or another status;
- the store sent a TLS alert rejecting the certificate;
- no answer at all (network, DNS, or local TLS).
Unreadable cert/key credentials are reported before bao runs.
bao.nix has the same fetch shape but no UMask=, so its mktemp file is
0600 and writable; it is untouched.
Closes#4735
The header comment grew to 39 lines across two audit-driven rounds,
over the 30-line comment-block-lint max. Trimmed to 30: merged the
value-as-argument rationale with its /proc/cmdline justification into
one paragraph, cut the usage example to one call instead of two, and
condensed the args paragraph — no content dropped, just restatement.
persistence.md's first-boot-migration marker paragraph used "it is"
and "there is not" — vale's Microsoft.Contractions rule (this repo's
config) wants the contracted forms, and "is not" also collides with
"is nothing" as a literal substring, which is what actually tripped
the error. Reworded to "it's" / "there's nothing", no meaning change.
Lint-only: no script logic changed, gates re-run below are all lint
checks (no module-eval, no cargo).
Refs #4723
The pipe contract had a gap: if a producer piped into
atomic_write_secret exited non-zero after writing partial output,
cat still saw a clean EOF and wrote that partial content through to
the live target via mv — pipefail only reported the failure
afterward, once the bad write was already committed. The helper now
takes the value as its 4th argument and writes it itself with printf
(a shell builtin, so the value never touches an external process's
own argv/environ, same as a function argument never does), so there
is no pipe left to fail silently.
Callers that compute the value with a command now capture it into a
variable first (`value=$(cmd)`), which fails under `set -e` before
atomic_write_secret is ever called — swarm-bao.nix's pin.env site is
the one that needed this (`pin_env_value="BAO_HSM_PIN=$(cat ...)"`).
All seven call sites converted; output is byte-identical (same
printf '%s\n' framing, now applied inside the helper instead of by
each caller).
Refs #4723
swarm-bao-token's pin.env write (the BAO_HSM_PIN EnvironmentFile for
openbao's pkcs11 seal) had the same write-then-chmod-on-live-path
shape as the sites already converted: printf > path directly on the
live file, chmod after. Same fix, same helper. Content
("BAO_HSM_PIN=<user-pin>\n") and final mode (0400, root-owned — no
chown, same as before) are unchanged; pin.env stays at the same path,
so openbao's EnvironmentFile= reference needs no change.
Refs #4723
atomic_write_secret's cleanup trap used RETURN, which never fires when
set -e aborts the function mid-body (a failing cat/chmod/chown), so a
secret-bearing temp file was left behind instead of being removed.
The write now runs in a subshell with its own EXIT trap, invoked via a
named handler (so `local rc=$?` is a normal, shellcheck-visible
assignment) that only removes the temp file when the subshell's exit
status is nonzero — the subshell's trap table is private, so a calling
unit's own EXIT trap is untouched. Reproduced the leftover-tmp bug
against the prior commit, confirmed it's gone, and confirmed both the
success path and a caller's own EXIT trap still work as before.
swarm-bao.nix's swarm-bao-forwarder-oidc unit had the identical
write-then-chmod-on-live-path defect as the four sites already fixed
here (fetches an OIDC client secret from swarm-bao, printfs it to the
live host path, chowns/chmods after) and was missed by the original
sweep. Converted it to atomic_write_secret; content and final
owner/mode (root:root, 0400) are unchanged.
Refs #4723
Four host glue units fetched a secret from swarm-bao and rendered it
with `> path; chmod`: a reader racing the write could see a truncated
file, and briefly one at the wrong mode before the chmod landed.
glue-matrix-bao-token.nix, glue-queue-agent-credential.nix (both
files), swarm-grafana.nix and swarm-otel.nix now write to a same-
directory temp file, set its final mode/owner, then `mv -f` it over
the target — a shared `atomic_write_secret` helper
(nix/host-modules/lib/atomic-write-secret.nix) so the five call sites
share one implementation.
The first-boot `/root/.claude` migration in nix/agent-modules/user.nix
wrote its done-marker unconditionally, so a failed `cp` (disk full,
permission error) left the marker behind and no boot ever retried the
copy. The marker is now written only when there was nothing to
migrate or the copy succeeded; `cp -an`'s no-clobber semantics already
make a retry after a partial copy safe.
Refs #4723
set_nspawn_flags propagated has_cap's error, so one corrupt
capabilities.json failed every agent's Start, spawn and Swap. It now
goes through holds_manage_root_agent, which logs the error (agent and
file) and treats the capability as absent: the agent starts without the
cross-agent, /applied and /meta mounts. caps_for/has_cap take the file
path so that seam is testable against a tempfile.
- meta.rs: a comment at the render_flake reads records why they
propagate (an empty tool-groups map renders toolGroups = null, i.e.
AGENT_DEFAULT, which fails open for narrower explicit entries).
- capabilities::read doc: states when set_caps/remove_agent rewrite
the file instead of saying remove_agent repairs it.
- set/remove corrupt-file tests assert ErrorKind::InvalidData.
tool_groups::read and capabilities::read returned an empty map when
their file existed but didn't parse. Every set_*/remove_agent is a
read-modify-write, and write() rewrote the file in place, so a crash or
ENOSPC mid-write left a truncated file, and the next write (e.g. the
manager-spawn seed of ruth's tool groups) replaced it with a map holding
only one agent. The scheduling and approval gates then denied every
other agent, recoverable only from meta git history.
- Both registries now read through agent_config::read_map: a missing
file is still the empty map, any other read failure or a parse
failure is an io::Error. set_groups / set_caps / remove_agent fail
without writing.
- Writes go through agent_config::write_map: temp file in the same
directory, fsync, rename, fsync the directory. hive-c0re had no
shared atomic-write helper (the existing tmp+rename sites are inline
and don't fsync).
- Callers of read / groups_for / has_cap now handle the error:
* dashboard GET /api/tool-groups, /api/capabilities,
/api/permissions/stale return 500 instead of an empty table;
* the SSE permission snapshots are skipped with a warn;
* render_flake returns Result, so sync_agents fails instead of
rendering every agent without its tool groups / capabilities;
* set_nspawn_flags propagates has_cap's error;
* the socket tool-group gates deny with the read error as message;
* seed_manager_tool_groups logs and does not seed.
- capabilities::write had no callers left once set_caps writes through
write_map, and is removed.
Closes#4719
`checks.module-eval-core-toggle` is red on main: "a gateway's services
leaf asks for only the swarm names that host fronts" read its leaf
request off `bare`, on the premise that `bare` fronts forge. That was
true where the property was written, before 978164dc (#4705) defaulted
`deploy.forgejo.enable` to false. The two merged in sequence, and on main
`bare` renders only the `_` and `h1.t.local` vhosts, so its
`localServiceDomains` is [] and the script carries `alt_names=''`,
never `alt_names=forge.t.local`.
The narrowing itself is right: a host that runs no forge fronts no
swarm name and asks for no services leaf. Only the fixture was stale.
The property now reads a hive with `deploy.forgejo.enable = true`,
whose request is `common_name=forge.t.local alt_names=forge.t.local`
while `auth.t.local` stays in the swarm-wide set.
The renewal timer from 0649673e is not involved: the check's derivation
is identical at 0649673e and at its parent, and ed2ec52f replayed onto
978164dc^ holds while replayed onto 978164dc it fails.
Refs #4587
`services.hyperhive.enable` and `services.hyperhive.c0re.enable` are gone.
One switch, `services.hyperhive.deploy.hive-controller.enable` (default
false, as the old toggle was), now gates hive-c0re and hive-priv. Both old
paths are `mkRenamedOptionModule` shims in deploy.nix, so a host config
that still sets either evaluates as before and gets a rename warning.
Every other read of the old toggle is resolved, including the 29 made
through the `hyperhiveCfg`/`hiveCfg` aliases:
- Dropped: each swarm service and its glue keeps only its own deploy
toggle (authelia, bao and its PKI glue, grafana, victorialogs,
victoriametrics, the secret publisher, swarm-ca, the OIDC client rows,
the controller/nats/matrix-ctl/publisher/services-issuer identities),
the forge, and the `domain` deprecation warning.
- To deploy.hive-controller.enable: the queue-agent credential reader and
its assertion, which feed hive-c0re and write under its state dir, plus
their policy-order entry; the network identity assertions; hive-tls's
two writes into hive-c0re's environment.
- hive-tls runs where the gateway runs self-signed
(`gateway.enable && useSelfSigned`), not on every host.
- The matrix appservice-token reader and its assertion stay on
`deploy.matrix.enable` plus their client-identity checks. They read
deploy.matrix's token file and registration script; their deploy.bao
inputs are the client-half options a hive sets to read a store it does
not run, so gating on deploy.bao.enable would drop the tested
remote-reader case.
- The `hiveName` assertion moves from hive-network.nix to hyperhive.nix
and fires wherever the hive, the store or the homeserver runs: each
turns the name into an identifier with no fallback.
On a host with `deploy.allSwarmServices` and no hive, the documented
services-host recipe, authelia, bao, grafana, victorialogs,
victoriametrics, the OIDC client rows and the hive CA now render; before,
the old toggle being off left them out.
Refs #4500
Every gateway asked the store's `pki/issue/swarm-services` for the whole
swarm's service set, so a private key on any gateway host could serve a
valid certificate for services that host does not front and never has.
`swarm.localServiceDomains` derives the per-host subset by filtering
`swarm.serviceDomains` against the vhosts this host actually renders —
the deploy flags those vhosts are already guarded on, read once rather
than copied into a second filter. The leaf request and the coverage
guard that decides whether to re-issue both read it, so they cannot
disagree about which names the leaf owes.
The sub-CA's name constraint and the role's `allowed_domains` stay the
swarm-wide set: every host's subset is inside it, and narrowing the
constraint per host would turn one signing into N.
The store's `swarm-services` role issues the services leaf for 720h, and
`swarm-services-cert` only ever ran at boot or rebuild: it is a
`RemainAfterExit` oneshot wanted by `multi-user.target` and no timer
targeted it. A hive not rebuilt within 30 days served an expired leaf.
`swarm-services-cert-renew` runs the same script from a daily timer. It
is a unit of its own because a timer starting the `RemainAfterExit` unit
is a no-op, and restarting that unit instead would propagate through
`hive-gateway-self-signed-cert`'s `Requires=` to nginx, so a sealed store
would take the gateway down over a still-valid leaf. Nothing requires or
orders against the new unit; it has no `Restart=`, so a failure stays in
`systemctl --failed` until the next tick, and the script only moves files
into place after the store has answered.
The re-issue threshold was `checkend 2592000`, the whole 30-day
lifetime, so every run re-issued. It is now half the role's lifetime,
read from a new internal option `deploy.bao.servicesPkiLeafTtlHours`
that the role's `ttl`/`max_ttl` also read. Boot and timer share the
script and so the threshold. The services-root re-check reads the
same option, at the store's own replacement threshold (hours × 3600),
so the hive asks for a new leaf when the store replaces its root. A
`flock` keeps the two runs from interleaving one issuance's key with another's leaf.
`checks.module-eval-hive-tls` pins the timer, that the unit it starts
re-runs the issuance without `RemainAfterExit`, that nothing depends on
it, and that both the leaf and root thresholds move with the option.
Closes#4587
#4708 defaulted deploy.forgejo.enable to false, and the hive-forge module's
whole config block (including the mirror-declared-without-a-controller
warning) now hangs off it. controllerWithMirrors and mirrorsNoController
both relied on the old always-on default to reach that warning path.
The orgs agent-configs/internal/agents (plus mirror owners), the
operators team in agents and agent-configs, the pull-mirrors,
internal/docs, internal/knowledge (public, README-seeded) and the
agent-configs org avatar are one set per forge. hive-c0re ensured them in
its boot sweep, as the core admin, and only on the hive co-located with
the forge container.
swarm-controller now reconciles them at start and every 5 minutes
(forge/objects.rs: observe -> pure plan -> apply). A failed object logs
a warn line plus a pass summary and is retried next tick. create_repo
ensures the agent-configs org and its operators team first, so a config
repo's merge gate never depends on the periodic pass having run.
hive-c0re drops ensure_org, SEEDED_ORGS, ensure_mirrors/ensure_mirror_repo,
ensure_operators_team, ensure_shared_docs_repo, ensure_knowledge_repo/
set_repo_public, seed_readme, ensure_config_org_avatar and the one-shot
knowledge::remove_webhook cleanup, with their now-unused helpers.
nix: the mirror list moves from the hive-c0re unit
(HYPERHIVE_FORGE_MIRRORS) to the swarm-controller unit
(SWARM_CONTROLLER_FORGE_MIRRORS), with an eval warning when mirrors are
declared on a host that runs no controller. c0re.orgAvatarPng is renamed
to deploy.swarm-controller.configOrgAvatarPng.
Refs #3782
The integration, tool, persistence and setup pages described hive-c0re
minting each agent's token into its state dir. They now describe the
swarm's appservice, its admin sender, the controller's mint and pass, the
daemon's store read and re-start timer, and the two new credential rows.
ruth's matrix account comes from the same pass once she holds the store
identity the setup page already has the operator mint.
The swarm mints each agent's `main` account now, so the hive's own mint
goes: `ensure_user_for`, `finish_user_provisioning`, `sync_agent`,
`sync_agent_standalone`, `token_path`, `legacy_password_path`,
`auto_reset_password` and `token_file_present`, and the calls from the startup sweep and the
rebuild bookkeeping. Both mints pinned the device `hyperhive-<agent>`, so
leaving this one would have each re-login kill the other's token.
`hivectl matrix create-user` refuses an agent's name and says where its
account comes from. Everything that still uses the hive's appservice token
stays: the hive's own account, the Space and chat room, and operator
accounts.
The daemon reads each account's token from `swarm/agents/<agent>/matrix/`
as the agent itself, inside its own container, and falls back to the file
only when the store has none. This is #4519's read, without its `main`
carve-out: the swarm now mints `main` there and no hive writes the file.
The daemon unit gets the agent's store identity, spelled the way
forge-token.nix spells it. A timer re-starts it while it is down: a token
the swarm mints or replaces in the store changes no file, so the path
watcher never fires for it, and a daemon that exited on a replaced token
would otherwise stay down until the container restarts.
A `MintAgentMatrixAccount` node creates the agent's account on the swarm's
homeserver with the swarm appservice token, stores its token at
`swarm/agents/<agent>/matrix/main`, and reads it back with whoami before
reporting success. It is a root of agent creation, `after_any` into the
deploy, and a five-minute backfill over every agent with a store identity
queues the same node — the shape of the forge-token mint.
The decision reads the stored token back rather than only checking that one
is stored: the swarm and a hive both pin the device `hyperhive-<agent>`, so
each login replaces the other's token. A failed read plans nothing, so an
outage never rotates every agent's token.
`matrixHomeserverUrl` now defaults to the swarm's `chat.` vhost, since the
mint is what consults it.
The matrix container renders the swarm registration before tuwunel starts
(`requiredBy` it, no network), tuwunel loads it as a second `.yaml`
credential, and a publish unit hands its token to the store. tuwunel 1.9.1
refuses only a duplicate id or as_token, not overlapping non-exclusive
namespaces, so it sits beside the hive's `hyperhive` registration.
`admin_execute` promotes exactly `@swarm` at boot. The module-eval pin
narrows from "no account" to that one list, compared whole so a second
entry fails; `admin_execute_errors_ignore` is pinned too. matrix-ctl may
write the swarm token's leaf, and the controller may only read it.
Everything is gated on matrix-ctl's store identity: with nobody to publish
the token, the registration would be an admin credential nobody reads.
The swarm gets an appservice identity of its own, separate from each hive's
`hyperhive` registration. `swarm-matrix-ctl appservice render` mints its
tokens inside the matrix container when they are absent and renders the
registration tuwunel loads; `appservice publish` writes its as_token to
`swarm/controller/swarm-controller/matrix/appservice-token`, the one kind no
hive's policy grants.
The homeserver calls move out of swarm-matrix-ctl into swarm-matrix-client,
with a `whoami`, so swarm-controller can mint agents' accounts through the
same pinned device id instead of a copy of them.
bao's own in-container collector labelled every log line and metric it
forwards `service.name=swarm-bao` (`processors.resource.attributes`,
keyed off `swarm.bao.machine`), while every panel in the shipped
Grafana bao dashboard queries the literal `service.name="bao"` — no
panel has matched since the scrape moved into that collector.
Rename it in the collector instead of templating the dashboard: a new
`collectorServiceName` binding in swarm-bao.nix, deliberately not
`cfg.machine` (that value names the container/receiver, not bao's
display identity), stamps `service.name="bao"` directly. bao.json is
back to its origin/main shape, unchanged.
Adds a module-eval case to checks.module-eval-grafana that reads the
collector's own evaluated config and asserts its service.name matches
every selector the shipped dashboard uses; verified invert-proof by
setting the value back to cfg.machine and confirming that specific
case (and only it) fails.
setup.md said Swarm SSO creates the operator's forge account, which was
not true until the previous commits. It now says how: sign in to the
forge once through authelia, then `swarmctl forge make-admin <you>`.
sso.md says what that first login does and why ACCOUNT_LINKING is
`login`. README, hivectl.md and forge.md drop `hivectl forge
create-user`, and the swarmctl README gains `forge make-admin`.
Refs #3782
Calls POST /api/forge/users/{name}/admin and prints what it found. It
fails with the controller's message when the user has not logged in via
SSO yet, and when the name is an agent's.
Refs #3782
POST /api/forge/users/{name}/admin reads the account and, when it is not
already a site admin, sets `admin` with admin_edit_user. It never
creates one: a human's account is made by their first authelia login,
so a missing one answers 404, saying the user has not logged in via SSO
yet. An existing admin is a success with nothing sent.
The edit carries `admin` alone. repo_creation_lockdown's login_name +
source_id = 0 would turn an SSO-made account into a local one: in
Forgejo 16 a source_id sets the login type.
An agent's name is refused, and so is any name when the roster can't be
read: a site admin ignores max_repo_creation, the lockdown that keeps an
agent's token from creating a repo and self-merging in it.
Refs #3782
The forge now creates a human's account on their first authelia login,
so the verb has no job left. Deletes it, HostRequest::ForgeCreateUser,
its handler, provision_user_token, change_user_password and the hive's
TOKEN_SCOPES. change_user_password also passed the password as an
argument to `forgejo admin user change-password`, so it showed in the
container's process list.
ensure_user_exists and mint_token stay for the `core` bootstrap, their
one caller now. ensure_user_exists loses its password parameter: only the
deleted path set one.
Refs #3782
[oauth2_client] turns on auto-registration through the authelia login
source, with the account named after authelia's preferred_username.
DISABLE_REGISTRATION stays true: forgejo 16's auto-registration checks
only ALLOW_ONLY_INTERNAL_REGISTRATION, so local sign-up stays off.
ACCOUNT_LINKING is `login`, forgejo's default, set explicitly. With
`auto`, an SSO login whose name matches an existing local account would
be handed that account, and agents, `core` and `swarm-controller` all
have one. `login` asks for that account's own password instead.
Refs #3782
argus review on #4713: the comment said every container block sets
--link-journal=host, and warned that dropping it silently blinds this
receiver. #4713 does exactly that for swarm-bao (its journal stays
inside the container for its own collector instead), so read on its
own the comment told a debugger to revert the fix. State the
exception.
--link-journal=host bind-mounts a host directory journald never
writes into for this container (empty, root:nogroup, confirmed on
the live host — #4527). Dropping it falls back to nixpkgs' default
--link-journal=try-guest, the same shape every agent container
already uses, so the in-container collector's journald receiver
(directory = /var/log/journal) now reads a journal that is actually
written.
merge stays true and services.hyperhive.swarm.otel.journaldUnits is
untouched so this deploy changes exactly one thing; comments that
described the old host-linked shape are rewritten to match.
Adds a module-eval assertion (bao-otel-collector.nix) that the
container's extraFlags never re-add --link-journal=host.
Refs #4527, #4499.
`swarm-bao-forwarder-oidc` fetches the store container's collector secret,
one path, and was the last reader still logging in with
`deploy.bao.clientCertFile`: the hive's own leaf, whose policy reads every
agent's credentials, the hive's tree and every service's OIDC secret. The
four-way split gave grafana's and the swarm collector's readers leaves of
their own and left this one behind.
It now holds `forwarder-oidc.pem`, minted by `swarm-bao-pki`, and logs in
under the `swarm-forwarder-oidc` cert-auth role, whose policy reads
`secret/data/swarm/services/<store forwarder client id>/oidc/client` and
nothing else. The role is written by `swarm-bao-forwarder-oidc-policy`
from the bootstrap token, which gains the two grants that unit calls, and
the reader is ordered after it. The subject is reserved as a hive name. A
store host whose pair is null is refused at eval rather than falling back
to the hive's leaf.
The hive's own role and `client.pem` are untouched; nothing is revoked.
hive-priv writes an agent's matrix-token 0600 and owned by the agent,
so hive-c0re, running as hive-core, cannot read it. The read-based
token-present guard in ensure_user_for therefore never fired, and every
matrix sweep (boot, every 30 minutes, every rebuild) re-minted each
agent's token through the appservice login and restarted its
hive-matrix-daemon.
Decide presence with a stat instead: a non-empty regular file counts as
present. hive-core can stat the file through the 0755 state dir.
Closes#4665
Every hive with hyperhive enabled ran its own hive-forge container, and
its gateway answered forge.<swarm> with its own bridge IP, so on a
multi-host swarm each hive talked to its own forge.
deploy.forgejo.enable defaults to false and allSwarmServices sets it with
mkDefault, like authelia and bao; singleHostSwarm gets it through that.
The forge's OIDC client moves to a glue module gated on authelia, so a
split authelia/forge swarm still registers it. CI now requires the forge
on the same host, and the controller's forgeTokenFile defaults to null
where the forge is not.
Closes#4705
Refs #3782
The stub-client helper in forge.rs built Client { api } after #4703
added a url field, breaking compilation of swarm-controller's test
target on main. Clone the URL the helper already builds for
Forgejo::new so it can also populate Client.url.
Closes#4709
crash_watch's 10s poll and auto_update's ensure_root_agent both read
lifecycle::list().await.unwrap_or_default(), which turned a failed read
into 'zero containers'. In crash_watch that made every previously-running
agent look like it crashed simultaneously (prev.difference(current) over
an empty current), and left prev empty for the next cycle too, so a
second wave of false 'agent logged in' / 'agent needs login' events fired
against the next successful read. In ensure_root_agent it read as
'manager container missing' and called lifecycle::spawn on a manager
that might already exist.
Both sites now treat a list error as its own outcome: log it at warn and
skip the cycle's decision entirely. crash_watch leaves prev exactly as
the last good read produced it. ensure_root_agent attempts no spawn.
Factors each site's decision into a pure helper (plan_cycle /
plan_root_agent) matching the check_not_live / confirm_gone_after_failed_destroy
pattern, with unit tests for the error case, a control for the readable
case, and (for crash_watch) an invert-proof run locally against the old
unwrap_or_default logic before reverting.
The rewritten ruth step skipped her store identity, without which
neither the backfill nor her container's fetch can reach her token. With
the backfill now creating a missing forge user, two commands cover her:
swarmctl agent mint-identity ruth --hive <hive>, then hivectl agent ruth
rebuild so hive-c0re hands the identity to her container.
Refs #3782
An agent with a hive-agent-* store identity but no forge user was
observed as NoForgeUser and dropped by plan(), so it never got a token.
plan() now keeps it, and queue_forge_token_mints inserts CreateForgeUser
ahead of MintAgentForgeToken with after_ok, the edge declare_agent_job
already uses. ensure_agent_user folds an existing user into success, so
the extra node is a no-op for agents that have one.
Refs #3782
credentials.md gains the forge-token row and drops the claim that the
forge token never passes through the store. setup.md says plainly that
an agent spawned on the hive alone, ruth's bootstrap included, now gets
no forge user from anything. CLI references regenerated.
Refs #3782
Delete ensure_user_for and mint_and_persist_agent_token, the user step
of sync_agent (the per-rebuild re-mint, #4644) and of
forge_after_first_spawn, and the hive-priv WriteAgentForgeToken request
that wrote the token into the agent's state dir. hivectl forge
create-user now refuses an agent and points at swarmctl agent
mint-forge-token. mint_token, ensure_user_exists and TOKEN_SCOPES stay:
provision_user_token and the core bootstrap still call them.
Refs #3782
forge-token.nix fetches swarm/agents/<agent>/forge-token under the
agent's own store identity into /run/hive-agent-forge-token/token, and
re-fetches on a timer so a rotation lands. hive-forge, the git
credential helper, hive-forge-notify, forge-avatar-sync and the web UI
read that file first and fall back to <state>/forge-token.
tea-login is deleted: it copied the token into ~/.config/tea, which
docs/swarm/credentials.md forbids for a store secret. hive-forge covers
the same verbs. swarmctl gains agent mint-forge-token.
Refs #3782
A MintAgentForgeToken node mints a fixed-name swarm-agent token with the
admin API, keeps it when the stored value's last eight and the normalised
scopes match the forge's list, and otherwise deletes and re-creates it.
The token is stored at swarm/agents/<agent>/forge-token. Agent creation
inserts the node, and a pass at start and every five minutes inserts it
for every agent holding a store identity whose token is missing or stale.
Refs #3782
disable_repo_creation's ownership guard now matches either the aligned
{agent}@hyperhive.local email or the legacy {agent}@hive.local one
hive-c0re::forge::users::ensure_user_exists used before its own
ensure_user_email alignment pass existed. That pass only runs on the
forge-host hive, and only once it has a core token and has ticked, so a
real pre-existing agent can still carry the old email when this node reads
it. Without this, the guard would bail "not this agent's" on a genuine
agent during that rollout window (argus, round 2).
Verified separately (not a code change): swarm-controller's forge account
is a site admin (created with --admin), and Forgejo's
convert.toUser/ToUser only hides an account's email when the caller isn't
the admin and isn't the account itself (services/convert/user.go), so
user_get already returns the real email regardless of hide_email on the
target account. No endpoint change needed for that half of the review.
disable_repo_creation now reads the account once before PATCHing
max_repo_creation/source_id, and fails the node if the email isn't the
{agent}@hyperhive.local marker create_agent_user itself sets. The 409/422
create-fold (#4681) only proves some account with that name exists, not
that this node created it, so a pre-existing non-agent account sharing an
agent's chosen name could otherwise get locked onto local auth with repo
creation disabled. Leaves the fold untouched (Option B, per atlas/argus on
#4693); the read moves into disable_repo_creation instead.
Also fixes the nix-sandboxed cargo-test check: forgejo_api::Forgejo::new
builds a reqwest client that eagerly resolves TLS roots via
rustls-native-certs even for the tests' plain-http loopback stub server,
which panics with "No CA certificates were loaded from the system" in the
CA-less build sandbox. Gives that check's nativeBuildInputs pkgs.cacert and
sets SSL_CERT_FILE, same pattern this repo's runtime deployment already
uses for the same reqwest/rustls resolution.
`ensure_agent_user` (the `CreateForgeUser` node) created the agent's
forge account but never set `max_repo_creation = 0`, relying on
hive-c0re's per-hive `ensure_repo_creation_disabled` pass to lock it
down later. That pass is being removed (#3507, #4669), and it is the
only guard against an agent token creating, owning and self-merging in
its own repo.
The node now PATCHes `max_repo_creation = 0` via `admin_edit_user`
after the create, on both the created and the already-exists path, and
a refused PATCH fails the node with Forgejo's message. The body mirrors
hive-c0re's `sparse_edit_user_option` (`login_name` + `source_id = 0`,
everything else unset).
The PATCH is unconditional: Forgejo's API never returns
`max_repo_creation` (it is in `EditUserOption` only, not `User`), so
there is no current value to verify against first.
Closes#4689
list_repos_with_open_issues only read one repo_search page (forgejo's
30-row default), so a repo past position 30 in the default alpha sort
silently dropped out of the issue-report/repo-dropdown data source, with
no truncation signal. The comment claiming RepoSearchQuery has no page
field was wrong -- Request::page()/page_size() are generic builder
methods independent of the query struct.
Adds page_search_results, a small paging loop over a fetch closure for
search-shaped (data: Option<Vec<T>>, no header) responses that can't use
the existing .all() helper (that needs a (Headers, Vec<T>) response
shape). Pages until a short page or a 40-page bound, erroring on the
bound rather than truncating again.
Closes#4675
A remote hive dialled nothing until an operator copied the queue's URL
into it, though the URL is the same string everywhere. statusPublish.natsUrl,
queue.agentNatsUrl and controller.queue.natsUrl now default to
tls://<swarm.nats.domain>:<port> unconditionally.
The statusPublish assertion treated a URL without a secret as a half
config. With the URL a default on every hive, only the secret claims
publishing: the assertion now refuses a secret without a URL or token
endpoint, and hive-c0re's status environment is gated on the secret too,
so a hive without one publishes nothing instead of reading a missing
credential.
swarm-bao-nats-tls-policy acts with the bootstrap token, and main's
module-eval-bao-grants now fails any such unit whose calls the policy
file does not grant. Adds its three paths and counts it among the units
the check must see.
The queue listened in plaintext on 4222, reached by bridge IP or loopback,
and nothing in-tree opened it to another hive. It now has a name, serves a
certificate for that name alone, and refuses clients that do not speak TLS.
- `swarm.nats.domain`, default `nats.<swarm.domain>`, a sibling name like
`swarm.bao.domain`. The queue host answers it via `gateway.localNames`;
every other hive resolves it through the operator's DNS, as for bao.
- `pki/roles/swarm-nats` allows that one name (bare domain, no subdomains,
IPs or localhost, server flag). A `swarm-nats` cert-auth role and policy
may only `update` `pki/issue/swarm-nats`, written by
`swarm-bao-nats-tls-policy`. The login leaf is minted by glue-bao-tls and
paired by glue-nats-bao-identity. `deploy.bao.natsCommonName` is reserved
as a hive name.
- `swarm-bao-nats-tls` issues the leaf into a directory bound read-only into
the container, restarts nats when it rotates, and re-runs daily.
It joins glue-bao-readers-policy-order, so it is ordered after its policy
unit (`after` and `wants`, never `requires`) where the store is on the
same host. The policy unit joins the store's journald list.
- nats gets `tls {}`, with the key via `LoadCredential`, and no
`allow_non_tls`. `validateConfig` is now off in every mode, because the
build-time check loads a leaf that only exists at runtime.
- 4222 is also open on `wg-hive` when the host is on the mesh, never
host-wide.
- `statusPublish.natsUrl`, `queue.agentNatsUrl`, the controller's URL under
`singleHostSwarm`, and the auth responder all dial
`tls://<swarm.nats.domain>:<port>`. swarm-queue-client hands its CA file
to the NATS connection too, so hive-c0re and the controller trust the
leaf's root.
- docs/swarm/README.md: the queue URL and the one DNS record a multi-host
swarm needs.
module-eval-nats-tls pins the role, the policy, the served leaf, the
firewall, the ordering, and a scan of every `*_NATS_URL` and the
responder's URL across the host and its containers.
Closes#4626
get_issue_report mapped every error from Client::issue_report to 503
(StatusUnavailable) unconditionally. issue_report calls
issue_list_issues(org, repo, ...) first, and a Forgejo 404 there — a
typo'd or deleted repo — was flattened into the same 503 a genuine
forge outage produces, which tells a client to retry a request that
will never succeed.
Downcast the anyhow error back to forgejo_api::ForgejoError (same
pattern main.rs's wanted_error_status uses for
swarm_queue_client::Error) and check it structurally against the
ApiErrorKind::NotFound / UnexpectedStatusCode(404) shapes forgejo's
generated client produces for a 404, rather than string-matching the
rendered message. Only that case answers 404, naming the org/repo;
every other forge failure still answers 503. Updates the route's
OpenAPI response list to document the 404.
Closes#4701