Skip to content

fix: security hardening — bind, validation, atomic writes, deploy rollback - #1

Closed
fernandogzp wants to merge 1 commit into
jaylfc:masterfrom
fernandogzp:fix/security-hardening
Closed

fix: security hardening — bind, validation, atomic writes, deploy rollback#1
fernandogzp wants to merge 1 commit into
jaylfc:masterfrom
fernandogzp:fix/security-hardening

Conversation

@fernandogzp

Copy link
Copy Markdown

Summary

Security and reliability improvements based on a thorough code review. All changes are backwards-compatible and all 127 existing tests pass.

  • Default bind 127.0.0.1 — server, install.sh, and config.yaml all defaulted to 0.0.0.0, exposing the unauthenticated dashboard to the entire network. Changed to localhost-only; users who need network access can explicitly set host: 0.0.0.0 in config
  • Agent name validation — added ^[a-z0-9][a-z0-9-]{0,62}$ regex check on create and deploy endpoints, preventing malformed names from reaching incus CLI
  • Framework allowlist — deploy endpoint now rejects unknown frameworks instead of pip install-ing arbitrary user input inside containers
  • Atomic config writessave_config() now writes to .yaml.tmp then renames, preventing corruption if the process crashes mid-write
  • Deploy rollback — if any step fails after container creation, the container is automatically destroyed instead of being left half-configured
  • QMD serve hardening — added NoNewPrivileges=true to the systemd unit template
  • Removed hardcoded IPs — Tailscale IPs and LAN addresses removed from committed data/config.yaml

Test plan

  • All 127 existing tests pass (pytest tests/ -v)
  • Manual test: deploy an agent with an invalid name → should get 400
  • Manual test: deploy with framework: "malicious-package" → should get 400
  • Manual test: verify server only listens on localhost after fresh install

🤖 Generated with Claude Code

…es, deploy rollback

- Default server bind: 0.0.0.0 → 127.0.0.1 (config.py, install.sh, data/config.yaml)
- Agent name validation: regex ^[a-z0-9][a-z0-9-]{0,62}$ on create/deploy
- Framework allowlist: deploy rejects unknown frameworks instead of pip-installing arbitrary input
- Atomic config writes: write to .tmp then rename to prevent corruption on crash
- Deploy rollback: destroys container on any failure after creation
- QMD serve inside containers: added NoNewPrivileges=true to systemd unit
- Removed hardcoded Tailscale/LAN IPs from committed config.yaml

All 127 existing tests pass.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@jaylfc

jaylfc commented Apr 5, 2026

Copy link
Copy Markdown
Owner

Thanks for the thorough review and solid PR @paralizeer. I've cherry-picked the security fixes I wanted and committed them in 824f8cf.

Merged (with credit):

  • Atomic config writes — write to .yaml.tmp then rename. Clean and prevents corruption.
  • Agent name validation — the regex check is important since names go into incus CLI commands.
  • Deploy rollback — try/except with container cleanup on failure. No more orphaned containers.
  • NoNewPrivileges=true on the qmd serve systemd unit inside containers.
  • QMD serve binds 0.0.0.0 inside container — correct, the host needs to reach it.
  • Removed hardcoded IPs — good catch, these should never have been committed. I've also moved config.yaml to config.yaml.example and gitignored the actual config along with hardware.json and installed.json.

Not merged (with reasoning):

  • Default bind 127.0.0.1 — TinyAgentOS is designed to run on LAN devices (SBCs, home servers) where users access the dashboard from another machine on the network. Binding to localhost would break this for most users out of the box. Since there's no auth yet (trusted LAN assumption), the UX tradeoff favours accessibility. Auth is on the roadmap and will address this properly.
  • Hardcoded framework allowlist — instead I validate against the app catalog registry dynamically. This means new frameworks added to the catalog are automatically allowed without code changes.

Closing this PR since the relevant changes have been applied. Appreciate the contribution — the atomic writes and deploy rollback in particular are the kind of reliability fixes that matter on embedded hardware where crashes happen.

@jaylfc jaylfc closed this Apr 5, 2026
jaylfc added a commit that referenced this pull request Apr 11, 2026
Cherry-picked from @paralizeer's PR #1:
- Atomic config writes (write to .tmp then rename)
- Agent name validation (alphanumeric + hyphens, 1-63 chars)
- Deploy rollback on failure (destroys container on error)
- NoNewPrivileges=true on qmd serve systemd unit
- QMD serve binds 0.0.0.0 inside container (host needs access)

Modified from PR:
- Keep 0.0.0.0 as default bind (this is a LAN device, not a public server)
- Dynamic framework validation from catalog registry instead of hardcoded allowlist

Additional fixes:
- Move config.yaml to config.yaml.example (template)
- Gitignore data/config.yaml, hardware.json, installed.json
- Auto-copy example config on first run
- Remove hardcoded Tailscale IPs from deployer defaults
- Remove committed agent configs and backend URLs
jaylfc added a commit that referenced this pull request Apr 18, 2026
…gration

Running prisma db push from our bootstrap created tables without seeding
_prisma_migrations, so LiteLLM's own prisma migrate deploy at startup
tried to apply migration #1 against an already-populated schema and
looped on "type JobStatus already exists", leaving the proxy unhealthy.

Our helper's only job now is to make prisma.client importable so
LiteLLM can run its shipped migrations itself. Drop the db push and
the psql/psycopg probe; keep the systemd PATH fix for prisma generate.
jaylfc added a commit that referenced this pull request Apr 18, 2026
…race, pre-built base image (#225)

* refactor(adapters): defer uvicorn imports so modules load without it

Adapters imported uvicorn at module top, so anything that imported
them for structural checks (tests, health-endpoint probes) would
crash with ModuleNotFoundError when uvicorn wasn't installed.
uvicorn.run is only needed when an adapter is run as a standalone
process — move the import into the __main__ guard.

Clears 19 pre-existing test failures across test_new_adapters.py
and test_channel_hub_new.py.

* feat(containers): add rename_container to the backend abstraction

Required by the new agent archive lifecycle: the delete path stops
the container then renames it to a dated `taos-archived-{slug}-{ts}`
bucket so a later restore can rename it back. Implemented for both
LXC (incus rename) and Docker (docker rename).

* feat(deployer): expanded container deps and manifest-aware framework install

Broken before: `pip3 install openclaw` ran unconditionally, its failure
was logged as a warning and the deploy continued, and the container
came up missing the deps most agent frameworks need.

Now:
- apt install includes nodejs, npm, build-essential, python3-dev,
  ca-certificates, gnupg, wget, with DEBIAN_FRONTEND=noninteractive
  and --no-install-recommends (timeout 15m for slow arm64 apt).
- Framework install dispatches on the manifest's install.method:
  pip uses manifest.install.package, script pushes + runs
  manifest.manifest_dir / install.script. Missing script files,
  unsupported methods, and non-zero exits all raise RuntimeError
  so the outer try/except rolls back the container and the agent
  shows status=failed instead of misleadingly 'running'.
- TAOS_MODEL env var is injected so the in-container runtime knows
  which model to send to LiteLLM.

* feat(litellm): kilocode + openrouter support, per-agent model registration, hot reload

generate_litellm_config now:
- Registers openrouter (openrouter/ prefix, native LiteLLM support)
  and kilocode (openai-compatible, explicit api_base) in the
  backend type maps.
- Expands each cloud backend's declared models into their own
  model_list entries keyed on the real model id, so agents can
  request a specific model. The 'default' alias is still appended
  as a fallback.

routes/providers.py: add/patch/delete now call proxy.reload_config
instead of the stale proxy.write_config, so the running LiteLLM
subprocess actually picks up config changes.

* feat(openclaw): install.sh and in-container agent runtime

The manifest declares method: script -> scripts/install.sh, which
didn't exist. The deployer has no way to install openclaw, so the
agent came up with no runtime and the chat path had nothing to hit.

The new script, run once inside a fresh Debian bookworm LXC:
- Creates /opt/openclaw with a pinned venv (fastapi, uvicorn,
  httpx, openai).
- Writes a minimal FastAPI runtime at /opt/openclaw/server.py that
  listens on 0.0.0.0:8100, accepts POST /message {text, from,
  thread_id?} and forwards to LiteLLM using the injected
  OPENAI_BASE_URL, OPENAI_API_KEY, and TAOS_MODEL env vars.
- Installs a systemd unit so the runtime survives restarts.
- Polls /health up to 20s and fails the install if the server
  didn't come up.

No memory, no tools, no persistence — the host owns all of that.
This is the minimum for the end-to-end chat pipeline to land
messages on an agent and get a reply back.

* feat(agents): uuid identity, persisted fields, DM channel, archive lifecycle

Several related changes to the agents API and config model that
together make agent creation survive the full round trip:

- Every agent gets a stable 12-char uuid (agent['id']), backfilled
  for existing config entries by normalize_agent.
- body.model and body.framework land on the agent row at create
  time; llm_key lands after the background deploy succeeds.
- A 1:1 DM channel is auto-created on successful deploy and its
  id persisted as chat_channel_id so the Messages app sees the
  agent immediately.
- extra_config to deploy_agent now always includes the app
  registry so the manifest-aware framework install can resolve.

Delete is now archive, not destroy. DELETE /api/agents/{name}:
stops the container, renames it to taos-archived-{slug}-{ts},
moves workspace/memory dirs under data_dir/archive/{slug}-{ts}/,
revokes the LiteLLM key, flags the DM channel archived, and moves
the config entry from config.agents to config.archived_agents.

New endpoints:
  GET  /api/agents/archived           -> list archive entries
  POST /api/agents/archived/{id}/restore -> reverses the archive
  DELETE /api/agents/archived/{id}    -> true permanent purge

Restore handles slug collisions (if a new agent has taken the
original name) by suffixing -2, -3, etc. Purge is what the old
hard-delete used to do: destroy container, rm -rf archive dir,
delete chat channel, drop the archived entry.

This also fixes 'can't re-create a deleted agent with the same
name' -- the old delete path left the LXC container around; the
new archive path renames it out of the way.

* feat(chat): route user DM messages to in-container agent runtime

User messages in a DM channel now reach the agent's FastAPI runtime
on port 8100 inside its LXC container; the reply is persisted as
an agent-authored message in the same channel and broadcast over
the chat hub so both the webapp and the PWA update in real time.

Wiring:
- AgentChatRouter (new, tinyagentos/agent_chat_router.py):
  fire-and-forget dispatch(message, channel). Skips non-user
  messages, looks up each non-user channel member as an agent,
  skips agents that aren't running (posts a short system reply
  instead), and POSTs to http://{agent.host}:8100/message with
  {text, from, thread_id}. Response content is written back via
  chat_messages.send_message. All errors caught -- broken agents
  don't crash the chat path.
- routes/chat.py: one-line dispatch call at the end of the HTTP
  post_message path and the WebSocket 'message' branch, so both
  entry points route identically.
- app.py: router instantiated in the lifespan after chat_hub.

No subscription plumbing, no retries -- the router is a direct
adapter between two owned stores. Timeouts and connect errors
become visible agent replies so the user sees what went wrong.

* feat(agents-ui): Archived section with Restore and Delete Permanently

Adds a collapsible 'Archived' panel below the live agents list in
AgentsApp. Shows each archived entry's display name, model, and
relative archive time; per-row Restore and Delete Permanently
buttons call the new backend endpoints with confirmations.

- parseArchiveTimestamp / relativeTimeFromTs helpers convert the
  YYYYMMDDTHHMMSS format the backend writes.
- ArchivedAgentsPanel is inlined (matches AgentRow / DeployWizard
  living in the same file) and self-hides when there are no
  archived entries.
- handleDelete's confirm copy now mentions archiving so users
  know it's recoverable.
- fetchArchived is called alongside fetchAgents at every existing
  refresh point.

Unit tests for the new helpers under
desktop/src/apps/__tests__/AgentsApp.archived.test.tsx.

* build: rebuild desktop bundle with Archived section and agent wiring

* feat(containers): stop --force flag + add_proxy_device across backends

Incus refuses to rename a running container, so every archive call now
has a hard dependency on the container being stopped first. The new
stop-force path sends --force (LXC) or kill (Docker) so archive can
guarantee the container is down before it attempts the rename. The
add_proxy_device method is added to the abstract base, LXC backend, and
Docker stub so the deployer can attach incus proxy devices for host-side
port forwarding when setting up an agent home.

* feat(deployer): agent-home mount, incus proxy devices, taos_host=127.0.0.1

Agents now get a dedicated home directory mounted at /root inside the
container so their runtime state (env file, model config, logs) persists
across container recreation. Proxy devices are attached via the new
add_proxy_device backend method so the host-side LiteLLM process can
reach the in-container agent port. The taos_host default is hardened to
127.0.0.1 so freshly deployed agents always resolve back to the host
loopback rather than relying on a potentially incorrect network variable.

* fix(litellm): is_running honours adopted instances so deployer mints keys

Previously is_running only checked the subprocess handle, which is None
for processes the deployer did not start itself (adopted instances). The
method now also checks the _adopted flag so that a pre-existing LiteLLM
process is correctly reported as running and a fresh API key is minted
rather than the deployer trying to start a second instance. The companion
reload_config path also skips process management when adopted.

* feat(openclaw): write env to /root/.openclaw/env under mounted home

The install script previously wrote configuration to a path that was
wiped on container recreation. Now the env file is written to
/root/.openclaw/env which sits inside the persistent agent-home mount,
so credentials and model config survive container restarts and upgrades
without reinstalling. The script also accepts values from environment
variables so the deployer can inject them at provision time.

* feat(agent-env): host-side helper to rewrite the agent env file in place

A small module that locates and rewrites the env file inside the
agent-home directory on the host without entering the container. This
is used by the restore path so a freshly issued LiteLLM key and updated
endpoint can be injected into the persistent /root/.openclaw/env without
having to reinstall the framework, which would risk breaking the agent's
installed state.

* feat(archive): force-stop, abort on rename failure, carry home dir, rewrite env on restore

Archive now force-stops the container before rename — incus refuses to
rename a running instance, and silently leaving it running produced
orphan containers. If the rename itself fails, the config entry is
left in the live list rather than being moved to archive, keeping the
system consistent. The agent home directory now travels with workspace
and memory into the archive bucket so the full /root is preserved.
On restore, the new host-side env rewrite helper updates
/root/.openclaw/env with the freshly issued LiteLLM key and endpoint
rather than reinstalling, which avoids breaking the installed framework.

* feat(auth): local-token Bearer auth for programmatic access

Adds a stable local token that is bootstrapped once at startup and
written to a known path with mode 0600. Any request carrying it in an
Authorization: Bearer header is granted full access without a session
cookie, allowing automated agents and local scripts to call the API
without going through the browser-based login flow. The middleware sits
before the session check so it has no impact on normal browser sessions.

* docs: per-agent home mount in framework-agnostic-runtime

Documents the agent-home directory layout and mount strategy so it is
clear what lives inside the container, what is persisted on the host,
and how the env-rewrite helper fits into the restore flow.

* fix(lifecycle): arm keep-alive timer on image generation

notify_task_complete was never called from the image generation route,
leaving the RKNN SD server running indefinitely after requests completed.
The _legacy_generate path (used when the resource scheduler is absent)
now wraps the backend HTTP call in try/finally so notify_task_complete
fires on both success and error paths. Chat and embedding traffic routes
through LiteLLM and does not hit this endpoint; that keep-alive path is
handled separately via a LiteLLM callback.

Tests added for both success and failure paths.

* feat(models): /api/models/loaded probes rknn-sd + sd-cpp backends

Image-gen backend types were silently skipped in loaded_models, causing
the Activity widget to show "Loaded Models (0)" even when the RKNN SD
server was active. Two new branches added to the probe loop:

- rknn-sd: GET {url}/v1/models (rknn_sd_server.py speaks OpenAI-compat),
  emits one entry per model with purpose=image-generation.
- sd-cpp: GET {url}/sdapi/v1/options, reads sd_model_checkpoint for the
  active checkpoint name, falls back to "unknown" if absent.

Both branches follow the existing ConnectError/Timeout/HTTPError swallow
pattern. Tests cover success, missing-checkpoint fallback, and offline
(connection refused) for both backend types.

* feat(trace): per-agent hourly-bucketed trace store for the librarian

Captures go per-agent in the bind-mounted home folder so archive,
restore, backup, and cross-worker migration all work via the existing
"move the home folder" rule. Each agent's .taos/trace/ directory holds
one SQLite bucket per UTC hour (YYYY-MM-DDTHH.db). Bucket routing is
driven by the event's created_at, not wall-clock at write time -- a
14:59:59.999 event routed at 15:00:00.001 lands in the T14 file, so
rollover never drops events.

Zero-loss: every write lands in the SQLite or is appended to a sibling
YYYY-MM-DDTHH.jsonl. Nothing is ever silently dropped. The librarian
merges both sources at read time.

The envelope is v1 and stable: v, id, trace_id, parent_id, created_at,
agent_name, kind, channel_id, thread_id, backend_name, model,
duration_ms, tokens_in, tokens_out, cost_usd, error, payload. Kinds are
enumerated (message_in/out, llm_call, tool_call/result, reasoning,
error, lifecycle); each has a documented payload shape so consumers
parse without guessing. trace_id + parent_id enable cross-event linkage
for reconstructing a full turn end-to-end.

POST /api/trace writes; GET /api/agents/{name}/trace reads with filter
+ limit. POST /api/lifecycle/notify lets the LiteLLM callback reset the
keep-alive timer for whichever backend served a request.

* feat(litellm): CustomLogger callback posts llm_call traces + keep-alive

Registered in generated litellm_config.yaml under
general_settings.custom_callbacks. Runs inside the LiteLLM subprocess
with no access to taOS's Python state, so it authenticates to taOS via
the local token file on disk and posts over HTTP to /api/trace and
/api/lifecycle/notify.

Agent name is derived from the virtual key alias ("taos-<slug>") that
the deployer sets when minting per-agent keys. This is how the per-
agent trace store knows which bucket to route to for a given completion.

Failure-mode is swallow-and-log: a broken callback must never fail a
real LLM request. A litellm-not-installed environment gets a no-op stub
so tests pass without the dep.

* feat(trace): wire TraceStoreRegistry in lifespan + pass auth env to containers

app.py: instantiate the registry on the data_dir, include the trace
router, close all connections on shutdown.

deployer.py: inject TAOS_LOCAL_TOKEN (read from data_dir/.auth_local_token
at deploy time) and TAOS_TRACE_URL into the container env. Any in-
container runtime that wants to post traces (or that we later replace
with real openclaw and tap via gateway events) has the credential and
endpoint ready.

* docs(design): update framework-agnostic-runtime for proxy networking + archive/trace

Switches env-snippet from host.docker.internal to 127.0.0.1 and explains
incus proxy devices. Drops Docker-only qualifier from workspace/memory status
table (LXC now has parity). Adds Per-agent trace capture, Agent archive/restore,
and Programmatic access (local token) sections. Extends Related and adds a
Related code list pointing at the new modules.

* docs(design): cross-reference per-agent trace layer from user-memory

Adds a section distinguishing user memory (long-lived user context) from
per-agent trace capture (event log inside agent-home). Explains how the
taOSmd librarian bridges both layers and links to the trace design in
framework-agnostic-runtime.md.

* docs(runbook): add agent archive, restore, and purge runbook

Step-by-step procedures for archiving a live agent, listing archives,
restoring with slug collision handling and LiteLLM key rotation, and
permanent purge. Covers failure modes including container rename failure,
archive dir collision, and restore container conflicts.

* docs(runbook): add trace querying runbook for API, SQLite, and cost attribution

Covers the three-endpoint trace API surface with curl examples, query
filter parameters, envelope field table, kind/payload reference, direct
SQLite access pattern, cost attribution recipe, and librarian consumption
pattern. Links to trace_store.py, routes/trace.py, and litellm_callback.py.

* docs(design): update plan-agent-deployer API table to reflect archive semantics

DELETE /api/agents/{name} now archives rather than hard-deletes. Updates the
endpoint table to show the archive path and the new purge endpoint.

* docs(design): openclaw-integration.md, bridge adapter as MVP

Primary reference for the real openclaw integration: gateway protocol
breakdown, install + runtime, config schema, extension model, known
limitations, 35-row capability map, and a 4-phase MVP-to-full roadmap.

MVP path is the bridge adapter from the 2026-04-11 framework-integration
-bridge-design spec, not the operator-client (raw v3 WS from taOS). The
operator-client is kept as a documented fallback only, because it
couples taOS to openclaw's gateway protocol version and any upstream
bump can break the fleet; the bridge isolates coupling to a single
~200 LoC patch inside our jaylfc/openclaw fork.

Review-gate refinements baked into Step 1: feature-flag the patch
entry (so an unset TAOS_BRIDGE_URL gives upstream-identical builds),
version-stamp the bootstrap, single coupling discipline, channels.kind
"external" + provider "taos" (upstreamable), 400 LoC patch ceiling,
automated persistence-audit as the trust anchor, parallel upstream PRs,
LiteLLM key rotation caveat (safe-on-restart today, reload RPC later).

Fixes the Debian-bookworm Node 18 gap (install Node 22.14+ via
NodeSource before npm install) and the stale 500MB manifest disk size
(real openclaw is 1-2GB on disk).

Appendix B lists 12 docs.openclaw.ai pages that 404'd at research
time; a follow-up pass using gh api on the repo docs/ tree fills those
gaps.

* docs(design): fill openclaw-integration gaps via gh api on the repo

Resolved 2 of 7 Appendix A open questions from primary source on
github.com/openclaw/openclaw. Struck through 9 of the 12 404'd
docs.openclaw.ai URLs in Appendix B where the repo had a mirror. Added
<!-- source: ... --> comments so future readers know which claims are
primary-sourced.

MVP impact: startup health-check loop updated from ss fallback to
`openclaw health --timeout` (Q5 resolved); gateway.bind: "lan" confirmed
as the correct key for container external binding (Q1 resolved).

* feat(trace): seal historic bucket files read-only after 2h rollover

Trace files older than 2h are chmod'd 0o400 during eviction so the
librarian's source-of-truth for historic agent activity is tamper-proof
on-disk. Rare late-arriving events (clock skew, deferred processing)
route to a sibling {bucket}.late.jsonl which stays writable -- zero-
loss guarantee preserved even for the extreme edge. list() merges
.db + .jsonl + .late.jsonl with dedup by event id (primary wins).

Sealing runs opportunistically inside _evict_old_buckets; no
background task.

* feat(openclaw): bridge endpoints — bootstrap, SSE events, reply ingestion

Three endpoints the openclaw fork patch (src/taos-bridge.ts) calls:

  GET  /api/openclaw/bootstrap           config snapshot at startup
  GET  /api/openclaw/sessions/{a}/events SSE stream of user messages
  POST /api/openclaw/sessions/{a}/reply  deltas, final, tool events, errors

BridgeSessionRegistry holds one queue per agent; chat router enqueues
user messages, openclaw subscribes and reads them, replies flow back
through /reply which writes to the per-agent trace store and broadcasts
via the chat hub (message_delta for streaming, edit_message+state
complete for final).

Bearer local-token auth on all three endpoints.

* feat(openclaw): install.sh uses real openclaw from jaylfc fork + Node 22

Replace Python FastAPI stub with real openclaw npm install from the
jaylfc/openclaw fork (taos-fork branch). Installs Node 22.x via NodeSource
since Debian bookworm ships Node 18. Bumps manifest disk_mb to 2000 to
accommodate the Node runtime. Pinned to upstream main SHA be7a415eb096.

* feat(deployer): write openclaw.json + .openclaw/env into agent-home at deploy

Before create_container runs, deployer writes the openclaw gateway config
and the bridge env file into the host-side agent-home directory. The
agent-home bind-mount carries them into /root/.openclaw/ inside the
container, where the systemd unit created by install.sh picks them up.

session_id == req.name (slug) for MVP; bridge endpoints already key on
agent name so no separate UUID is needed at this stage.

* refactor(chat): agent_chat_router enqueues to bridge session, not HTTP POST

Obsolete: raw HTTP POST to container :8100/message was the Python-stub
integration. Real openclaw talks to taOS via the bridge adapter: taOS
owns the SSE stream openclaw's fork patch subscribes to, and replies
come back through POST /api/openclaw/sessions/{agent}/reply (which
broadcasts to the chat hub and writes traces).

So the router shrinks to: on a user message in an agent's DM channel,
call registry.enqueue_user_message(slug, msg). That's it. All reply
plumbing lives in routes/openclaw.py.

* fix(deployer): use bind=instance for incus proxy devices to avoid host port conflict

When litellm is already running on 127.0.0.1:4000 on the host, adding
an incus proxy device with the default bind_mode tries to re-bind that
port on the host and fails with EADDRINUSE.  Setting bind=instance
makes incus bind the listen address inside the container instead, so
host services that already own the port are not disturbed.

* fix(install): chown npm cache before global install to fix EACCES in fresh containers

Debian's apt-installed nodejs leaves /root/.npm with mixed ownership,
causing npm install -g to fail with errno -13.  Fixing ownership before
the install is the documented npm fix for this condition.

* fix(install): remove stale npm cache before global install

Debian's apt npm (v8) creates /root/.npm with state that blocks the
newer npm shipped with Node 22.  rm -rf before the install is more
reliable than chown since the issue is the cache format, not just ownership.

* fix(install): pre-create npm cache dir and use --unsafe-perm for root installs

The Debian npm post-install creates /root/.npm with problematic ownership.
rm + mkdir ensures a clean dir; --unsafe-perm suppresses the root-cache
check in older npm versions that remain in the system PATH during install.

* fix(deployer): remap container root to host process uid via raw.idmap

Agent-home directories are owned by the taOS process user (uid 1000).
Without a UID mapping, incus containers run as an offset uid (100000+)
that cannot write to those host dirs, causing npm and other tools to
fail with Permission Denied when creating files under /root.

Setting raw.idmap 'both <host_uid> 0' before attaching mounts maps
container root to the host process owner so bind-mounted dirs are
writable.  Requires one stop/start cycle after setting the idmap.
Revert the earlier npm --unsafe-perm workaround; it was masking this.

* fix(install): use HTTPS URL for npm install instead of github: shorthand

The github: prefix makes npm resolve to git+ssh://... which requires
GitHub SSH keys that fresh containers do not have, causing hangs.
Using git+https:// avoids the SSH key requirement.

* fix(install): use tarball URL instead of git+https to avoid SSH fallback

npm's git+https:// handling still falls back to SSH for github.com repos.
Using the tarball URL (https://github.com/.../tarball/<branch>) downloads
a plain HTTPS tarball and avoids git transport entirely.

* feat(host-firewall): install systemd one-shot to ACCEPT incusbr0 through docker DROP

Docker installed on the same host as incus sets iptables FORWARD
policy to DROP and adds its own ACCEPT rules only for docker bridges.
Incus-created containers (taOS agents) fall through to the default
DROP for TCP sessions Docker's chains don't claim -- causing symptoms
like github.com unreachable from inside containers while npmjs.org
(cached via Cloudflare) works.

The Docker-blessed fix is to insert ACCEPT rules into the DOCKER-USER
chain. A one-shot systemd unit does this at boot, idempotently. The
fix is reversible (ExecStop removes the rules). install.sh on the Pi
drops the scripts into /opt/tinyagentos/scripts/ and enables the unit
before tinyagentos.service comes up, so containers have working
networking on first boot.

* feat(host-firewall): path unit, timer, subnet probe, connectivity check

Hardens the existing host-firewall oneshot against the realistic install
matrix with Docker in the picture:

  - tinyagentos-host-firewall.path: fires the oneshot whenever
    /var/run/docker.pid appears, so a user who apt-installs Docker
    after taOS gets their rules reapplied without rebooting.
  - tinyagentos-host-firewall.timer: 5-minute re-assertion
    (belt-and-braces for any Docker chain churn we didn't expect).
  - scripts/incus-bridge-probe.sh: detects incusbr0 IPv4 collisions
    with pre-existing bridges and reassigns to a free RFC1918 /24
    at install time; idempotent no-op on clean hosts.
  - install.sh: calls the probe after incus init, then runs a
    throwaway ephemeral container that curls github.com and
    registry.npmjs.org as a post-install connectivity smoke test.
    Warns but does not block on failure so users can diagnose and
    retry.
  - host-firewall-up.sh gains a --check mode (exit 1 when rules are
    missing, used by the timer for visibility) and skips cleanly on
    hosts without incus.
  - host-firewall.service now uses ConditionPathExistsGlob so
    iptables-nft-only systems skip instead of failing.
  - detect_runtime() in containers/backend.py logs the selection
    + alternatives every call and makes the LXC-preferred policy
    explicit in its docstring.

Policy: LXC is the preferred agent runtime; Docker coexists for the
app store's containerised services but never takes precedence over
LXC for agents.

* docs(coexistence): LXC / Docker coexistence policy + runbook

Companion to the host-firewall systemd unit / path / timer.
Documents:

  - Why LXC is the preferred agent runtime and Docker coexists
    without handicapping it.
  - Clash matrix (iptables FORWARD, subnet, chain re-ordering,
    first-boot race, host port, runtime install ordering, cgroups).
  - Install scenarios -- what happens end-to-end for fresh Debian,
    Docker-before-taOS, taOS-then-Docker, Docker restart.
  - Operational runbook: diagnosing a silent network failure,
    adding / removing Docker from a running taOS host, smoke
    testing coexistence.
  - Runtime selection policy: detect_runtime() prefers LXC,
    Docker path exists only as fallback.

framework-agnostic-runtime.md's existing Host firewall subsection
links to this doc for the full story.

* chore(gitignore): exclude AI-assistant artefacts + new superpowers plans/specs

Keeps the repo looking fully human-authored per project policy. CLAUDE.md,
GEMINI.md, AGENTS.md, OPENHANDS.md, .claude/, .aider*, .continue/,
.cursorrules, .cursor/, .copilot-instructions.md, .windsurf/,
.playwright-mcp/, and new docs/superpowers/plans/ + specs/ now stay local.

Existing tracked files under docs/superpowers/ remain tracked (changing
that would rewrite history); only new additions are ignored.

* fix(host-firewall): correct systemd condition syntax + remove path-unit cycle

- tinyagentos-host-firewall.service used a single-line space-separated
  ConditionPathExistsGlob which systemd evaluated to no-match, silently
  skipping the unit. Split into two ConditionPathExists= directives with
  the "|" OR-prefix so the unit runs when iptables exists at either path.

- tinyagentos-host-firewall.path declared After=tinyagentos-host-firewall
  .service alongside the implicit Wants= from [Path].Unit=, producing an
  ordering cycle that systemd rejected. Removed the After= — the path
  unit activates the service on path-change; systemd handles ordering.

* feat(messages): archived channels section + dead-agent grey-out

- Add ArchivedChannel type extensions (settings with archived_at, archived_agent_id)
- Fetch /api/chat/channels?archived=true alongside live channels on init
- Fetch /api/agents and /api/agents/archived for author resolution
- Collapsible "Archived" section in sidebar (desktop + mobile) with per-row
  Restore (RotateCcw) and Delete Permanently (Trash2) hover actions
- Restore calls POST /api/agents/archived/{id}/restore; disabled with tooltip
  when archived agent entry is missing
- Delete calls DELETE /api/chat/channels/{id} with confirmation
- Opening archived channel shows full message history
- Archived banner above composer; input + send disabled for archived channels
- resolveAuthorDisplayState() pure helper: maps author_id to active/archived/removed
- Greyed author names (opacity 0.55, strikethrough) with tooltip for dead agents
  message body remains fully readable at reduced opacity; no body strikethrough
- Accessibility: aria-expanded/aria-controls on collapsible, aria-label on all buttons

* build: rebuild desktop bundle with archived-chats UI

* feat(chat): archived filter on /api/chat/channels + ensure_message helper

channel_store.list_channels() gains an `archived: bool | None` param —
None = no filter (existing callers unchanged), True = only channels with
settings.archived truthy, False = only channels where it's falsy.

MessagesStore.ensure_message(msg) idempotently re-inserts a message by
id (INSERT OR IGNORE). Used by the restore path when re-importing a
chat-export.jsonl; callers can retry safely.

routes/chat.py::list_channels forwards the new query param through.

* feat(archive): export chat to agent-home + reimport on restore + purge channels

At _archive_agent_fully: iterate DM channels the agent is a member of,
dump every message to {agent-home}/{slug}/.taos/chat-export.jsonl (one
envelope per line, 0o600). Channel settings flagged with
{archived: true, archived_at, archived_agent_id, archived_agent_slug}.

At restore_archived_agent: if a chat-export exists, stream it back into
chat_messages via ensure_message() (idempotent by id); unflag every
channel where archived_agent_id matches this archive.

At purge_archived_agent: delete every channel's messages + the channels
themselves for channels flagged with this archive_id, then rm -rf the
archive bucket as before. Irreversible.

Chat history now travels with the agent's home folder — archive /
backup / restore / cross-worker migration all carry it automatically.

* docs(design): accept architecture pivot v2 — 10 decisions resolved

Turning the §10 open questions into §10 resolved decisions:
  1. btrfs storage pool (portable across the mixed-arch cluster via
     each host's Linux layer).
  2. archive.target configurable, default pool:
  3. chat history in both tarball + global DB (already shipped).
  4. snapshot preferred, rsync fallback.
  5. Garage primary S3 NAS + optional FUSE POSIX mount per agent.
  6. Recycle-bin Layer 1 + Layer 3 only (skip libtrash LD_PRELOAD
     in Phase 1).
  7. Forgejo (not Gitea).
  8. taos-archive-<ts> snapshot prefix; auto-snapshots untouched.
  9. Garage sled metadata backend (cross-arch portable).
  10. 40 GiB default per-agent quota; per-agent override.

Status flipped from Proposal to Accepted. Phase 1 (disk quota +
recycle bin) is the next code to land.

* feat(disk-quota): host-side monitor + resize API + threshold notifications

Adds DiskQuotaMonitor with btrfs/incus/df sampling priority chain,
threshold-transition notifications (ok/warn/hard), hard-threshold agent
pausing, and live quota resize via incus. HTTP surface at
GET /api/agents/{name}/disk, POST /api/agents/{name}/quota,
POST /api/disk-quota/scan. NotificationStore gains disk_quota event type.
31 new tests cover edges, transitions, pause behaviour, and routes.

* feat(disk-quota): systemd timer + install.sh wiring

Adds tinyagentos-disk-quota.service (oneshot calling disk-quota-scan.sh)
and tinyagentos-disk-quota.timer (every 5 min, starts 2 min after boot).
install.sh installs the script to /opt/tinyagentos/scripts/ and enables
both units after the existing host-firewall block.

* docs(recycle-bin): runbook for the soft-delete system

Covers how the per-container recycle bin works, trash-cli ops,
escape hatches, what is not covered (Layer 2/3), and admin ops.

* feat(admin-prompts): library of structured admin tasks + HTTP endpoint

tinyagentos/admin_prompts/*.md — initial library:
  - disk-audit: agent audits own disk, proposes [DELETE]/[MOVE-TO-NAS]
    /[KEEP] per item, user confirms
  - memory-audit: RSS + cache inspection with proposed cleanups
  - health-report: read-only status + error-log summary
  - weekly-summary: trace-driven self-report, tokens + cost + highlights

Each prompt forces the agent to check current date first, list
proposed actions before executing, and require user confirmation for
every destructive step. Reminders point at /usr/local/bin/rm (recycle-
bin soft-delete) over /usr/bin/rm.

GET /api/admin-prompts enumerates, GET /api/admin-prompts/<name>
returns the body so the Messages composer can prefill (Phase 1.B UI
lands separately).

* feat(fs-snapshot): Snapper backstop on btrfs pools — Layer 3 recycle-bin

Detects incus storage pool driver on install; if btrfs, installs and
configures Snapper with config name taos-containers, hourly x24 + daily x7
retention. ZFS and dir backends skip gracefully with a clear message.
Wired into install.sh after the host-firewall block; non-fatal if the
optional backstop fails. Includes probe script for operator diagnostics
and a bash structural test.

* docs(fs-snapshot): runbook for Layer 3 recycle-bin backstop

Covers what Layer 3 does versus Layers 1/2, how to verify snapper is
running, listing snapshots, file restoration from btrfs snapshot paths,
disabling the timers, and storage cost expectations.

* feat(recycle): list/restore/purge API routes for container recycle bins

GET  /api/agents/{name}/recycle — list an agent's /var/recycle-bin/
GET  /api/recycle                — aggregated view across all agents
POST /api/agents/{name}/recycle/restore  — restore one item
DELETE /api/agents/{name}/recycle/{id}   — permanent purge

Items are exposed with a base64url id derived from the original path
so the frontend doesn't need to hold a mapping table. All container
interactions go via the existing exec_in_container abstraction, so
the same code works for LXC and Docker backends transparently.

Offline containers return status=container_offline with an empty list
instead of a 5xx — the UI can render "agent is stopped" without
coupling to agent state.

* feat(frameworks): two-tier beta/alpha verification status; openclaw first

Consolidates framework verification statuses from four tiers (tested/beta/experimental/broken)
to two (beta/alpha/broken). openclaw is the only beta entry; all others become alpha.
Updates adapter registry, all 15 agent manifests, and framework route tests with regression
guard against the retired "experimental" status.

* feat(agents-ui): order openclaw first; Beta + Alpha labels in framework picker

Sorts the framework list so openclaw appears at the top. Updates the Framework
interface union type to beta|alpha|broken. Replaces the "beta"/"experimental"
pills with "Beta" (amber) and "Alpha · Testing" (neutral). The show-alpha toggle
and deselect-on-hide logic are wired to the new alpha status.

* build: rebuild desktop bundle with framework picker reorder

* feat(ui): disk-quota card + admin-prompt prefill + recycle-bin browser

Phase 1.B: per-agent disk quota pill (warn/hard) in AgentRow; notification
cards above agent list with Expand +10 GB and Audit with agent actions.
Cross-app navigation via taos:open-messages CustomEvent to MessagesApp.

Phase 1.C-frontend: MessagesApp listens for taos:open-messages, selects
channel, fetches GET /api/admin-prompts/{name}, stuffs body into composer,
shows dismissible prefill banner above input area.

Phase 1.E.2: Recycle Bin location in FilesApp sidebar; fetches GET /api/recycle,
grouped by agent, per-item Restore (POST) and Delete Permanently (DELETE) with
confirm dialogs; empty state and container-offline notice.

* build: rebuild desktop bundle with Phase 1 frontend

Rebuilt after disk-quota card, admin-prompt composer prefill,
and recycle-bin browser features.

* feat(containers): set_root_quota + root_size_gib on create_container

Add set_root_quota(name, size_gib) to the module-level API (__init__.py),
LXCBackend, and DockerBackend. Add root_size_gib param to create_container
in all three; quota is applied after launch, before mounts/env.

Docker overlay2-without-pquota returns success with a soft note rather
than a hard failure. Abstract base class updated with the new abstract
method signatures.

Tests cover success, overlay2 soft path, genuine failure, and pass-through
from create_container for both LXC (module-level) and Docker backends.

* refactor(deployer): snapshot-model -- single trace mount, no workspace/memory/home bind mounts

Replace the three host-side bind mounts (workspace, memory, home) with a
single trace mount: {data_dir}/trace/{slug}/ -> /root/.taos/trace/.

Remove _write_openclaw_bootstrap; install.sh now writes
/root/.openclaw/openclaw.json + .openclaw/env inside the container via
env vars injected at create_container time (TAOS_BRIDGE_URL, OPENAI_BASE_URL,
OPENAI_API_KEY, TAOS_LOCAL_TOKEN, TAOS_AGENT_NAME).

Add root_size_gib=40 to DeployRequest (default per arch pivot v2 S10.10)
and pass it through to create_container.

Update test_deployer.py: remove old three-mount + openclaw host-write
assertions, add test_one_trace_bind_mount, test_no_workspace_memory_home_mount,
test_root_quota_passed_through_default, test_root_quota_custom_value_honoured,
test_trace_dir_created_on_host, test_bridge_url_injected_into_env.

agent_env.py is NOT deleted: routes/agents.py (Phase 2.B) still imports
update_agent_env_file for the restore path. Phase 2.B will clean it up.

* feat(openclaw): install.sh writes /root/.openclaw config + env inside container

Add section 2a between npm install and the recycle-bin block. Uses env vars
injected by the deployer (TAOS_AGENT_NAME, TAOS_MODEL, OPENAI_BASE_URL,
OPENAI_API_KEY, TAOS_BRIDGE_URL, TAOS_LOCAL_TOKEN) to write:

  /root/.openclaw/openclaw.json  (mode 600) — gateway + LiteLLM provider config
  /root/.openclaw/env            (mode 600) — EnvironmentFile for systemd unit

Both files live inside the container rootfs and travel with snapshot archives.
Safe defaults via := fallback for dev/test environments where not all vars
are set.

Remove the old section 3 that only ensured the .openclaw dir existed (no
longer needed; section 2a creates it). Renumber comments for sections 3 and 4.

* feat(containers): snapshot_create + snapshot_restore + snapshot_list + set_env

* refactor(trace): store path moved to {data_dir}/trace/{slug} — bind-mount target

_agent_trace_dir now returns data_dir/trace/slug instead of
data_dir/agent-home/slug/.taos/trace. Aligns with the Phase 2.A
deployer bind-mount (data_dir/trace/{slug}/ → /root/.taos/trace/).
Updates module docstring and framework-agnostic-runtime.md to reflect
the new path and the rationale for separating trace from home-folder.

* feat(migrate): script to move legacy agent-home/*/.taos/trace bucket files

scripts/migrate-trace-paths.sh walks data_dir/agent-home/*/
for .taos/trace directories and moves .db/.jsonl files to
data_dir/trace/{slug}/. Idempotent: skips if source absent, no-clobber
merge if both old and new paths exist. install.sh runs it as a
non-fatal step on every install/upgrade after disk-quota setup.

* refactor(agents): archive/restore/purge use incus snapshot primitives

* feat(config): archive.target configurable (pool | path | s3)

* refactor(env): delete obsolete tinyagentos/agent_env.py -- functionality moved into install.sh + incus set_env

* docs(design): framework-agnostic-runtime thesis evolution — containers hold their own state

Rewrite framework-agnostic-runtime.md to reflect the Phase 2.A–2.C
post-pivot reality: containers hold their own state, hosts hold the
federation. The three bind mounts (workspace/memory/home) are gone;
the single trace bind mount remains. The "Per-agent home" section is
replaced by "Three bind mounts removed" and "agent_env.py removed"
migration notes. The rule application checklist gains a sixth question
on archive atomicity. The "Why the pivot" section summarises the
reasoning. The audit table and Related code section are updated to
match what shipped.

Add a "Post-landing status" banner to architecture-pivot-v2.md marking
the decision record complete with commit ranges for Phases 1 and 2.A–2.C.

Add a note to lxc-docker-coexistence.md that Docker's lack of incus
snapshot primitives means graceful fallback on that backend.

* docs(runbook): rewrite archive/restore/purge for snapshot primitives

Full rewrite of agent-archive-restore.md: archive flow now documents
incus snapshot create, chat export to host-owned path, and snapshot_name
in config. Restore flow documents incus snapshot restore, set_env for
new LLM key, systemctl restart openclaw. Purge documents incus delete
--force destroying snapshots atomically.

Adds Quick reference table, archive.target options table, and legacy
migration section for pre-Phase-2 entries that have no snapshot_name
(identify via jq filter, purge via existing DELETE endpoint, or contact
dev for manual re-snapshot). Troubleshooting covers snapshot-not-found,
rename collision, env-rewrite failure, and openclaw restart failure.

* fix(containers): env setter key=value form; root-quota via override for profile-inherited devices

Two related deploy blockers:

1. `incus config set <name> environment.<key> <value>` parses a value
   starting with `-` as a CLI flag. Local auth tokens from
   `secrets.token_urlsafe(32)` can legitimately start with `-` or `_`,
   so deploys failed with `unknown shorthand flag: 'X'` whenever the
   token landed on that character. Switch to the single-argument
   `environment.<key>=<value>` form, which is the documented canonical
   syntax and parses unambiguously.

2. `incus config device set <name> root size=...` rejects root
   devices inherited from a profile with
   `Device from profile(s) cannot be modified for individual instance.`
   Switch to `incus config device override` which creates a per-instance
   copy of the device if it doesn't exist, then falls back to `set` if
   an override is already present.

Tests cover: dash-prefixed token value, root quota on profile-inherited
device.

* fix(openclaw): npm install via github: shorthand so prepare lifecycle builds in container

We were using the tarball URL form which skips npm's prepare lifecycle.
That pushed the build burden onto the fork branch (committing prebuilt
dist/), which kept landing as incomplete and crashing openclaw.mjs:178.

Switch to `npm install -g github:jaylfc/openclaw#taos-fork`. This
triggers prepare → pnpm build:docker in the destination (container),
which is npm's standard mechanism for git-sourced packages that need
building. corepack-activate pnpm before install.

Adds 2-3 minutes to first deploy on arm64 Pi for the build step.
Reliable; removes the brittle committed-artefact dance.

* fix(openclaw): ensure git installed before npm github: shorthand install

npm install -g github: requires git to clone the repo. Fresh Debian
bookworm containers don't have git; add a guard to install it if
absent before the npm install step.

* fix(openclaw): install prebuilt tarball from GitHub Releases (no per-deploy build)

Per-deploy builds are not beginner-friendly (slow on arm64, fragile
because of pnpm workspace context, depends on Pi having full build
toolchain present and configured). Switch to downloading prebuilt
tarballs published by the fork's CI workflow as Release assets.

Architecture detection picks arm64 vs x64. URL is the GitHub-stable
'releases/latest/download/<asset>' redirect so we always grab the
freshest build with no version bumps in this repo.

Failure mode: if download fails, install.sh exits non-zero with a
clear error — there is intentionally no build fallback. The fix when
GitHub is unreachable is to fix connectivity, not to silently start
building.

* fix(openclaw): add --ignore-scripts to npm install-g from prebuilt tarball

Tarball already has dist/ built by CI. Running prepare at install time
tries to spawn git (for hook config) then falls back to pnpm build:docker;
both fail in a fresh container that has no git/pnpm. The bin entry
(openclaw.mjs) is wired directly — prepare output is not needed.

* fix(openclaw): re-add git prerequisite for libsignal transitive dep

libsignal (@whiskeysockets/baileys dep) has a git+https URL so npm
needs git at install time even when installing a prebuilt tarball.
No build happens — this is purely npm fetching a dependency.

* fix(openclaw): update openclaw.json config to match v2026.4 schema

- providers must be a record keyed by provider id, not an array
- gateway.mode: local must be set or gateway refuses to start
- models field must be an array (empty ok), not default_model string

* fix(openclaw): defer service start until llm_key written to config

openclaw gateway calls the bootstrap endpoint on startup which requires
llm_key in the agent config. Previously install.sh started the service
immediately, causing HTTP 409 crash-loop before the deployer had written
the key. Now:
- install.sh enables the unit but defers start (no --now)
- _background_deploy() starts openclaw.service after writing llm_key
  and saving config, so bootstrap succeeds on first attempt

* fix(openclaw): allow null llm_key in bootstrap, fix providers schema in response

- bootstrap now returns 200 when llm_key is null (no LiteLLM proxy case)
  instead of HTTP 409 which caused the gateway to crash-loop
- bootstrap response uses providers-as-record format to match openclaw
  v2026.4 config schema (same fix as install.sh)

* fix(openclaw): bootstrap uses built-in litellm provider + LITELLM_API_KEY

- Return models.providers.litellm (not taos) with correct openclaw schema
- models[] built from agent.model + fallback_models, each as {id,name,contextWindow,maxTokens,input,reasoning}
- agents.defaults.model.primary = "litellm/<agent.model>"
- Drop default_model field; use \${LITELLM_API_KEY} substitution instead of raw key
- Revert e356a98 empty-string fallback: null llm_key returns 409 with clear message
- Update bootstrap shape assertions + add fallback-models length test + null-key 409 test

* fix(deployer): inject LITELLM_API_KEY + TAOS_MODEL/TAOS_FALLBACK_MODELS env

- Add LITELLM_API_KEY env var for openclaw's litellm provider (same value as per-agent virtual key)
- Keep OPENAI_API_KEY set to same value as compat shim for smolagents and other frameworks
- TAOS_MODEL always set (empty string when unconfigured, not omitted)
- Add TAOS_FALLBACK_MODELS env var (comma-separated) so install.sh can build models[] at install time
- Add fallback_models field to DeployRequest dataclass
- Set LITELLM_API_KEY="" fallback when no proxy configured, matching OPENAI_API_KEY pattern

* fix(openclaw): install.sh writes openclaw.json with litellm provider schema

- Provider name changed from taos to litellm; baseUrl points to 127.0.0.1:4000 (no /v1 suffix)
- apiKey uses \${LITELLM_API_KEY} env-var substitution for openclaw's runtime resolution
- models[] array built at install time from TAOS_MODEL + comma-separated TAOS_FALLBACK_MODELS
- agents.defaults.model.primary set to "litellm/<TAOS_MODEL>" prefix
- env file extended with LITELLM_API_KEY and TAOS_FALLBACK_MODELS entries

* fix(deployer): correct DeployRequest field order (fallback_models after data_dir)

* fix(llm_proxy): own port 4000, drop adoption, log /key/generate failures

Adopting a pre-existing LiteLLM on :4000 silently reuses whatever config
and master key the foreign process booted with, so UI-added providers
never take effect and /key/generate rejects the Bearer sk-taos-master
with 401 (silent return None). Agents then deploy with llm_key=null and
openclaw crashes with LITELLM_API_KEY missing.

start() now terminates any foreign PID on the port (SIGTERM, 5s grace,
SIGKILL) before spawning its own process. is_running() reports ownership
only. reload_config() drops its adopted branch. Admin calls that can
return non-200 now log status + body so master-key mismatches surface
in logs instead of being swallowed.

* fix(llm_proxy): emit master_key + deployer fallback to shared key when no DB

LiteLLM's /key/generate requires Postgres; SQLite is unsupported. On the
default single-user Pi deployment there's no DB, so LiteLLM runs in
routing-only mode and cannot issue per-agent virtual keys. Previously
create_agent_key returned None in that mode and the deployer set
LITELLM_API_KEY="", which crashed the openclaw gateway on boot with
"LITELLM_API_KEY is missing or empty".

Routing-only mode is now the supported default path:
- general_settings.master_key added to the yaml config
- LITELLM_MASTER_KEY exported into the subprocess env
- Single source of truth TAOS_LITELLM_MASTER_KEY constant
- Deployer falls back to the master key when virtual-key issue fails,
  so the container always gets a usable auth token

Users who configure Postgres later still get per-agent virtual keys
through the same /key/generate path.

* fix(providers): postgres-backed virtual keys + generic provider catalog + model discovery

- LLMProxy accepts database_url; app reads data/.litellm_db_url at boot
  and exports it as DATABASE_URL into the litellm subprocess so
  /key/generate can mint per-agent virtual keys.
- Add Provider fills canonical URL from PROVIDER_URL_DEFAULTS and probes
  {url}/models to populate the model list when empty — generic across
  openai, anthropic, openrouter, kilocode (no per-type branching on the
  probe). Falls back to per-type seed list (kilocode → kilo-auto/free)
  when the probe returns nothing so the entry still registers at least
  one routable model.
- Deployer scopes the minted virtual key to the agent's primary + fallback
  models (models=[req.model, *fallback_models]) instead of defaulting to
  the unrestricted "default" alias.
- Deployer fails loudly when a DB is configured but /key/generate still
  returns None — hiding that class of failure is what shipped the
  broken kilocode path in the first place.
- generate_litellm_config now WARNs when a cloud-type backend is missing
  url or models, so silent drops surface in logs instead of showing up
  as a broken agent much later.
- scripts/repair_providers.py repairs legacy config.yaml entries that
  pre-date the autofill/discovery logic.

* fix(llm_proxy): resolve api_key_secret values into subprocess env

Generated LiteLLM configs use os.environ/<name> markers to reference
provider api keys, but nothing was actually exporting those names
into the subprocess env. Cloud providers therefore hit the litellm
OpenAIException "api_key client option must be set" even with a
correctly-configured backend list.

LLMProxy.start/reload_config now accept a secrets={name: value} map.
app.py resolves each backend.api_key_secret from the secrets store at
boot and again on catalog-change reload; routes/providers.py does the
same on add/patch/delete so newly-added or rotated keys take effect
without a full app restart.

* feat(litellm_migrate): auto-apply Prisma schema on boot when DB is configured

LiteLLM's /key/generate requires a Postgres-backed Prisma schema, but
LiteLLM does not run migrations itself. Fresh installs had to manually
run `pip install prisma && prisma generate && prisma db push` before
virtual keys worked.

New tinyagentos/litellm_migrate.py locates the bundled schema at
litellm/proxy/schema.prisma, probes for LiteLLM_VerificationToken in the
configured DB, and shells out to the venv's prisma CLI only when the
table is missing. Idempotent — safe on every boot. Called from the
lifespan hook before LLMProxy.start() so LiteLLM sees a ready schema.

Added prisma>=0.11.0 to the proxy optional dependency group so the CLI
lands in the venv on fresh installs.

* fix(litellm_callback): wire callbacks under litellm_settings + sibling shim for get_instance_fn

* feat(providers): /api/providers/models passthrough with refresh + ttl cache

* feat(agents): agent-creation model picker reads from LiteLLM passthrough

* fix(litellm_migrate): prepend venv bin to PATH so prisma-client-py resolves under systemd

* fix(litellm_migrate): psql probe fallback so boot doesn't wrongly rerun migration

* fix(litellm_migrate): only run prisma generate, let LiteLLM own DB migration

Running prisma db push from our bootstrap created tables without seeding
_prisma_migrations, so LiteLLM's own prisma migrate deploy at startup
tried to apply migration #1 against an already-populated schema and
looped on "type JobStatus already exists", leaving the proxy unhealthy.

Our helper's only job now is to make prisma.client importable so
LiteLLM can run its shipped migrations itself. Drop the db push and
the psql/psycopg probe; keep the systemd PATH fix for prisma generate.

* fix(llm_proxy): prepend venv bin to PATH and widen startup wait

LiteLLM's proxy_cli shells out ``subprocess.run([\"prisma\"])`` during
startup to detect whether Prisma is runnable. Under systemd the service's
default PATH doesn't include our venv's bin/, so the lookup raises
FileNotFoundError and LiteLLM prints "prisma package not found" and skips
DB setup entirely — leaving virtual-key issuance broken even though the
package IS installed in the venv.

Prepend the venv bin that already hosts the litellm binary so the child
process resolves ``prisma`` (and ``prisma-client-py`` for generate).

Also bump the startup wait from 30s to 120s: LiteLLM on a fresh Pi DB
runs ``prisma migrate deploy`` before opening its HTTP port, which takes
45-60s on ARM.

* fix(llm_proxy): capture LiteLLM stderr to sibling log file

stderr=DEVNULL silently swallowed proxy startup failures (prisma
migration errors, config parse errors, model-router failures),
turning "why is the proxy unhealthy?" into a 30-minute debugging
hunt. Write stderr to a file next to litellm_config.yaml so
operators can read it without attaching strace.

* fix(llm_proxy): poll health/readiness and drop SIGHUP reload

Two separate bugs kept LiteLLM from ever settling on the Pi.

1. Startup polling hit ``/health``, which gates on the master key and
   returns 401 for an unauthenticated client. LiteLLM was healthy within
   ~50s but ``start()`` kept polling until the 120s timeout, logged
   "failed to start within 120s", and returned False even though the
   subprocess was fine. ``/health/readiness`` is the public endpoint.

2. ``reload_config`` sent SIGHUP to trigger a config reload. LiteLLM
   runs as single-worker uvicorn (no ``--workers``), which does not
   register a SIGHUP handler, so the default action — terminate — fires.
   Every ``/api/providers/models?refresh=true`` was silently killing
   the proxy, then ``_fetch_litellm_models`` got connection-refused and
   returned []. Drop SIGHUP entirely; the existing stop+start path was
   already the fallback.

Also switch the foreign-process probe to ``/health/readiness`` for the
same 401 reason.

* fix(llm_proxy): forward TAOS_LOCAL_TOKEN to LiteLLM subprocess

The TaosLiteLLMCallback running inside the LiteLLM subprocess POSTs
llm_call events back to the taOS bridge at ``/api/trace``, which
requires the local auth token. The callback's token-discovery logic
checks ``TAOS_LOCAL_TOKEN`` env first, then ``/data/.auth_local_token``
and ``~/.taos/.auth_local_token``. Under systemd the real token lives
at ``{data_dir}/.auth_local_token`` — none of the candidate paths — so
every callback fired a POST without Authorization and taOS responded
401, leaving trace rows with no ``llm_call`` events despite LiteLLM
actually processing requests.

Read the token in app.py and forward it via the new ``local_token``
constructor kwarg on LLMProxy, which exports it into the subprocess env.

* fix(litellm_callback): extract agent slug from user_api_key_metadata

LiteLLM 1.83.4 surfaces the agent slug in litellm_params.metadata under
user_api_key_metadata.agent (matching what LLMProxy.create_agent_key
writes when minting the virtual key). The previous extraction read
metadata.key_alias which is no longer populated on success events, so
every llm_call trace was bucketed under the _unknown_ sentinel slug.

Walks four sources in priority order:
  1. user_api_key_metadata.agent
  2. user_api_key_auth_metadata.agent
  3. user_api_key_alias (strips the taos- prefix)
  4. key_alias (legacy, kept for older LiteLLM builds)

* feat(trace): record message_in events so transcript captures both sides

enqueue_user_message now writes a message_in trace event under the
agent's slug, following the ENVELOPE_V1_SCHEMA message_in shape
({from, text}) with extra informational fields (message_id,
author_type, delivery).

Guards against orphan _unknown_ or empty-slug entries.
Fails soft: trace write errors are logged, never raised.

* feat(agents): persist optional emoji on agent record + deploy API

* feat(agents): emoji picker in create flow + display in agent UI (rebuilt PWA)

* fix(agents): tolerant DELETE for orphan agents — skip snapshot/stop when container absent (#221)

Failed deploys leave behind a config row with no LXC container, which
caused DELETE /api/agents/{name} to error on snapshot_create. Probe
container_exists first; for orphans, skip stop/snapshot, revoke any
LiteLLM key, and either hard-delete the row (no history) or record a
tombstone (chat/trace present so purge is available from Archived).

Adds container_exists helper to tinyagentos.containers; four new tests
cover the orphan hard-delete, orphan tombstone, skipped-snapshot
assertion, and purge of a snapshotless tombstone.

* feat(agents): pre-built openclaw LXC base image for fast deploys

Adds a GitHub Actions workflow that builds per-arch Debian 13 LXC
base images with Node 22, openclaw, and recycle-bin scaffolding
already installed. Published as assets on the 'rolling-images'
Release tag.

The deployer now checks for the 'taos-openclaw-base' image alias
before launching; when present it uses the cached image and sets
TAOS_BASE_IMAGE_PRESENT=1 so install.sh skips the apt-get + npm
steps. Without the image the deployer falls back transparently
to images:debian/bookworm and install.sh does the full install.

tinyagentos.agent_image exposes is_image_present and
ensure_image_present helpers; the latter runs as a background task
on app startup to bootstrap the image on first boot.

Closes #220

* ci(agents): fix bridge forwarding + NAT for incus in GHA runner

* fix(agent_image): use os.pipe() so curl stdout actually reaches incus stdin

The previous impl passed curl.stdout (a Python StreamReader) as
stdin= to the incus subprocess, which asyncio cannot forward as
an OS-level FD. Curl would read the first ~90KB then block on a
pipe nobody was draining. Using an explicit os.pipe() pair with
the read end handed to incus and the write end to curl gives us a
real kernel pipe and the import completes.

* fix(agent_image): use temp file + positional alias for incus 6.x

Incus 6.x rejects '-' as stdin for image import and rejects bare
HTTPS URLs (expects an incus image server). Download to a temp
file then pass its path. Also fix image list query: positional
<alias> arg (--filter=alias=... is only valid for container list).

* docs(openclaw): rename provider to built-in litellm in integration tracker

The design doc still referenced models.providers.taos (a custom provider
that was abandoned mid-implementation in favour of openclaw's built-in
litellm provider type). Updated the bootstrap example, the integration
tracker table, and the openclaw.json shape to match what actually ships.
The channels-side "provider: taos" identifier is unchanged; that's the
channel-kind name, separate from the LLM provider.
hognek referenced this pull request in hognek/tinyagentos May 30, 2026
Address kilo-code-bot review on PR jaylfc#477:

- Add _activeCloudTypes module variable, updated from /api/providers/types
  so isCloud() and groupByCategory() use canonical types instead of
  hardcoded FALLBACK_CLOUD_TYPES (WARNING #2 + line 149/227 observations)

- Add AbortController with 5s timeout to the provider types fetch
  (WARNING #1 — missing timeout)

- Suppress AbortError log noise in the catch block
jaylfc added a commit that referenced this pull request May 31, 2026
…477)

* refactor: consolidate provider-type lists to single source of truth

Provider types were duplicated across 4 files:
- tinyagentos/config.py — VALID_BACKEND_TYPES
- tinyagentos/routes/providers.py — CLOUD_TYPES
- tinyagentos/llm_proxy.py — BACKEND_TYPE_MAP, CHAT_BACKEND_TYPE_MAP,
  CLOUD_BACKEND_TYPES
- desktop/src/apps/ProvidersApp.tsx — CLOUD_TYPES, LOCAL_TYPES

Adding a type meant touching all four; #350 missed llm_proxy.py and
would have silently routed to api.openai.com instead of the user's URL.

Changes:
- New: tinyagentos/providers/__init__.py — canonical types + LiteLLM maps
- config.py, llm_proxy.py, routes/providers.py — import from providers
- New: GET /api/providers/types endpoint — returns canonical types as JSON
- ProvidersApp.tsx — fetches types from API at boot, hardcoded values
  demoted to fallback. LOCAL_TYPES now correctly includes sd-cpp/rknn-sd
  (was missing from frontend hardcoded list).

Adding a provider type now touches one place.

Closes #351

* fix(frontend): pass localTypes as prop to ProviderForm, drop unused cloudTypes state

ProviderForm is a child component — it cannot access ProvidersApp state.
localTypes is now passed as a prop. cloudTypes was unused (module-level
isCloud/groupByCategory can't access component state).

* fix: address CodeRabbit nitpicks — NEEDS_API_BASE_TYPES, drop redundant import, log fetch failures

- Add NEEDS_API_BASE_TYPES / IMAGE_GEN_TYPES to providers/__init__.py
- llm_proxy.py: use NEEDS_API_BASE_TYPES instead of hardcoded tuple
- routes/providers.py: drop redundant CLOUD_TYPES local import
- ProvidersApp.tsx: log fetch failures instead of silent catch

* fix(providers): use fetched cloud types + add fetch timeout

Address kilo-code-bot review on PR #477:

- Add _activeCloudTypes module variable, updated from /api/providers/types
  so isCloud() and groupByCategory() use canonical types instead of
  hardcoded FALLBACK_CLOUD_TYPES (WARNING #2 + line 149/227 observations)

- Add AbortController with 5s timeout to the provider types fetch
  (WARNING #1 — missing timeout)

- Suppress AbortError log noise in the catch block

* fix(providers): abort provider-types fetch + clear timeout on unmount

Return a cleanup from the boot-time useEffect so the fetch is aborted and the
timeout cleared on unmount — prevents setState-after-unmount. Addresses the
CodeRabbit review on this PR.

* fix(providers): drop unused cloudTypes binding (TS6133), keep setter as re-render trigger

---------

Co-authored-by: Hogne <hogne@tinyagentos.dev>
Co-authored-by: jaylfc <jaylfc25@gmail.com>
jaylfc added a commit that referenced this pull request Jul 16, 2026
jaylfc pushed a commit that referenced this pull request Jul 18, 2026
…890 C2) (#1903)

* feat(cluster): worker-initiated graceful pause + drain protocol (#890 C2)

Add worker-initiated status transitions via heartbeat:
- heartbeat() accepts optional status and drain_reason fields
- Workers can self-report 'update-available', 'draining', or 'updating'
- Status 'update-available' workers remain routable (still serving)
- Status 'draining' triggers same behavior as controller-initiated drain
- Monitor loop handles stale update-available workers

Controller side (ClusterManager):
- _worker_for_resource, get_workers_for_capability, aggregate_catalog
  now include 'update-available' alongside 'online'
- Worker-initiated drain emits notifications to activity feed

Worker side (WorkerAgent):
- report_update_available(), initiate_self_drain(), notify_drain_complete()
- Convenience methods wrapping heartbeat() with appropriate status

Tests: 11 new unit tests covering state transitions, routability,
catalog inclusion, lease claims, monitor loop staleness, and the
full online->update-available->draining pipeline.

State machine:
  online -> update-available -> draining -> updating -> (restart) -> online

* fix(cluster): rebase onto dev + address Kilo findings (1C/1W/2S) for worker drain

Rebased feat/worker-self-drain onto origin/dev, resolving conflicts in
manager.py, routes/cluster.py, and worker/agent.py from recently-landed
registration-drift refresh fields (#1538).

Kilo fixes:
- CRITICAL: Protect 'update-available' from status-less heartbeat reversion.
  The periodic ~15s heartbeat with no status was silently flipping
  update-available workers back to online, defeating the update signal.
  Now 'update-available' is protected alongside 'draining' — only explicit
  status transitions are honoured.
- WARNING: Handle 'updating' in _monitor_loop. A worker stuck/crashed in
  'updating' state was never re-onlined or offlined, permanently excluded
  from routing/catalog. Now they are force-offlined after HEARTBEAT_TIMEOUT.
- SUGGESTION: Validate status against allow-list and sanitize drain_reason
  in notification block to prevent operator-facing UI injection.
- SUGGESTION: notify_drain_complete() now sends status='updating' heartbeat
  so the controller transitions from 'draining' to 'updating'.

Test update: test_heartbeat_without_status_reonlines_update_available
renamed to test_heartbeat_without_status_preserves_update_available and
now asserts the corrected protection behavior.

72/72 targeted tests pass.

* fix(cluster): protect updating status from heartbeat reversion and cancel arbiter tasks on stale-update

Add 'updating' to the protected status tuple in heartbeat so a periodic
status-less heartbeat doesn't revert an updating worker to 'online', which
would re-enter it into routing/catalog/lease-claims mid-update and prevent
the monitor loop's stale-updating branch from firing.

Mirror the stale-draining branch's GPU arbiter task cancellation in the
stale-updating path so that in-flight arbiter work for force-released
leases isn't orphaned when an updating worker goes stale.

Fixes: PR #1903 Kilo review findings (WARNING manager.py:239, SUGGESTION manager.py:835)

* fix(cluster): hoist status allow-list to module frozenset + guard against invalid status heartbeat bypass

Kilo found two issues post-rebase on #1903:
1. The drain/update status allow-list was duplicated in three places
   (guard condition, assignment check, notification block). Hoist to
   module-level _VALID_STATUSES frozenset and reference in all spots.
2. An invalid status value via heartbeat ('online', 'offline', garbage)
   could revert a protected (draining/update-available/updating) worker
   back to 'online', defeating drain protection. Tighten the guard to
   require status in _VALID_STATUSES when the worker is protected.

Closes #1903.

* fix(cluster): persist lifecycle status across controller restart, cancel arbiter on stale-update-available, gate online event

- worker/agent.py: store lifecycle_status/lifecycle_reason in WorkerAgent,
  resend on periodic heartbeats so controller restart doesn't revert
  draining/updating workers to 'online' (CodeRabbit MAJOR #1)
- manager.py: cancel GPU arbiter tasks on stale-update-available worker
  timeout, matching the cancellation already done for stale-drain and
  stale-update paths (CodeRabbit MAJOR #2)
- manager.py: gate recovery 'worker.online' event on
  worker.status == 'online' to suppress false events for workers
  recovering into draining/updating state (CodeRabbit MINOR #3)
- tests/test_worker.py: accept **kwargs in fake_heartbeat mocks to
  match expanded heartbeat() signature

---------

Co-authored-by: Hogne <227774406+hognek@users.noreply.github.com>
jaylfc pushed a commit that referenced this pull request Jul 18, 2026
…s-user-collab spec (#2024)

CR #1 — Add replay protection, canonical serialization, audience binding,
freshness window, and persistent replay cache to envelope handshake (line 124).

CR #2 — Add peer-token vs. project-level authorization: contact identity
authenticates the route family, but per-operation project membership,
policy, and channel membership are independently enforced (line 131).

CR #3 — Add fail-closed runtime checks on every delegated request: verify
current sponsor-contact status and project membership at call time, not
just at token issuance; purge outbox and unclaim in-flight tasks on revoke
(line 255).

CR #4 — Require non-null, owner-validated, immutable project_id claim in
minted JWTs for project_tasks and canvas_read scopes (line 251).

CR #5 — Require atomic uniqueness constraint on (contact_id, channel_id,
remote_msg_id) in chat_messages, insert-before-ack, treat conflicts as
successful duplicates (line 328).

Kilo #1 — Clarify project.visibility as a future schema addition
(projects.visibility TEXT DEFAULT 'private'), not an existing column on
hub posts (line 524).

Co-authored-by: Hogne <227774406+hognek@users.noreply.github.com>
jaylfc added a commit that referenced this pull request Jul 18, 2026
* docs(design): cross-user collaboration spec (contacts, human collaborators, agent delegation, human DM)

Flagship epic: add another taOS user as a contact (hub friend edge), invite
them as a human project collaborator, receive their delegated agents (normal
external-selfjoin identities tagged with a sponsor contact), and chat
human-to-human over an instance-to-instance peer channel. Rides the shipped
invite/consent/registry rails; rejects registry-abuse and bus-federation.
Phased plan with a pilot-first path (hogne) and lanes for lead/hognek/fleet.
Awaiting owner decisions (section 10).

* docs(design): add Milestone F - collaborator community view (stats, leaderboard, live kanban, community chat)

When a user collaborates on a remote project, their Projects app gets a rich
read-mostly community view (a mini social-GitHub for the shared project) so
they see their contribution/involvement and feel part of the team. Introduces
a scoped read-only project projection served over the peer channel (never
direct API access, never sensitive fields) - the read counterpart to the
write-delegation in Milestone D. Refines Decision 2: rich read surface, still
zero board/file write for the human. Member-scoped in v1; public view a later
toggle on the same machinery.

* docs(design): address 1 Kilo + 5 CodeRabbit security findings on cross-user-collab spec (#2024)

CR #1 — Add replay protection, canonical serialization, audience binding,
freshness window, and persistent replay cache to envelope handshake (line 124).

CR #2 — Add peer-token vs. project-level authorization: contact identity
authenticates the route family, but per-operation project membership,
policy, and channel membership are independently enforced (line 131).

CR #3 — Add fail-closed runtime checks on every delegated request: verify
current sponsor-contact status and project membership at call time, not
just at token issuance; purge outbox and unclaim in-flight tasks on revoke
(line 255).

CR #4 — Require non-null, owner-validated, immutable project_id claim in
minted JWTs for project_tasks and canvas_read scopes (line 251).

CR #5 — Require atomic uniqueness constraint on (contact_id, channel_id,
remote_msg_id) in chat_messages, insert-before-ack, treat conflicts as
successful duplicates (line 328).

Kilo #1 — Clarify project.visibility as a future schema addition
(projects.visibility TEXT DEFAULT 'private'), not an existing column on
hub posts (line 524).

Co-authored-by: Hogne <227774406+hognek@users.noreply.github.com>

---------

Co-authored-by: hognek <hognek@gmail.com>
Co-authored-by: Hogne <227774406+hognek@users.noreply.github.com>
jaylfc pushed a commit that referenced this pull request Jul 20, 2026
…orage (#2036)

* feat(secrets): per-agent GitHub token grant UI + short-lived token storage

- Add github-installation secret category to SecretsStore
- Create tinyagentos/github_token.py: lifetime-aware token minting cache
  with mint_installation_token() and get_agent_github_token()
- Add SecretsStore.get_agent_github_installations() for per-agent queries
- Add GET /api/secrets/agent/{name}/github endpoint
- Update SecretsApp.tsx with github-installation category support
- Add per-agent access toggles for each GitHub App repo in GitHubConnect.tsx
- Tests: 7/7 github_token tests, 6/7 secrets tests pass

* fix(secrets): address 6 Kilo findings on GitHub token grants

Fix 4 WARNING + 2 SUGGESTION findings from Kilo review on PR #2036:

WARNING fixes:
- routes/secrets.py: Validate agent_name path parameter (reject empty,
  path separators, traversal) with docstring noting CSRF auth middleware
- github_token.py: Iterate all granted installations instead of only
  installations[0]; document per-repo vs per-installation token scope
- GitHubConnect.tsx + github_oauth.py + github.ts: Pass real permissions
  from GitHub App installation instead of hardcoded []
- test_secrets.py: Add TestSecretsEncryption class restoring XOR/Fernet
  round-trip coverage (encrypt/decrypt, list masking, update re-encrypt)

SUGGESTION fixes:
- github_token.py: LRU eviction by oldest expiry timestamp
- GitHubConnect.tsx: Per-repo savingGrants Set<string> state

* fix(secrets): add auth guard to GitHub grants endpoint, fix route converters for names with /

- Add Depends(get_current_user) to GET /api/secrets/agent/{agent_name}/github
  so unauthenticated callers receive 401 instead of leaking installation IDs
  and repo names (jaylfc MUST-FIX #1, Kilo WARNING)
- Change {name} to {name:path} on GET/PUT/DELETE /api/secrets/{name}
  so secrets whose names contain / (e.g. GitHub repo full names like
  owner/repo) are correctly routed
- Add test_github_grants_requires_auth (401 for unauthenticated caller)
- Fix test_update_secret_agents (uses literal / path now that {name:path}
  handles multi-segment names)
- Remove useless path-traversal guard on agent_name (Kilo: irrelevant
  since agent_name is not a filesystem path)

* fix(github): normalize permissions to list form across the stack

The GitHub API returns permissions as a dict {contents: read} but the
rest of the pipeline (secrets.py, tests) expects a list [contents:read].
Convert at the API-response boundary in github_oauth.py and align
TypeScript types (Record<string,string> → string[]).

Fixes Kilo WARNING: permissions type mismatch across the stack.

* fix(github-grants): add owner/admin auth check, guard non-dict JSON, surface save errors

- routes/secrets.py: verify caller owns agent (or is admin) before
  returning GitHub installation grants (CodeRabbit Major)
- secrets.py: guard against json.loads returning non-dict payload
  in get_agent_github_installations (CodeRabbit Minor)
- GitHubConnect.tsx: surface save-failure as transient inline error
  message instead of silently ignoring (CodeRabbit Minor)

---------

Co-authored-by: Hogne <227774406+hognek@users.noreply.github.com>
jaylfc pushed a commit that referenced this pull request Jul 26, 2026
…limiter docs (#2032)

* fix(peer): address Kilo WARNINGs — align design doc, FIXME rate limiter

- Design doc: outbound_token comment now matches code (plaintext; deferred
  to post-MVP) instead of misleading 'encrypted at rest'.
- Rate limiter: add explicit FIXME for shared-store backing across workers,
  document per-worker aggregate limit semantics.

Note: monkeypatch.setenv fixture fix was already applied in upstream merge
of #2025, so this commit carries only the two remaining warnings.

* fix(peer): address 5 Kilo findings — centralized auth, nonce replay, rate-limit LRU, prune commit, token docs

Fix #1 (WARNING): Document outbound_token plaintext threat model in
contacts_store schema comment — token must be presented on outbound
requests; at-rest encryption deferred to post-MVP.

Fix #2 (WARNING): Centralized /api/peer/ authentication via router-level
_peer_auth_dep dependency. Previously EXEMPT_PREFIXES bypassed all auth
middleware and correctness depended on every route calling
_authenticate_peer. Now every route under /api/peer/ gets bearer-auth
automatically; route handlers read contact_id from request.state.

Fix #3 (SUGGESTION): Commit the opportunistic nonce prune in its own
transaction so a replay (IntegrityError) rollback does not undo it.

Fix #4 (SUGGESTION): Add record_nonce(…, kind='ack') to /api/peer/ack
for replay protection. A replayed ack now returns 409 Conflict, matching
the /inbox and /chat contract.

Fix #5 (SUGGESTION): Add LRU fallback eviction to the rate limiter.
When all 2000+ entries have active windows (no expired entries to sweep),
the oldest entry is evicted to prevent unbounded dict growth under
sustained load.

---------

Co-authored-by: Hogne <227774406+hognek@users.noreply.github.com>
jaylfc added a commit that referenced this pull request Jul 26, 2026
* fix(desktop): dialogs render above windows, and mint shows the URL and PIN (#2092)

Two bugs on the same screen, both reported from live use.

Window z-index was an unbounded counter: every open, focus, restore and
recenter incremented nextZIndex forever, while portal overlays sit at a fixed
z-[10001]. After enough focus switches in one session, windows rendered on top
of modal dialogs. The stack is now renumbered to 1..N on each change, so window
z stays far below the overlay layer regardless of session length, and relative
order is preserved by sorting on the existing values first. This affected every
portal overlay, not only the invite dialog.

ProjectMembers closed the invite dialog in its onMinted handler, unmounting it
before the result rendered. The invite URL and PIN are shown exactly once and
cannot be recovered, so a successful mint looked like a silent failure. The
parent now refreshes its member list only and the user closes the dialog once
they have copied the credentials.

* feat(agents): project_tasks_create scope so an external agent can author cards (#2098)

An agent holding project_tasks could claim, close and comment on existing
cards but never open one, so an approved grant bought nothing on that route.
Rather than widen project_tasks, which is documented and tested as read plus
lifecycle plus comments (Invariant 2 + 5) and would retroactively grant
authoring to every agent already approved for it, authoring gets its own
narrower scope that an owner opts into per agent.

create_task now authorises through the same _authorize_task_actor as the other
task routes, parameterised on scope, so existence-hiding 404s behave
identically. The middleware allowlist admits POST .../tasks, which lets the
token reach a handler that then verifies JWT, project binding and the narrower
scope; project_tasks alone is still refused.

Tests keep the original invariant (project_tasks alone cannot create) and add
the halves that make it meaningful: the new scope DOES allow authoring, it is
project-bound so a grant on A cannot create on B, and it does not widen member
management. Authorship is attributed to the agent, not the project owner.

* feat(library): item card component with thumbnail, status, artifacts, collection link (#2097)

Add LibraryItemCard component per docs/design/library-app.md sections 2-4.
Card shows thumbnail (or placeholder), title, kind badge, media duration,
pipeline status per stage (jobs shape), artifact list (text, transcript,
description, ocr) with preview, link-to-collection action, and a disabled
Download stub until P3. Failure states are always visible -- no silent
empties for missing thumbnails, pipeline stages, artifacts, or errors.

Includes lib/library.ts with types and API client for the library store
(items, artifacts, jobs) and 25 component tests covering pending,
processing, ready, and error states.

* feat(agents): enforce files_read/files_write so member agents can access project Files (#2100)

* feat(agents): enforce files_read/files_write scopes so member agents can access project Files

Project-files routes (/api/projects/{slug}/files*, mkdir, trash, stats) had no
membership or scope gate and were absent from the agent middleware allowlist, so
an agent token could not reach them at all while the files_read/files_write
scopes existed but were never enforced. This wires them up, mirroring the canvas
pattern:

- _authorize_files_actor resolves slug -> project and authorizes a session
  owner/admin (unchanged) OR an agent holding files_read (reads) / files_write
  (writes) grant bound to that project. A token bound to another project, or an
  unknown slug, collapses into an existence-hiding 404; a missing scope is 403.
- _AGENT_FILES_ROUTES added to auth_middleware so agent JWTs pass through to
  the routes, which verify the grant.
- InviteAgentDialog offers files_read (default on) and files_write, so an owner
  can grant file access at invite time.

Agents that are project members can now read the project's Files and add files
via the API. Grant creation already grants these scopes generically, and
membership is added via the always-on project_tasks scope.

Adds tests/test_routes_project_files_agent_scope.py (10 cases covering read/write
allow, missing-scope 403, cross-project 404, unknown-slug 404, session owner
unchanged).

* feat(agents): surface Files in the invite bundle + fix agent API-surface docs

- build_connection_bundle now advertises the project Files endpoints and adds a
  Files capability section to the join guide when files_read/files_write are
  granted, plus task_create when project_tasks_create is granted, so a joining
  agent is told the Files API exists and how to reach it (slug-keyed paths).
- docs/agent-coordination.md: drop the non-existent project_doc_review scope,
  add files_read/files_write, project_tasks_create, and decisions_write to the
  agent API-surface list, and correct the doc-gate note (it fires only on file
  add/delete, so it does not catch allowlist edits).
- README: replace the understated read-only agent-surface sentence with the
  real scoped surface (tasks, canvas, files, decisions, a2a).

Verified against VALID_SCOPES / _ALLOWED_SCOPES and the auth_middleware
allowlist. Invite tests pass (36).

* feat(providers): add Nous Portal as a cloud model provider (#2102)

Nous Portal (Nous Research) is an OpenAI-compatible inference API serving the
Hermes 4 family and frontier models. Wire it up as a first-class cloud provider
so it can be added from the Providers app instead of a hand-configured
openai-compatible endpoint.

- providers/__init__.py: add 'nous' to ALL_TYPES + CLOUD_TYPES, and map it to
  the OpenAI LiteLLM prefix (api_base set explicitly, like kilocode).
- routes/providers.py: default base URL https://inference-api.nousresearch.com/v1
  and a seed model list (flagship Hermes models) for the case where /v1/models
  cannot be listed without a working credential.
- backend_adapters.py: 'nous' uses the CloudAPIAdapter probe.
- Frontend: add 'nous' to the cloud provider type lists and the Providers app
  metadata (label 'Nous Portal', default URL, description, key placeholder).

Base URL and OpenAI-compatibility verified against Nous Portal docs. Backend
provider suite passes (68); frontend tsc clean.

* feat(desktop): Assistant Studio - a workspace for a personal-assistant agent (#2103)

A new studio app where the user picks a registered agent to be their PA and
works out of one hub. Left rail: Overview, Journal, Calendar/time, Tasks, Comms,
Canvas, and a Deliverables (files/reports) area. The PA picker defaults to
Hermes when present and persists the choice. Journal, Tasks, Calendar events and
Deliverables persist locally per PA so switching PA swaps the whole workspace;
Comms opens the live agent chat and Canvas points at the project canvas.

MVP scope: self-contained, no new backend (localStorage-backed), so it is
additive and safe. Accessible (labels, aria-current, keyboard add). Registered
as an optional studio app. Backend wiring (real calendar, PA-scoped board/files)
is a follow-up.

tsc clean; frontend build passes.

* feat(agents): request additional scopes for an existing agent identity (#1921)

* feat(agents): request additional scopes for an existing agent identity

Add a scope-request flow so an already-registered agent can gain more
scope grants on its SAME canonical_id, instead of the auth-request flow
which mints a new identity on approval (and 409s on an active-handle
collision).

Endpoints (in routes/agent_auth_requests.py):
- POST /api/agents/registry/{cid}/scope-requests            (create)
- POST /api/agents/registry/{cid}/scope-requests/{id}/approve
- POST /api/agents/registry/{cid}/scope-requests/{id}/deny

Auth (security-critical): creation is gated to the agent's OWN registry
bearer token (sub == canonical_id) OR the owning user / an admin, because
the agent already holds credentials; an anonymous caller can never
escalate an existing identity. The middleware allowlist exposes only the
create path to a registry JWT; approve/deny are owner/admin only.

Approval writes add_grant(cid, scope, project_id) per granted scope
(idempotent via the UNIQUE key), never registers a second identity, and
lets the admin narrow but not widen the requested scopes.
decisions_read/decisions_write are grantable globally or per-project;
project_tasks and canvas scopes still require an explicit project_id.

Adds AgentScopeRequestsStore, a scope-agnostic check_agent_identity
helper, and full route tests. VALID_SCOPES stays in sync with
_ALLOWED_SCOPES.

Fixes #1920

* fix(agents): fold scope-request approval security findings (#1921)

Addresses the Kilo + CodeRabbit findings on the approve/create scope-request paths:

- Major (project binding): approve_scope_request bound global-capable scopes
  (decisions_*) to effective_project, which fell back to the agent-named
  req.project_id when the operator gave no explicit project_id. Since an agent
  can self-request, that let a global scope bind to any project the operator
  never validated (cross-project escalation). Now grants bind ONLY to the
  operator's explicit body.project_id (None = global); the agent-named value is
  never a binding.
- Atomicity + races: approve_scope_request wrote grants + membership before the
  set_decision flip with no lock, so concurrent approvals could double-grant.
  Wrapped the whole approval in the same per-request _get_approve_lock the
  consent path uses, with a pending re-check inside the lock (grant-before-flip
  is safe under the lock + idempotent add_grant).
- Info leak: create_scope_request now authorizes BEFORE scope-vocabulary
  validation, so an unauthorized caller cannot probe whether a scope name is
  valid.

Adds tests: global scope ignores an agent-supplied project_id (binds global),
and create authorizes before vocab (403 not a 400 vocab leak). 16 tests pass.

* fix(agents): take the per-request lock in deny_scope_request too (#1921)

Kilo review: approve_scope_request now serializes concurrent approvals under
_get_approve_lock; deny lacked the same lock, so a concurrent approve+deny of
the same request was unserialized. Wrap the deny body in the same per-req lock
with a pending re-check, matching approve and the sibling consent path.

* fix(apps): register assistant-studio as an installable optional app (#2104)

Assistant Studio (#2103) shipped in the frontend registry as optional but was
not in the server-side optional-app catalog, so getLaunchableApps hid it (it is
only shown once installed) and POST /api/apps/optional/assistant-studio/install
returned 'not an optional app'. Add it to OPTIONAL_FRONTEND_APPS + the version /
trust / provenance dicts, matching the other Creative Studios.

* fix(shortcuts): resolve the container's real incus project (and start it) for terminal shortcuts (#2105)

The agent terminal/TUI shortcuts opened an incus PTY with no --project flag, so
incus used the client's default project (a per-user one like user-999). An agent
whose container lives in a different project (e.g. a legacy container in
'default') failed with 'Failed to fetch instance taos-agent-<name> in project
user-999: Instance not found', even though the container exists. Every other
container op already resolves the real project via _resolve_container_project
(--all-projects); the PTY path was the lone exception.

_open_incus_pty now resolves the container's actual project and passes
--project, and starts the container if it is stopped (incus exec fails on a
non-running instance), so a dev-access shortcut works regardless of project or
run state. Adds a sync _resolve_project_and_state_sync sibling to the async
resolver.

Tests: exec targets the resolved project + stopped container is started; a
running container is not restarted.

* test(scope-requests): add grant-enforcement E2E, tighten approve tolerance, add deny negative test (#2108)

Three test gaps filled following PR #1921 merge:

1. E2E grant-enforcement test (test_approved_scope_grant_unlocks_route_e2e):
   agent self-requests decisions_write → admin approves → agent uses token
   on POST /api/decisions → assert 200. Proves the full grant-enforcement
   chain is wired end to end. Regression guard for issue #2095.

2. Tighten test_agent_cannot_approve_its_own_request:
   replace (401, 403) tolerance with exact 401. The middleware does not pass
   a registry JWT through to the approve handler (only the create path is
   allowlisted), so it falls to the session gate → 401 exactly.

3. Deny negative test (test_agent_cannot_deny_its_own_request):
   same auth model as approve — middleware does not allowlist the deny
   endpoint for registry JWTs, so an agent token on the deny endpoint
   falls through to the session gate → 401.

Co-authored-by: Hogne <227774406+hognek@users.noreply.github.com>

* chore(deps): bump actions/setup-python from 6 to 7 (#2109)

Bumps [actions/setup-python](https://github.com/actions/setup-python) from 6 to 7.
- [Release notes](https://github.com/actions/setup-python/releases)
- [Commits](actions/setup-python@v6...v7)

---
updated-dependencies:
- dependency-name: actions/setup-python
  dependency-version: '7'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

* chore(deps): bump the python-deps group with 2 updates (#2110)

Updates the requirements on [matrix-nio](https://github.com/matrix-nio/matrix-nio) and [litellm[proxy]](https://github.com/BerriAI/litellm) to permit the latest version.

Updates `matrix-nio` from 0.25.2 to 0.26.0
- [Changelog](https://github.com/matrix-nio/matrix-nio/blob/main/CHANGELOG.md)
- [Commits](matrix-nio/matrix-nio@0.25.2...0.26.0)

Updates `litellm[proxy]` to 1.93.0
- [Release notes](https://github.com/BerriAI/litellm/releases)
- [Commits](BerriAI/litellm@v1.92.0...v1.93.0)

---
updated-dependencies:
- dependency-name: matrix-nio
  dependency-version: 0.26.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: python-deps
- dependency-name: litellm[proxy]
  dependency-version: 1.93.0
  dependency-type: direct:production
  dependency-group: python-deps
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

* chore(release): v1.0.0-beta.44 version bump + changelog (#2121)

* chore(release): v1.0.0-beta.44 version bump + changelog

* docs(changelog): complete the beta.44 entry (dialog/mint fix, task-create scope)

* test(desktop): add unit tests for AssistantStudioApp (#2116)

* test(desktop): add unit tests for AssistantStudioApp

* test(desktop): await the mount-time agents fetch in the first two Assistant Studio tests

Qodo review: the first two tests rendered the component without awaiting its
mount-time /api/agents fetch, so the resulting setState could land outside
React Testing Library's act() and flake in stricter environments. The later
tests in the same file already waited, so this was an inconsistency as much as
a latent flake. Both now drain the fetch before asserting.

* fix(library): wire up the unused source ingest option (#2117)

* fix(library): wire up the unused source ingest option

* fix(library): remove the unused source ingest option instead of sending it

Review found that serializing source only moved the silent drop server-side:
/api/library/ingest accepts file, url and title only, and LibraryStore has no
source column (its source_url is already derived from the url). So the field was
discarded either way, while now looking wired.

The card offered wire-it-in or remove-it and I recommended wire-it-in without
checking the backend, which was wrong. Removing it is the honest fix: no caller
passes source, so nothing breaks, and a dead option no longer implies a feature
that does not exist. Capturing real source metadata is a backend feature, not a
client nit.

* chore(deps): bump the spa-deps group in /desktop with 17 updates (#2111)

Bumps the spa-deps group in /desktop with 17 updates:

| Package | From | To |
| --- | --- | --- |
| [@radix-ui/react-dialog](https://github.com/radix-ui/primitives/tree/HEAD/packages/react/dialog) | `1.1.19` | `1.1.23` |
| [@radix-ui/react-dropdown-menu](https://github.com/radix-ui/primitives/tree/HEAD/packages/react/dropdown-menu) | `2.1.20` | `2.1.24` |
| [@radix-ui/react-label](https://github.com/radix-ui/primitives/tree/HEAD/packages/react/label) | `2.1.11` | `2.1.15` |
| [@radix-ui/react-select](https://github.com/radix-ui/primitives/tree/HEAD/packages/react/select) | `2.3.3` | `2.3.7` |
| [@radix-ui/react-slot](https://github.com/radix-ui/primitives/tree/HEAD/packages/react/slot) | `1.3.0` | `1.3.3` |
| [@radix-ui/react-switch](https://github.com/radix-ui/primitives/tree/HEAD/packages/react/switch) | `1.3.3` | `1.3.7` |
| [@radix-ui/react-tabs](https://github.com/radix-ui/primitives/tree/HEAD/packages/react/tabs) | `1.1.17` | `1.1.21` |
| [@radix-ui/react-tooltip](https://github.com/radix-ui/primitives/tree/HEAD/packages/react/tooltip) | `1.2.12` | `1.2.16` |
| [@tiptap/extension-link](https://github.com/ueberdosis/tiptap/tree/HEAD/packages/extension-link) | `3.28.0` | `3.29.0` |
| [@tiptap/extension-underline](https://github.com/ueberdosis/tiptap/tree/HEAD/packages/extension-underline) | `3.28.0` | `3.29.0` |
| [@tiptap/pm](https://github.com/ueberdosis/tiptap/tree/HEAD/packages/pm) | `3.28.0` | `3.29.0` |
| [@tiptap/react](https://github.com/ueberdosis/tiptap/tree/HEAD/packages/react) | `3.28.0` | `3.29.0` |
| [@tiptap/starter-kit](https://github.com/ueberdosis/tiptap/tree/HEAD/packages/starter-kit) | `3.28.0` | `3.29.0` |
| [react](https://github.com/react/react/tree/HEAD/packages/react) | `19.2.7` | `19.2.8` |
| [react-dom](https://github.com/react/react/tree/HEAD/packages/react-dom) | `19.2.7` | `19.2.8` |
| [@playwright/test](https://github.com/microsoft/playwright) | `1.61.1` | `1.62.0` |
| [@vitejs/plugin-react](https://github.com/vitejs/vite-plugin-react/tree/HEAD/packages/plugin-react) | `6.0.3` | `6.0.4` |


Updates `@radix-ui/react-dialog` from 1.1.19 to 1.1.23
- [Changelog](https://github.com/radix-ui/primitives/blob/main/packages/react/dialog/CHANGELOG.md)
- [Commits](https://github.com/radix-ui/primitives/commits/HEAD/packages/react/dialog)

Updates `@radix-ui/react-dropdown-menu` from 2.1.20 to 2.1.24
- [Changelog](https://github.com/radix-ui/primitives/blob/main/packages/react/dropdown-menu/CHANGELOG.md)
- [Commits](https://github.com/radix-ui/primitives/commits/HEAD/packages/react/dropdown-menu)

Updates `@radix-ui/react-label` from 2.1.11 to 2.1.15
- [Changelog](https://github.com/radix-ui/primitives/blob/main/packages/react/label/CHANGELOG.md)
- [Commits](https://github.com/radix-ui/primitives/commits/HEAD/packages/react/label)

Updates `@radix-ui/react-select` from 2.3.3 to 2.3.7
- [Changelog](https://github.com/radix-ui/primitives/blob/main/packages/react/select/CHANGELOG.md)
- [Commits](https://github.com/radix-ui/primitives/commits/HEAD/packages/react/select)

Updates `@radix-ui/react-slot` from 1.3.0 to 1.3.3
- [Changelog](https://github.com/radix-ui/primitives/blob/main/packages/react/slot/CHANGELOG.md)
- [Commits](https://github.com/radix-ui/primitives/commits/HEAD/packages/react/slot)

Updates `@radix-ui/react-switch` from 1.3.3 to 1.3.7
- [Changelog](https://github.com/radix-ui/primitives/blob/main/packages/react/switch/CHANGELOG.md)
- [Commits](https://github.com/radix-ui/primitives/commits/HEAD/packages/react/switch)

Updates `@radix-ui/react-tabs` from 1.1.17 to 1.1.21
- [Changelog](https://github.com/radix-ui/primitives/blob/main/packages/react/tabs/CHANGELOG.md)
- [Commits](https://github.com/radix-ui/primitives/commits/HEAD/packages/react/tabs)

Updates `@radix-ui/react-tooltip` from 1.2.12 to 1.2.16
- [Changelog](https://github.com/radix-ui/primitives/blob/main/packages/react/tooltip/CHANGELOG.md)
- [Commits](https://github.com/radix-ui/primitives/commits/HEAD/packages/react/tooltip)

Updates `@tiptap/extension-link` from 3.28.0 to 3.29.0
- [Release notes](https://github.com/ueberdosis/tiptap/releases)
- [Changelog](https://github.com/ueberdosis/tiptap/blob/main/packages/extension-link/CHANGELOG.md)
- [Commits](https://github.com/ueberdosis/tiptap/commits/v3.29.0/packages/extension-link)

Updates `@tiptap/extension-underline` from 3.28.0 to 3.29.0
- [Release notes](https://github.com/ueberdosis/tiptap/releases)
- [Changelog](https://github.com/ueberdosis/tiptap/blob/main/packages/extension-underline/CHANGELOG.md)
- [Commits](https://github.com/ueberdosis/tiptap/commits/v3.29.0/packages/extension-underline)

Updates `@tiptap/pm` from 3.28.0 to 3.29.0
- [Release notes](https://github.com/ueberdosis/tiptap/releases)
- [Changelog](https://github.com/ueberdosis/tiptap/blob/main/packages/pm/CHANGELOG.md)
- [Commits](https://github.com/ueberdosis/tiptap/commits/v3.29.0/packages/pm)

Updates `@tiptap/react` from 3.28.0 to 3.29.0
- [Release notes](https://github.com/ueberdosis/tiptap/releases)
- [Changelog](https://github.com/ueberdosis/tiptap/blob/main/packages/react/CHANGELOG.md)
- [Commits](https://github.com/ueberdosis/tiptap/commits/v3.29.0/packages/react)

Updates `@tiptap/starter-kit` from 3.28.0 to 3.29.0
- [Release notes](https://github.com/ueberdosis/tiptap/releases)
- [Changelog](https://github.com/ueberdosis/tiptap/blob/main/packages/starter-kit/CHANGELOG.md)
- [Commits](https://github.com/ueberdosis/tiptap/commits/v3.29.0/packages/starter-kit)

Updates `react` from 19.2.7 to 19.2.8
- [Release notes](https://github.com/react/react/releases)
- [Changelog](https://github.com/react/react/blob/main/CHANGELOG.md)
- [Commits](https://github.com/react/react/commits/v19.2.8/packages/react)

Updates `react-dom` from 19.2.7 to 19.2.8
- [Release notes](https://github.com/react/react/releases)
- [Changelog](https://github.com/react/react/blob/main/CHANGELOG.md)
- [Commits](https://github.com/react/react/commits/v19.2.8/packages/react-dom)

Updates `@playwright/test` from 1.61.1 to 1.62.0
- [Release notes](https://github.com/microsoft/playwright/releases)
- [Commits](microsoft/playwright@v1.61.1...v1.62.0)

Updates `@vitejs/plugin-react` from 6.0.3 to 6.0.4
- [Release notes](https://github.com/vitejs/vite-plugin-react/releases)
- [Changelog](https://github.com/vitejs/vite-plugin-react/blob/main/packages/plugin-react/CHANGELOG.md)
- [Commits](https://github.com/vitejs/vite-plugin-react/commits/plugin-react@6.0.4/packages/plugin-react)

---
updated-dependencies:
- dependency-name: "@radix-ui/react-dialog"
  dependency-version: 1.1.23
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: spa-deps
- dependency-name: "@radix-ui/react-dropdown-menu"
  dependency-version: 2.1.24
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: spa-deps
- dependency-name: "@radix-ui/react-label"
  dependency-version: 2.1.15
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: spa-deps
- dependency-name: "@radix-ui/react-select"
  dependency-version: 2.3.7
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: spa-deps
- dependency-name: "@radix-ui/react-slot"
  dependency-version: 1.3.3
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: spa-deps
- dependency-name: "@radix-ui/react-switch"
  dependency-version: 1.3.7
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: spa-deps
- dependency-name: "@radix-ui/react-tabs"
  dependency-version: 1.1.21
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: spa-deps
- dependency-name: "@radix-ui/react-tooltip"
  dependency-version: 1.2.16
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: spa-deps
- dependency-name: "@tiptap/extension-link"
  dependency-version: 3.29.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: spa-deps
- dependency-name: "@tiptap/extension-underline"
  dependency-version: 3.29.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: spa-deps
- dependency-name: "@tiptap/pm"
  dependency-version: 3.29.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: spa-deps
- dependency-name: "@tiptap/react"
  dependency-version: 3.29.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: spa-deps
- dependency-name: "@tiptap/starter-kit"
  dependency-version: 3.29.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: spa-deps
- dependency-name: react
  dependency-version: 19.2.8
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: spa-deps
- dependency-name: react-dom
  dependency-version: 19.2.8
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: spa-deps
- dependency-name: "@playwright/test"
  dependency-version: 1.62.0
  dependency-type: direct:development
  update-type: version-update:semver-minor
  dependency-group: spa-deps
- dependency-name: "@vitejs/plugin-react"
  dependency-version: 6.0.4
  dependency-type: direct:development
  update-type: version-update:semver-patch
  dependency-group: spa-deps
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

* fix(agents): bind project_tasks_create and files scopes to a project on approval (#2127)

* fix(agents): bind project_tasks_create and files scopes to a project on approval

Review of the beta.44 promotion found that _SCOPE_PROJECT_SCOPES listed only
project_tasks and the canvas scopes, so project_tasks_create, files_read and
files_write could be approved with no project_id. The grant was then written
global (project_id=None), and check_agent_scope_for_project only matches a grant
bound to the project, so the operator believed they had granted access while the
agent silently had none. Fails closed, but silently wrong is its own bug.

Also replaces the em dashes in docs/agent-coordination.md with commas and colons
per the house style, and records in project_files.py why the session path
deliberately allows an unknown slug (lazily-created, slug-addressed files tree,
documented by test_list_unknown_slug_returns_empty) while the agent path stays
strict.

42 tests pass (project files, files agent scope, scope requests).

* fix(projects): setting an agent as project lead now sets lead_member_id (#2113)

add_agent_to_project wrote role='lead' on the member row but never called
set_lead, which is the only writer of projects.lead_member_id. The member row
is just a label; the pointer column is the actual lead. So the agent read as
lead in the UI while every lead-gated check refused them. This is what happened
to Hermes on taOSrabbit: role='lead', is_lead=0, lead_member_id NULL.

Best-effort like the rest of that block, since the membership and grant already
stand on their own.

* fix(agents): one definition of which scopes require a project binding

Qodo caught that the auth-request approval path still granted files_* and
project_tasks_create globally. The project-scope set existed as three parallel
copies (two function-local, one module-level), and the earlier fix only
corrected the module-level one -- leaving the path an invite actually takes
still writing those grants with project_id=None.

check_agent_scope_for_project only matches a grant bound to that exact project,
so a global grant never matches. The approval returns 200, the operator
believes access was granted, and the agent silently has none.

Now defined once and referenced everywhere; _SCOPE_* are plain aliases rather
than rebuilt literals, since re-listing the members is how the copies drifted.

Adds the regression test that was missing: nothing pinned this set, which is
why three copies could disagree unnoticed. Checks alias identity (not equality)
so a re-introduced copy fails even while it still happens to agree, asserts a
single assignment per name in the module source, and asserts every project-bound
scope is in VALID_SCOPES -- a typo there fails open, granting globally.

* ci: raise the test timeout above the actual suite runtime (#2134)

The cap was 45 min with a comment claiming 3.12/3.13 finish in ~16. Measured
over the last 12 job records that is no longer true: 3.13 takes 31-41 min and
3.12 takes 38-41, so the cap sat roughly 4 minutes above the slowest normal run.

Two of those 12 were killed mid-suite with nothing actually wrong, including
the one gating #2127. A cap that close to the median does not catch hangs, it
manufactures red PRs, and a timeout kill is indistinguishable from a real
failure until you check the clock against timeout-minutes. That is the worst
property a merge gate can have.

75 keeps a bound on a genuinely hung job while leaving real headroom. The suite
growing from ~16 to ~40 min is its own problem and is filed separately; this
stops it corrupting merge decisions in the meantime.

* feat(feedback): add per-user 24h submission cap (#2131)

* feat(feedback): add per-user 24h submission cap

* fix(feedback): make the 24h cap atomic (Qodo)

The cap did a SELECT COUNT then an INSERT as two awaited calls. Every
sequential test passes and the limit still does not hold: concurrent requests
all read the same count, all see room, and all insert. Verified against the
pre-fix code, 8 racers with one slot left all got 201 and took a cap of 20 to
27. A user with a few tabs open trips this without trying, and a scripted
client defeats it outright.

Moved enforcement into the store as a single INSERT ... SELECT ... WHERE
(SELECT COUNT(*) ...) < ?, so SQLite's write lock makes the check and the
insert indivisible, and rowcount says which way it went.

Adds the concurrency test that was missing, plus one asserting a rejected
submission leaves no row behind: a partially-written reject would tighten the
cap on every retry.

* refactor(feedback): drop count_recent, orphaned by the atomic cap

This PR added count_recent for the route to call before inserting. Moving the
cap into create_within_cap left it with no callers anywhere in the tree, so it
is dead on arrival rather than pre-existing code worth keeping.

Removing it also removes the tempting wrong path: a future caller reaching for
count_recent would reintroduce exactly the check-then-insert race the atomic
statement exists to close.

* test: replace always-true assert with issubclass check in test_installer_class_available (#1979)

* fix(install): route LXC port allocation through centralized allocator

Replace the standalone _find_free_port() in lxc_installer.py with
allocate_host_port(app_id) from port_allocator.py so the centralized
allocator is the single source of truth for all app host-port
assignments.

- Remove _find_free_port(), socket, and closing imports from lxc_installer
- Import allocate_host_port instead of RESERVED_PORTS
- Accumulate failed ports in exclude set across TOCTOU retry loop
- Fix stale _docker_published_port docstring (DockerInstaller now
  maps {allocated_host_port}:{container_port}, not {p}:{p})
- Add test class verifying _find_free_port is gone and
  allocate_host_port is the only import

Refs: #695

* test: strengthen allocator import assertion per Kilo suggestion

Add assert not hasattr(mod, 'RESERVED_PORTS') to verify the old import
is truly removed, not just that allocate_host_port is present.

* test: replace always-true assert with issubclass check in test_installer_class_available

---------

Co-authored-by: Hogne <227774406+hognek@users.noreply.github.com>

* feat(wallpaper): add Wallhaven proxy route + sectioned picker integration (#1902)

* feat(wallpaper): add Wallhaven proxy route + browse-online picker section

- Add wallhaven_api_key config field (env-only, never in repo)
- Create GET /api/wallhaven/search proxy route to wallhaven.cc API
- Keyless by default; optional X-API-Key header when WALLHAVEN_API_KEY set
- Handle rate limiting (429), timeouts (504), and Wallhaven errors (502)
- New WallhavenBrowser component: debounced search, thumbnail grid, pagination
- Integrate WallhavenBrowser into WallpaperPicker as collapsible section
- WallpaperPicker: "Browse online" toggle expands search UI, selecting a
  Wallhaven image applies it as a remote wallpaper
- Backend tests (test_wallhaven.py, 12 tests) with respx mocking
- Frontend tests (WallhavenBrowser.test.tsx, 8 tests; WallpaperPicker 14 tests)

Fixes #864

* fix(wallpaper): escape CSS url, persist wallpaperIdByTheme, guard JSONResponse, validate categories/purity, move os import

- Escape single-quotes and backslashes in remote wallpaper URLs to prevent
  CSS injection and render breakage (WallpaperPicker.tsx onSelect).
- Route through a state updater that sets wallpaperIdByTheme, and include
  light/mobile/fallback variants so theme switch preserves remote wallpapers.
- Use the received label as wallpaperOverlayText instead of discarding it.
- Guard resp.json() with try/except ValueError and cap response at 1 MB
  (wallhaven.py:71).
- Validate categories/purity with ^[01]{3}$ before forwarding to Wallhaven.
- Move import os as _os from inside load_config to module top (config.py).

* fix(wallpaper): harden CSS url() escaping — escape parens, strip control chars, validate http(s) scheme

Kilo WARNING: CSS url() escaping was incomplete. Only backslashes and
single-quotes were escaped, allowing ')' in query strings to prematurely
terminate url('...') and inject CSS. Now also:

- Escape '(' as %28 and ')' as %29
- Strip control characters (\x00-\x1f, \x7f)
- Validate scheme is http(s) before proceeding

Also removes a duplicate 'Browse online' section left behind during
the sectioned-picker rebase — changes are low-risk and build/vitest pass.

---------

Co-authored-by: Hogne <227774406+hognek@users.noreply.github.com>

* fix(gpu-arbiter): clamp drain_tick_seconds floor + remove double _signal_capacity wake (#1987)

* feat(gpu-arbiter): event-driven admission wakeup replacing 2s poll tick (#1864 A3)

Replace the hardcoded asyncio.sleep(2) in _process_queue with an
asyncio.Event-based wake path.  The drain loop now blocks on
asyncio.wait_for(self._wake.wait(), timeout=drain_tick_seconds) so a
reservation release or task completion immediately wakes the drain
loop.  The 2 s constant becomes the configurable fallback timeout
(drain_tick_seconds, default 2.0).

- __init__: new drain_tick_seconds param, self._wake Event
- _signal_capacity(): new method — calls self._wake.set()
- _process_queue: event-driven wait + clear, fallback timeout
- _release_reservation: calls _signal_capacity on actual release
- _run_gpu_task finally: calls _signal_capacity on completion
- tests/test_gpu_arbiter_wakeup.py: 2 new tests
  * test_release_triggers_immediate_drain (60 s tick, < 0.5 s admit)
  * test_poll_tick_still_drains_without_signal (0.1 s tick)

* fix(gpu-arbiter): clamp drain_tick_seconds floor + remove double _signal_capacity wake

Kilo review on #1986:
1. Clamp drain_tick_seconds to >= 0.01 in __init__ to prevent
   asyncio.wait_for(timeout=0) ValueError crash in _process_queue.
2. Remove redundant _signal_capacity() in _run_gpu_task finally —
   _release_reservation already calls it at line 250, making the
   line-448 call a double-wake on every task completion.

32/32 GPU arbiter tests pass.

---------

Co-authored-by: Hogne <227774406+hognek@users.noreply.github.com>

* fix(peer): address Kilo WARNINGs — os.environ leak, design doc, rate limiter docs (#2032)

* fix(peer): address Kilo WARNINGs — align design doc, FIXME rate limiter

- Design doc: outbound_token comment now matches code (plaintext; deferred
  to post-MVP) instead of misleading 'encrypted at rest'.
- Rate limiter: add explicit FIXME for shared-store backing across workers,
  document per-worker aggregate limit semantics.

Note: monkeypatch.setenv fixture fix was already applied in upstream merge
of #2025, so this commit carries only the two remaining warnings.

* fix(peer): address 5 Kilo findings — centralized auth, nonce replay, rate-limit LRU, prune commit, token docs

Fix #1 (WARNING): Document outbound_token plaintext threat model in
contacts_store schema comment — token must be presented on outbound
requests; at-rest encryption deferred to post-MVP.

Fix #2 (WARNING): Centralized /api/peer/ authentication via router-level
_peer_auth_dep dependency. Previously EXEMPT_PREFIXES bypassed all auth
middleware and correctness depended on every route calling
_authenticate_peer. Now every route under /api/peer/ gets bearer-auth
automatically; route handlers read contact_id from request.state.

Fix #3 (SUGGESTION): Commit the opportunistic nonce prune in its own
transaction so a replay (IntegrityError) rollback does not undo it.

Fix #4 (SUGGESTION): Add record_nonce(…, kind='ack') to /api/peer/ack
for replay protection. A replayed ack now returns 409 Conflict, matching
the /inbox and /chat contract.

Fix #5 (SUGGESTION): Add LRU fallback eviction to the rate limiter.
When all 2000+ entries have active windows (no expired entries to sweep),
the oldest entry is evicted to prevent unbounded dict growth under
sustained load.

---------

Co-authored-by: Hogne <227774406+hognek@users.noreply.github.com>

* fix(notifications): show approve/deny for scope-request notifications in bell (#2107)

* fix(notifications): show approve/deny buttons for agent_scope_requests in bell and toast

PR #1921 (agent scope-requests) emits notifications with source
'agent_scope_requests' but NotificationCentre and NotificationToast
only rendered ConsentActions for source === 'auth_requests'. Scope-
request approve/deny was invisible in the UI.

- Branch source check in NotificationCentre and NotificationToast to
  also accept 'agent_scope_requests'
- ConsentActions now accepts optional source + canonicalId props and
  routes approve/deny to the scope-request endpoints when source is
  'agent_scope_requests' (/api/agents/registry/{id}/scope-requests/...)
- consentPayload() extracts canonical_id from notification data for
  the scope-request endpoint path
- Backward compatible: source defaults to 'auth_requests' for existing
  callers (DecisionsApp, tests)

Ref: #1921

* fix: guard missing canonicalId for scope-request consent actions

Per CodeRabbit review: the canonicalId ?? requestId fallback could produce
malformed URLs like /registry/{requestId}/scope-requests/... when
canonical_id is missing from the notification data. Now explicitly
rejects scope-request approve/deny with a clear error when the
canonicalId is absent.

---------

Co-authored-by: Hogne <227774406+hognek@users.noreply.github.com>

* fix(csrf): merge CSRF token into effectiveInit for Request inputs (bot-fix for #1991) (#1999)

* fix(csrf): merge CSRF token into effectiveInit for Request inputs, not rebuild

- Compute effective method from init?.method || input.method
- Merge token into effectiveInit headers instead of rebuilding Request
  (avoids body stream consumption and respects init-provided headers/method)
- Update test to inspect init.headers instead of Request.headers

Addresses Kilo finding (Request inputs excluded from CSRF) and
CodeRabbit finding (broken token injection when init overrides
method/headers + body stream consumption).

Refs: PR #1991

* fix(csrf): merge Request headers with init headers instead of choosing one

Kilo's WARNING on this PR was correct. `init?.headers || input.headers` picks
one source, so whenever a caller supplied headers via BOTH the Request and the
init, the Request's own headers were discarded entirely:

    fetch(new Request(url, {headers: {Authorization}}), {method, headers})

lost the Authorization header. init also wins at the fetch layer for a Request
plus init, so nothing downstream restored it.

Now both are merged, init winning on conflict, which matches how fetch itself
resolves the two.

The regression test is verified to fail on the pre-fix code, and it exposes the
bug as slightly worse than reported: with the old line, that call lost the CSRF
token too, so the header this PR exists to attach was itself dropped in exactly
the case it was being extended to cover.

Test assertions live inside the existing it() block on purpose: installAuthGuard
has a module-level `installed` flag, so a second install in a fresh it() is a
no-op and window.fetch would be the bare spy rather than the wrapper. I lost
time to that before spotting it, hence the note.

---------

Co-authored-by: Hogne <227774406+hognek@users.noreply.github.com>
Co-authored-by: jaylfc <jaylfc25@gmail.com>

* fix(canvas): make the .tldr export a file tldraw can actually open (#2133)

* fix(canvas): make the .tldr export a file tldraw can actually open

The .tldr snapshot is the data-recovery escape hatch for #2132: it is what a
user falls back on to get their canvas out and open it elsewhere. It did not
work, and nothing tested that it did.

Two independent faults, both verified against tldraw 4.5.12 rather than assumed:

1. Wrong container entirely. We emitted {schema, store{}}, the shape of an
   in-memory store snapshot. A .tldr FILE is
   {tldrawFileFormatVersion, schema, records[]}. tldraw's tldrawFileValidator
   requires all three, so parseTldrawJsonFile threw and returned notATldrawFile
   -- a flat refusal to open, before it ever looked at the content. The declared
   schemaVersion 2 also lacked the required sequences dict.

2. Shapes stock tldraw cannot render. note/link/image were typed taos-note /
   taos-link / taos-image, our own shape utils. Nobody else's tldraw has them.
   On the live Pi that is 55 of 64 live elements. The rest were geo shapes
   carrying taos_* keys in props, which tldraw rejects as unknown props.

Fixed by emitting the real envelope with a serialized schema captured from
tldraw, mapping every kind onto native note/text/geo shapes, and moving taOS
provenance into meta, which tldraw carries through untouched.

Deliberately does NOT pass through a user_shape's literal tldraw_shape blob,
despite that being the highest-fidelity option. One invalid record makes tldraw
reject the WHOLE file, so a single stale blob would cost the user their entire
board in the one file whose only job is recovery. That risk is concrete here:
we declare a fixed schema constant, so tldraw runs no migration on a blob
written by an older version. The lossless copy belongs in the canonical JSON
export instead. The raw blob stays untouched in the DB payload either way.

Index keys are generated properly rather than hardcoded to "a1": tldraw
validates them, so "a10" is rejected outright (a fraction may not end in 0),
and reusing one key discards z-order.

Also sets Content-Disposition, without which the browser renders the JSON
inline and the user never gets a file.

Verified on real data, not just fixtures: all three live Pi boards exported and
loaded into a stock tldraw with every shape intact (22/22, 25/25, 17/17).

* test(canvas): update the two tests that pinned the old .tldr shape

Both asserted "store" in body, which encoded the bug this branch fixes: the
in-memory {schema, store{}} store-snapshot shape rather than the
{tldrawFileFormatVersion, schema, records[]} file envelope tldraw actually
accepts. They now assert the file format and read shapes out of records.

My miss: these live under tests/projects/ and I only ran the files I had
touched, so CI found them instead of me. Swept the tree for any other consumer
of the old shape and there are none.

* fix(canvas): make every field read in the .tldr export non-fatal

CodeRabbit and Qodo independently caught the same bug, and they were right.
_element_text did payload.get(...), but the `or {}` guard only catches falsy
payloads: a truthy non-dict (a list or bare string decoded from the DB) reached
.get() and raised AttributeError. el["id"], el["x"], el["kind"] could equally
raise KeyError on an absent column.

Any of those aborts _build_tldraw_snapshot for the WHOLE project. This is the
recovery export, so it runs precisely on the boards whose rows are already
ragged, and the failure mode was that one odd row costs the user every other
shape they were trying to rescue. That directly contradicts the rule this file
already states, which makes it a bug in the code rather than in the intent.

Every read now degrades to a safe default instead of raising. A string payload
becomes the shape's label, since if that is all the row holds it is the most
useful thing to show. _num_or excludes bool deliberately: bool is an int
subclass, so True would otherwise sail through as a width of 1.0 and silently
misplace a shape rather than fall back.

Verified against tldraw's real parser, not just unit tests: the live Pi board
with malformed rows spliced in exports 67 elements and loads all 67 shapes in a
stock tldraw.

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: hognek <hognek@gmail.com>
Co-authored-by: Hogne <227774406+hognek@users.noreply.github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants