Skip to content

Latest commit

 

History

History
271 lines (207 loc) · 26.3 KB

File metadata and controls

271 lines (207 loc) · 26.3 KB

Decision Log

One short entry per meaningful decision. Newest at the top. Cite sources for ideas borrowed from papers or other projects (ideas only — implementations here are original).

2026-07-07 — Landing video: HyperFrames (HTML→MP4), not Higgsfield

Higgsfield is out: its 7-Day Unlimited pass covers the web app but not MCP/API generation (which bills 0-balance workspace credits — verified, submits failed). Pivoted the cinematic assets to HyperFrames: author each as an HTML/CSS/GSAP composition and render to MP4 locally and free (Playwright Chromium + ffmpeg). This fits the project's free-forever ethos better than a paid video API, and the sources are committed (reproducible/editable) under landing/motion/<slot>/. The flagship "a skill is born" (arc-reactor core → emits skill card → scanline self-test → green check → settle) is rendered: 1080p, 12s, 1.16 MB, wired into the landing page and frame-verified (core region 94% cyan, near-black bg). Requirements learned: vendor GSAP locally (CDN caused render nav-timeout), transforms-only motion (not top), renderer-bundled fonts, AA contrast — all enforced by hyperframes check. Installed ffmpeg via winget (Gyan.FFmpeg) since HyperFrames' encode needs ffmpeg+ffprobe on PATH.

Plugin/skill installs the user listed (/plugin ..., npx skills add ...) were NOT run: /plugin is an interactive CLI dialog unavailable in this session, and both install into global/user Claude scope, which the project constitution (§12) forbids the agent from doing — that's the owner's call in an interactive shell. HyperFrames was already available as project skills, so it covered the need.

2026-07-07 — Premium front door lives on a landing page, not in the app

Asked to make Jarvis a premium experience with Higgsfield-generated video. Key scope call: cinematic video does not belong inside the desktop app — the constitution requires offline-first, 60fps, free-forever, "premium means powerful not bloated," so autoplaying heavy clips in the HUD is a regression by our own standard. The genuine missed opportunity is that the project has no public landing page — the front door that turns a visitor into a stargazer, and the one place every performance rule (WebM/AV1, lazy-load, Core Web Vitals, prefers-reduced-motion) actually applies. Built it as a standalone static site under landing/ (no framework, no build) so it can never touch the app bundle or offline guarantees; instant load is the premium signal (~28 KB uncompressed, video-free baseline).

Design principle carried over from the app: one glow. The hero is the live arc-reactor canvas (vanilla port of ArcCore) — genuinely premium, zero assets, zero bytes, so the first impression needs nothing generated. Generated video is reserved for moments that earn it, mapped model-to-job in landing/STORY.md: Veo 3.1 for atmospheric loops, Seedance 2.0 for the single flagship "a skill is born" sequence (keyframes first), WAN 2.6 to restyle a real screen-recording of undo (authentic-but-stylized, no fabricated footage/testimonials), MiniMax for cheap drafts. Kling deliberately unused (no recurring character needed).

Blocked on credits, not code: Higgsfield balance is 0 credits on a Plus plan — every generation (even keyframes) draws credits, and topping up / starting a paid trial is the owner's financial call, so nothing was generated. Instead every slot ships a designed SVG poster (the current experience and the reduced-motion/load fallback) with lazy AV1/WebM+MP4 plumbing already wired; when credits exist, each asset is one command per landing/assets/ASSETS.md and flipping data-has-asset="true".

2026-07-07 — Voice v0: OS voices out, feature-detected recognition in

§6.4 first slice. Output: speechSynthesis in the webview — WebView2 exposes the Windows voices, so spoken replies are free, offline, and require zero downloads; replies are sanitized for speech (code blocks summarized, URLs and markdown stripped, length capped at a sentence boundary — pure, tested) and a pickVoice heuristic prefers natural/neural English voices. Barge-in v0: sending a message or starting the mic cancels playback; the core gets speaking and listening motion states. Input: SpeechRecognition is feature-detected, never assumed — WebView2 historically doesn't ship it, so the mic button explains honestly and points at the plan instead of failing silently. Voice v1 (ROADMAP) is fully local Whisper STT in Rust — candidate paths are whisper-rs (cmake/C++ build, risky on this ARM64 setup) vs. candle's pure-Rust whisper (heavier deps, no native build); decide after a build spike. Wake word, VAD, and continuous conversation come after input works locally. Everything voice is optional and off by default.

2026-07-07 — Replay & undo v1: capture inverses, revert don't rewrite, audit determinism

§5.4 closes the hero-feature set on three principles. (1) Inverse state is captured at write time, not reconstructed later: note.saved events carry the previous content, chat events carry their memory row ids — undo reads the event, never guesses. (2) Undo appends, never rewrites: skill undo is a git-revert-style rollback (previous sources come back as a NEW version, history intact), and every undo logs its own undo.* event, so the timeline stays a truthful append-only record even about regrets. Wipes and reflections are declared irreversible rather than half-reversed. (3) Replay is verified, not narrated: instead of a replay animation, v1 ships a replay audit — rebuild the conversation purely from the log (honoring undos and wipes) and diff it order-preserving against the live database, reporting matched/missing/extra. Deterministic means the two agree; drift is shown, not hidden. Grounded in the determinism-faithfulness / replayable-agent literature. Step-through session player deferred to v2.

2026-07-07 — Confidence v0: verbalized self-rating, surfaced not trusted

§5.3 shipped as a single-call verbalized rating: the chat system prompt requires a trailing [confidence: NN] line, extracted and stripped by the command layer so it travels as data, never as text. Below 40 the model is instructed to ask ONE clarifying question instead of answering. Grounded in "Agentic Uncertainty Reveals Agentic Overconfidence": verbalized confidence is poorly calibrated (live llama3.2 rated a trivial fact at 100 and its own clarifying question at 80) but directionally useful — v0 therefore surfaces it (gauge dial on the core, per-message label, event log) rather than gating anything hard on it. Live probes confirmed both behaviors: factual → answer + parseable marker; ambiguous ("fix it, it's broken") → clarifying question. A second-call self-critique would rate better but doubles latency/free-tier spend; rejected for v0. Calibration tracking (predicted vs. actual via the reflection pass) is the v1 path in ROADMAP.

2026-07-07 — Reasoning-memory v0: the event log is the experience stream

§5.2 implemented as a reflection pass over the existing event log rather than a parallel trace store — the log already records what was tried and what happened (chat outcomes, authoring attempts, skill runs, failures). A pass digests events past a watermark (reflection.last_event_id fact), asks the model for at most 3 one-sentence lessons (JSON contract, tolerant parser — live llama3.2 probe revealed it returns a bare object instead of an array, now accepted and regression-tested), stores them in a new insights table (schema v2, migration tested against a hand-built v1 database), and rides the freshest lessons into both the chat and skill-authoring system prompts. Periodicity without an autonomous loop: the frontend pings reflect_if_due after each turn; it fires only when 20+ new messages accumulated. Watermark advances even on an empty harvest so the same events are never re-digested. Idea credits: ReasoningBank, "Hindsight is 20/20", Reflexion. Scoring/decay and selective forgetting deferred to v1 (tracked in ROADMAP).

2026-07-06 — Skill authoring v1: trust the harness, not the model

The assistant now authors skills from a natural-language request: strict JSON contract (name/description/code/test), parse tolerant of fences/chatter, save, run the bundled test, and on failure loop the error back to the model (Reflexion) for up to 3 attempts; a never-passing draft is saved flagged and reported honestly. Validated against the owner's live llama3.2 before shipping: the 3B model produces well-formed JSON but buggy Rhai (${} interpolation habits, off-by-one tests) — which confirmed the design premise that correctness must come from executing the test, never from trusting the model. The system prompt gained a concrete correct example and an explicit anti-${} rule after those live probes. Authoring quality scales with the model: local 3B will flag more drafts; a free Groq 70B key one-shots most requests. Follow-up in ROADMAP: Ollama structured output (format: json) to guarantee parseability.

2026-07-06 — Skill engine v0: Rhai as the skill runtime

Skills (§5.1) are Rhai scripts, not Python/WASM/JS: Rhai is pure Rust (no system deps, compiles in CI unchanged), sandboxed by construction (scripts see only language built-ins — no fs, no network, nothing we don't register), and hard-capped per execution (200k operations, bounded call depth / string / collection sizes) so a runaway loop terminates instead of hanging the assistant — there's a test proving it. Contract: fn run(input) plus a bundled fn test(); the test runs on every save and re-save; a failing skill is flagged and refused at run time, never used blindly (Reflexion-style refinement loop). Versions are integers; the previous source is archived to history/v<N>/ on every update — cheap provenance for the future replay/undo story. Idea credits: Voyager's skill library, MUSE-Autoskill, Reflexion. v1 (LLM authors skills from chat) is next in ROADMAP; v0 deliberately ships the engine + manual authoring UI so the harness is proven before the model writes code into it.

2026-07-06 — Router hardening: cache + cooldown, local model exempt

Free-tier hygiene (§2) as pure clock-injected bookkeeping in core/reliability.rs: a TTL+capacity response cache (dedupes identical requests — double-sends and retries — 10 min TTL, 64 entries) and a per-provider cooldown tracker (exponential 30s→10min backoff, honors Retry-After) that the router consults before calling cloud providers. Deliberate asymmetry: Ollama is exempt from both penalties — local inference is unlimited and private, so penalizing it only hurts the user. Cache hits are marked cached: true and shown in the reply meta, because silently serving stale answers would violate the honesty principle. Provider call errors became structured (CallError::Http{status, retry_after}) so backoff decisions don't parse error strings.

2026-07-06 — Event log v0: JSONL, append-only, tolerant reader

Chose plain JSONL over SQLite for the event log even though SQLite is already in the app: an append-only text file is trivially greppable, corruption-isolated per line (the reader skips bad lines instead of failing — tested with a simulated torn write), and matches the replay literature's framing of an immutable event stream. Ids are monotonic and resume across restarts. Logging is best-effort by design (log_event swallows errors): the log must never take the assistant down. Full chat text is logged because deterministic replay (§5.4) needs it; the file lives in the same local app-data dir as memory and is covered by the same export/wipe story (wiring the log into export is a follow-up). The EVENTS tab is read-only on purpose — no replay/undo buttons until the engine exists.

2026-07-06 — Multi-view HUD: tabs + palette, honest telemetry only

Added tab navigation (chat / notes / memory, Ctrl+1-3) and a Ctrl+K command palette (subsequence-scored filtering, pure + unit-tested) instead of pulling in a router dependency — three views don't justify react-router yet; revisit when deep-linkable views land (§6.3). The reference wallpapers are full of gauges, so the rule for live data is: only real numbers — CPU/RAM/uptime come from a new get_telemetry command backed by sysinfo, plus actual message/fact/note counts and a wall clock. No invented readouts. New views surface the backend that already existed (notes tool, memory export/wipe) rather than faking future features (skill library, replay) — those stay in the roadmap until their engines exist.

2026-07-06 — HUD visual language: instrument panel, one glow

Owner supplied film-Jarvis reference imagery (wallpapercave.com/jarvis-wallpapers). Reading the references closely: they are restrained — mostly monochrome steel hairlines and tiny uppercase mono labels, with exactly one bright element (the circular core). Adopted as a hard rule: glow is reserved for the arc-reactor core (canvas ArcCore, the app's signature element — its offline/idle/thinking states are the primary trust signal) and everything else stays hairlines and type. Light theme reinterprets the same instrument as a blueprint on paper rather than a dimmed dark theme. System fonts only (offline-first, no font downloads). Original implementation; imagery used as mood reference only.

2026-07-05 — Bootstrap stack: Tauri v2 + React/TS + Rust core

  • Desktop shell: Tauri v2 (default per the project brief): light installers, low idle RAM, Rust backend suits the future low-latency voice pipeline. Electron kept as a documented fallback only.
  • Frontend: React 19 + TypeScript + Vite. Mainstream ecosystem, works with Framer Motion/Motion One for the §6.2 motion system later. Design tokens as CSS variables from day one, themed via data-theme on the root element.
  • Core logic in Rust, Tauri-independent. src-tauri/src/core/ (router, memory, tools) has no Tauri types so every module unit-tests without a webview. Tauri commands in lib.rs are a thin adapter layer.
  • Memory v0: SQLite via rusqlite (bundled) — durable, zero-dependency install, schema-migration table from the start so upgrades never lose data. Vector recall deferred to M1 (candidate: sqlite-vec to stay single-store). Data dir: OS app-data dir, overridable with JARVIS_DATA_DIR (used by tests and dev).
  • Router v0 providers: Ollama → Groq → OpenRouter :free. Ollama first because local is unlimited and private. Providers implement one trait; the router walks them in priority order and returns a structured "no provider" onboarding message rather than an error when nothing is configured. Gemini/Cerebras adapters deferred until the trait proves itself.
  • First built-in tool: local notes scoped to the app data dir — real utility, zero external side effects, exercises the tool interface without needing the approval-gate machinery yet.
  • Licensing: Apache-2.0 (already in repo from initial commit).
  • CI on ubuntu-latest with Tauri system deps for the Rust job (fastest runners); cross-platform packaging matrix deferred to release milestones per brief §8.
  • Prompt/agent files are gitignored (.claude/, CLAUDE.md, docs/agentbrief*) at the owner's request — continuity files stay local to this machine.

Idea credits for planned hero features (tracked in ROADMAP): Voyager skill library; MUSE-Autoskill; Reflexion; ReasoningBank; "Hindsight is 20/20" agent memory; Mem0 selective memory; "Agentic Uncertainty Reveals Agentic Overconfidence"; 2026 replayable-agent/determinism literature. To be re-read before each respective implementation.

Reflection v2: meaning-based duplicates, and forgetting you can undo

Duplicate detection was token overlap, which cannot see a paraphrase that shares no vocabulary — and reflection produces exactly those. Lessons are now embedded when they're created (backfilled by the existing index pass) and compared by cosine, falling back to word overlap when either side has no vector. Mismatched vector widths also fall back rather than producing a confident wrong number: different embedding models are not comparable.

The cosine floor (0.92) is far stricter than the token one (0.6) on purpose. Sentence embeddings put any two English sentences about a shared subject at 0.7-0.85, so a threshold that reads as strict for word overlap merges lessons that merely share a topic. The logged reason names the method it used, because 0.93 overlap and 0.93 cosine mean very different things.

Forgetting became a soft delete (schema v5). The app deciding what matters to someone is a guess, and a decay curve is a crude one, so a drop now lands in a "forgotten" tab with its reason and can be restored in one click. A hard DELETE gave the user no way to disagree with a judgement the app made on its own.

One UI trap worth recording: dimming a dropped card with opacity silently did nothing, because .reflect-card carries an enter animation with fill mode both, and a filled animation beats a normal declaration. The dimming moved to the card's contents, which aren't animated.

Voice v3 barge-in: calibrate against the echo, never transcribe

Voice v2 closed the microphone while the assistant spoke, which is the right default — an open mic during playback is how a hands-free loop starts answering its own voice — but it also made interrupting impossible, and being unable to cut off a wrong answer is the worst moment to be stuck.

The obstacle is echo. With no AEC, a mic next to a speaker hears the assistant at least as loudly as the user, so a fixed threshold either fires on the assistant's own voice or needs a shout. core::bargein measures instead: for the first half-second of playback it only calibrates, learning how loud this assistant is in this room at this volume, and the threshold is a multiple of the mean-and-peak blend of that measurement. A sustain window separates a person starting a sentence from a cough or a door.

Two properties were treated as non-negotiable. It fires at most once per answer, because the tail of the same sentence would otherwise produce a stream of duplicate interruptions. And it never transcribes: audio is read for loudness and dropped frame by frame, so Phase::wants_audio still excludes Speaking and the assistant cannot hear itself into a request. wants_barge_monitor is a separate flag for exactly that reason, and isCapturing deliberately does not include it — the capture indicator is a privacy signal, and claiming a recording during loudness monitoring would be a lie.

Where it genuinely does not work is documented rather than papered over: on a laptop at high volume the echo floor can exceed normal speech, and no amount of threshold tuning fixes that. Headphones make it near-perfect.

The frontend races playback against the watcher with a settled guard, because a natural ending and an interruption could otherwise both advance the session and the second would move a session that had already gone on to something else. The watcher's budget is estimated from answer length (speechBudgetMs) and deliberately over-estimates: overshooting keeps the mic open a few seconds too long, while undershooting means the tail of a long answer cannot be interrupted.

Lesson utility: measure whether reflection helps, and demote rather than delete

Reflection had been tested everywhere except the place that matters: whether the answers got better. Published work finds an inverted-U where consolidated lessons help, then stop, then drag performance below a no-memory baseline, and the failure is silent by construction — nothing errors, the lessons still read sensibly, and the answers just get worse.

The baseline is the whole design. A per-lesson helpfulness rate on its own is meaningless: 60% is excellent if answers without lessons manage 40%, and alarming if they manage 90%. So the measurement's first-class citizen is the rate for graded answers that carried no lessons at all, and no verdict is reached until that baseline itself has enough evidence.

Two mistakes were available and both were avoided deliberately. Answers written before this shipped have no insight_ids field; reading that absence as "carried no lessons" would have filled the baseline with history that predates the feature and made every lesson look bad. Those turns are skipped instead. And a lesson with two ratings at 0% is not evidence of anything, so it stays unknown rather than being condemned — most lessons will never clear the minimum, which is the correct outcome given that grading is voluntary.

Acting on the result is deliberately gentle. A harmful lesson gets its score multiplied down in prompt selection, so it falls out of the prompt and fades sooner. It is not deleted, because the evidence is correlational: several lessons ride along in one answer, so credit is shared, and a lesson selected for hard questions will look worse than one selected for easy ones. top_for_prompt_weighted takes a weight function rather than importing core::utility, which keeps core::forgetting pure arithmetic over the bookkeeping and lets the caller supply the evidence.

One correction worth recording: the proposal claimed chat_send already logged which lessons rode along. It did not. mark_insights_used bumps a counter, which says a lesson was used at some point but not which answer it shaped, and the join needs the latter. The ids are now on the chat.assistant event.

Cost-weighted confidence: put the risk sensitivity in code, not the prompt

One ask-threshold covered everything, which treats being wrong in conversation and executing generated code as equally serious. The obvious fix is to tell the model to be more careful when the stakes are higher. That is the one thing the literature says it reliably will not do: models state their uncertainty adequately and then show no sensitivity to the penalty for being wrong, not even where abstaining is mathematically optimal (arXiv:2601.07767).

So required_confidence is a total function over ActionKind, ordered by how hard an action is to undo rather than how complicated it is. Totality is the safety property: a new capability is a compile error here, so it cannot be added and quietly inherit a permissive floor. The same reason classify is total.

The irreversible set is deliberately given a floor above 100 rather than 99. A 99 floor is still satisfiable by a confident model; 101 encodes "never on confidence alone" as arithmetic instead of as a comment someone can miss.

trust_ceiling is the other half: what a claim of 100 is actually worth given the measured record. A model running 40 points hot tops out at 60, and everything above that goes out of reach until the record improves. Two properties were treated as non-negotiable and both have tests. Self-maintenance stays at the base threshold, because a poor record must never gate off the reflection and indexing that would let the loop recover. And with no calibration record, nothing changes at all — the ceiling is 100 and planning is identical, so the feature cannot silently rewrite the loop's behaviour the day it ships.

assess_for is separate from assess on purpose. Raising the bar for answering because the user might act on the answer would make the assistant hedge constantly, which is exactly what the ask-threshold was tuned to avoid.

Two of my own test assertions were wrong here and the code was right, which is worth recording. A 30-point bias puts the ceiling at exactly 70, the authoring floor, so authoring is in reach — the boundary is inclusive. And a raw 60 against a floor of 75 is not "demoted": it never cleared the bar, so calibration changed nothing. demoted stays reserved for the case actually worth explaining to the user, where the raw number would have passed and the calibrated one did not.

Skill retirement: reuse the scoring engine, but not naively

A passing test proves a skill runs. It says nothing about whether running it helps, and the research this came from measured LLM-authored skills at +0.0pp against a no-skill baseline where human-curated ones gave +16.2pp (arXiv:2605.19576). Libraries accumulate and drift below baseline with no error signal.

The instruction was to reuse core::forgetting rather than write a second scoring engine, and that was right — but two obvious ways to do it are wrong, and both only showed up as failing tests.

Runs cannot map onto uses. uses raises a score, so a skill run thirty times and failed thirty times comes out looking well-established. Only successes count as evidence.

And the clock cannot run from created_at. A lesson goes stale as the situation it described recedes; a skill does not rot on the shelf. Decaying from creation retires a skill that works perfectly for the crime of being old, which is the opposite of the intent. It decays from the last run instead, so an actively-succeeding skill stays fresh however old it is.

That left decay unable to express "it runs and it doesn't work", because a recently-run failing skill looks fresh. So retirement has two independent reasons rather than one number: a success rate below half, and staleness. Keeping them separate means the logged reason can say which applied, and 30-of-30-failed reads very differently from unused-for-months.

Three protections, each load-bearing. New skills are exempt. Fewer than ten runs is not a record — their ablation found retirement without an evidence minimum scored worse than never retiring, below the no-skill baseline, so removing this would reproduce that. And capacity pressure only removes skills that have a record, or library size rather than evidence would decide what survives.

Retirement deactivates and never deletes, matching the soft delete Reflection v2 introduced. A skill is the user's code; the library forming an opinion about it does not entitle it to throw the code away. Running a retired skill fails loudly with the reason and how to restore it, so anything that referenced it breaks visibly instead of silently doing nothing. Saving a new version un-retires with a clean record, because fixing a skill is the user overruling the retirement and the fix should not inherit the failures that retired its predecessor.

last_run_at is stored explicitly rather than inferred from updated_at, which also moves on save and re-test. A re-test is not a use, and conflating them would keep an unused skill looking fresh indefinitely.

Adding ActionKind::RetireSkills made required_confidence fail to compile, which is the deny-by-default property from #43 working on its first real test: a new capability could not be added without someone deciding what confidence it needs.