diff --git a/docs/00_overview/BACKLOG_DASHBOARD.md b/docs/00_overview/BACKLOG_DASHBOARD.md index 4955cfab..c4770360 100644 --- a/docs/00_overview/BACKLOG_DASHBOARD.md +++ b/docs/00_overview/BACKLOG_DASHBOARD.md @@ -2,7 +2,7 @@ # RelyLoop BACKLOG Dashboard -_Reflects feature-folder state as of **2026-06-02** (latest mtime of any planned/implemented feature `.md` file). Regenerated by `make dashboard` and the `mvp1-dashboard-regen` pre-commit hook. For the rich local view (filter chips, type colors), open [`backlog_dashboard.html`](backlog_dashboard.html) in a browser._ +_Reflects feature-folder state as of **2026-06-03** (latest mtime of any planned/implemented feature `.md` file). Regenerated by `make dashboard` and the `mvp1-dashboard-regen` pre-commit hook. For the rich local view (filter chips, type colors), open [`backlog_dashboard.html`](backlog_dashboard.html) in a browser._ ## Next up diff --git a/docs/00_overview/DASHBOARD.md b/docs/00_overview/DASHBOARD.md index 6aa2ffeb..f589de86 100644 --- a/docs/00_overview/DASHBOARD.md +++ b/docs/00_overview/DASHBOARD.md @@ -1,13 +1,13 @@ # RelyLoop — Release Roadmap -_Top-level index across MVP1 → GA v1+ as of **2026-06-02**. Click a release name to drill into the per-release dashboard. Theme labels sourced from [`docs/01_architecture/tech-stack.md` §"Canonical release matrix"](../01_architecture/tech-stack.md). For the rich local view, open [`dashboard.html`](dashboard.html) in a browser._ +_Top-level index across MVP1 → GA v1+ as of **2026-06-03**. Click a release name to drill into the per-release dashboard. Theme labels sourced from [`docs/01_architecture/tech-stack.md` §"Canonical release matrix"](../01_architecture/tech-stack.md). For the rich local view, open [`dashboard.html`](dashboard.html) in a browser._ ## Releases | Release | Theme | Progress | Status | |---|---|---|---| | [MVP1 / v0.1](MVP1_DASHBOARD.md) | The Loop | 94 / 94 scoped done | **Complete** | -| [MVP2 / v0.2](MVP2_DASHBOARD.md) | Three-Engine + Real Signals | 13 / 24 scoped done · 24 remaining | **In progress** | +| [MVP2 / v0.2](MVP2_DASHBOARD.md) | Three-Engine + Real Signals | 13 / 24 scoped done · 25 remaining | **In progress** | | MVP3 / v0.3 | Observable | — | **Not yet scoped** | | GA v1 / v1.0 | Production-ready | — | **Not yet scoped** | diff --git a/docs/00_overview/MVP1_DASHBOARD.md b/docs/00_overview/MVP1_DASHBOARD.md index 838b4912..0831cefc 100644 --- a/docs/00_overview/MVP1_DASHBOARD.md +++ b/docs/00_overview/MVP1_DASHBOARD.md @@ -2,7 +2,7 @@ # RelyLoop MVP1 Dashboard -_Reflects feature-folder state as of **2026-06-02** (latest mtime of any planned/implemented feature `.md` file). Regenerated by `make dashboard` and the `mvp1-dashboard-regen` pre-commit hook. For the rich local view (filter chips, type colors), open [`mvp1_dashboard.html`](mvp1_dashboard.html) in a browser._ +_Reflects feature-folder state as of **2026-06-03** (latest mtime of any planned/implemented feature `.md` file). Regenerated by `make dashboard` and the `mvp1-dashboard-regen` pre-commit hook. For the rich local view (filter chips, type colors), open [`mvp1_dashboard.html`](mvp1_dashboard.html) in a browser._ ## Next up diff --git a/docs/00_overview/MVP2_DASHBOARD.md b/docs/00_overview/MVP2_DASHBOARD.md index 04ff6c10..82814c99 100644 --- a/docs/00_overview/MVP2_DASHBOARD.md +++ b/docs/00_overview/MVP2_DASHBOARD.md @@ -2,7 +2,7 @@ # RelyLoop MVP2 Dashboard -_Reflects feature-folder state as of **2026-06-02** (latest mtime of any planned/implemented feature `.md` file). Regenerated by `make dashboard` and the `mvp1-dashboard-regen` pre-commit hook. For the rich local view (filter chips, type colors), open [`mvp2_dashboard.html`](mvp2_dashboard.html) in a browser._ +_Reflects feature-folder state as of **2026-06-03** (latest mtime of any planned/implemented feature `.md` file). Regenerated by `make dashboard` and the `mvp1-dashboard-regen` pre-commit hook. For the rich local view (filter chips, type colors), open [`mvp2_dashboard.html`](mvp2_dashboard.html) in a browser._ ## Next up @@ -20,15 +20,15 @@ Plan approved; run /impl-execute to ship | Metric | Value | |---|---| -| Filed under MVP2 | **43** folders total (done + specced not-done + idea backlog + bugs) | +| Filed under MVP2 | **44** folders total (done + specced not-done + idea backlog + bugs) | | Specced features done | **13 / 24** (54%) — of features *past the idea stage* (those with a spec); the idea backlog below is NOT in this denominator, so 100% ≠ release complete | -| Pending work | **28** items (every not-done feat/infra/chore/bug across all priorities) | +| Pending work | **29** items (every not-done feat/infra/chore/bug across all priorities) | | → P0 — do next | **0** unblocking / paying daily cost | | → P1 | **0** high-value, ready when P0 clears | -| → P2 (default) | 24 important to file, not blocking | +| → P2 (default) | 25 important to file, not blocking | | → Backlog | 4 captured for record, not planned | | Open bugs | 9 | -| Legacy "Path to MVP2" | 24 items — scoped-not-done + bugs + chore-ideas only (excludes feat/infra ideas) | +| Legacy "Path to MVP2" | 25 items — scoped-not-done + bugs + chore-ideas only (excludes feat/infra ideas) | | Backlog ideas | 4 idea-only feat/infra (not yet scoped into MVP2) | | In flight | 0 feature(s) actively shipping | @@ -80,25 +80,26 @@ _None._ _None._ -### Idea (15) +### Idea (16) | # | Priority | Feature | Type | One-liner | Depends on | Status | |---|---|---|---|---|---|---| | 1 | P2 | [infra_openapi_types_freshness_gate](planned_features/02_mvp2/infra_openapi_types_freshness_gate/idea.md) | Infra | Idea — Phase 2 of [`infra_generated_artifact_freshness_gate`](../infra_generated_artifact_freshness_gate/feature_spec.md), extracted to its own folder | — | Idea — Phase 2 of [`infra_generated_artifact_freshness_gate`](../infra_generated_artifact_freshness_gate/feature_spec.md), extracted to its own folder | | 2 | P2 | [infra_smoke_fork_pr_secret_skip](planned_features/02_mvp2/infra_smoke_fork_pr_secret_skip/idea.md) | Infra | `.github/workflows/pr.yml` triggers on `pull_request:` ([pr.yml:43](../.github/workflows/pr.yml)) — **not** `pull_request_target`. GitHub deliberately withholds repository secrets from workflows trigg | — | Idea — tangential discovery while merging PR #387 (`chore_arq_pool_aclose_deprecation`) | | 3 | P2 | [chore_demo_reseed_partial_completion_fast_test](planned_features/02_mvp2/chore_demo_reseed_partial_completion_fast_test/idea.md) | Chore | `infra_solr_ci_readiness` made the demo reseed engine-tolerant: when an engine is unreachable, its scenario is skipped, the reseed completes with `status="complete"` and a non-empty `scenarios_skipped | — | Idea — tangential discovery during `infra_solr_ci_readiness` Story 1.2 implementation | -| 4 | P2 | [chore_solr_post_pipeline_followups](planned_features/02_mvp2/chore_solr_post_pipeline_followups/idea.md) | Chore | The 13-story `infra_adapter_solr` execution surfaced several follow-on items that fit neither the original spec nor any sister feature folder. None block the MVP2 Solr release — they're operator-exper | — | Idea — tangential observations from `infra_adapter_solr` end-to-end | -| 5 | P2 | [chore_ubi_hybrid_template_render](planned_features/02_mvp2/chore_ubi_hybrid_template_render/idea.md) | Chore | Idea — contract decision deferred (NOT a worker bug) | — | Idea — contract decision deferred (NOT a worker bug) | -| 6 | P2 | [bug_e2e_teardown_chain_node_delete_500](planned_features/02_mvp2/bug_e2e_teardown_chain_node_delete_500/idea.md) | Bug | The E2E global-teardown deletes seeded rows in a fixed order (per `chore_e2e_test_rows_isolation` Story 1.2 cleanup registration). For auto-followup **chains**, the seeded nodes are `queued` studies c | — | Idea — tangential discovery during `feat_overnight_autopilot` (Story 4.2 E2E, PR forthcoming) | -| 7 | P2 | [bug_relyloop_spec_ubi_section_drift](planned_features/02_mvp2/bug_relyloop_spec_ubi_section_drift/idea.md) | Bug | [`docs/00_overview/relyloop-spec.md`](relyloop-spec.md) §"Click-derived judgments — OpenSearch UBI as the engine-neutral primary path" (line ~706) carries two staleness bugs from the 2026-05-27 releas | — | Idea — captured during `feat_ubi_judgments` preflight (2026-05-29) | -| 8 | P2 | [bug_reseed_failure_blocks_retry_arq_singleton_dedup](planned_features/02_mvp2/bug_reseed_failure_blocks_retry_arq_singleton_dedup/idea.md) | Bug | `run_demo_reseed` is enqueued with a fixed Arq job id `demo_reseed:singleton` (the singleton concurrency guard). When a run reaches a terminal state, Arq stores its **result** under `arq:result:demo_r | — | Idea — tangential discovery while verifying `fix(demo): add Solr (8983) to the reseed engine host-URL mapping` (branch `feat_demo_reseed_solr_and_steplog`) | -| 9 | P2 | [bug_seed_meaningful_demos_silent_bulk_errors](planned_features/02_mvp2/bug_seed_meaningful_demos_silent_bulk_errors/idea.md) | Bug | [`scripts/seed_meaningful_demos.py:917-935`](../../scripts/seed_meaningful_demos.py#L917-L935) bulk-indexes 1000 Amazon ESCI products into a dedicated index per demo scenario: | — | Idea — captured during `bug_smoke_seed_es_unavailable_shards_race` Phase 2.5 tangential sweep | -| 10 | P2 | [bug_studies_detail_vitest_intermittent_timeout](planned_features/02_mvp2/bug_studies_detail_vitest_intermittent_timeout/idea.md) | Bug | Under the full `pnpm test` run (`vitest run`, default worker pool), the Study-detail-page render test sometimes blocks past the 5 s `testTimeout` default — but the test itself is data-driven from mock | — | Idea — captured during `chore_template_library_expansion` post-impl tangential sweep | -| 11 | P2 | [bug_webhook_concurrent_merge_race_timing_sensitive](planned_features/02_mvp2/bug_webhook_concurrent_merge_race_timing_sensitive/idea.md) | Bug | Idea — surfaced during `bug_demo_clusters_unreachable_in_healthz` PR #236 CI. | — | Idea — surfaced during `bug_demo_clusters_unreachable_in_healthz` PR #236 CI. | -| 12 | Backlog | [feat_fts_rank_ordering](planned_features/02_mvp2/feat_fts_rank_ordering/idea.md) | Feature | `feat_data_table_primitive` shipped filter-only FTS — `?q=foo` matches rows where `search_vector @@ plainto_tsquery('english', 'foo')` is true but orders results by `created_at DESC, id DESC` (the def | — | Idea — deferred from `feat_data_table_primitive` (MVP1) per spec §16. | -| 13 | Backlog | [infra_arq_subprocess_test](planned_features/02_mvp2/infra_arq_subprocess_test/idea.md) | Infra | Idea (deferred from `feat_study_lifecycle` Phase 2 / PR #25 final GPT-5.5 review). Still applicable as of 2026-05-14: the three in-process tests cited below still cover the resume contract correctly; | — | Idea (deferred from `feat_study_lifecycle` Phase 2 / PR #25 final GPT-5.5 review). Still applicable as of 2026-05-14: the three in-process tests cited below still cover the resume contract correctly; a subprocess test would add a narrow Arq-version-regression guard. | -| 14 | Backlog | [chore_auto_followup_parent_advisory_lock](planned_features/02_mvp2/chore_auto_followup_parent_advisory_lock/idea.md) | Chore | The shipped `feat_auto_followup_studies` worker uses a two-layer idempotency scheme: | — | Idea — captured as a standalone file to resolve broken cross-references in `feat_auto_followup_studies` D-11 + plan F2 + `bug_auto_followup_completed_parent_stop_chain_race/idea.md`. The slug was coined 2026-05-24 in D-11 but only existed as descriptive prose across other documents until now. | -| 15 | Backlog | [bug_chat_long_conversation_truncation](planned_features/02_mvp2/bug_chat_long_conversation_truncation/idea.md) | Bug | [`backend/app/services/agent_chat.send_user_message`](../../backend/app/services/agent_chat.py) defensively caps the OpenAI history at the most recent `HISTORY_MAX_MESSAGES = 100` messages… | — | Held for MVP2 (decided 2026-05-13). Folder renamed with `_mvp2` suffix to make the deferral visible at-a-glance in `ls docs/00_overview/planned_features/`. Resume work when MVP2 starts — no technical dependency on MVP2 infra (audit_log is N/A; Langfuse is convenience only); the deferral is scope discipline + zero current impact (latent bug, no operator has hit the 100-message cap). | +| 4 | P2 | [chore_pr_yml_parallelize_backend_job](planned_features/02_mvp2/chore_pr_yml_parallelize_backend_job/idea.md) | Chore | `.github/workflows/pr.yml` has a job named `backend (lint + typecheck + tests + coverage)` that runs four sequential things in one job: ruff/lint, mypy, the full pytest matrix (unit + integration + co | — | Idea — captured during PR #426 CI watch | +| 5 | P2 | [chore_solr_post_pipeline_followups](planned_features/02_mvp2/chore_solr_post_pipeline_followups/idea.md) | Chore | The 13-story `infra_adapter_solr` execution surfaced several follow-on items that fit neither the original spec nor any sister feature folder. None block the MVP2 Solr release — they're operator-exper | — | Idea — tangential observations from `infra_adapter_solr` end-to-end | +| 6 | P2 | [chore_ubi_hybrid_template_render](planned_features/02_mvp2/chore_ubi_hybrid_template_render/idea.md) | Chore | Idea — contract decision deferred (NOT a worker bug) | — | Idea — contract decision deferred (NOT a worker bug) | +| 7 | P2 | [bug_e2e_teardown_chain_node_delete_500](planned_features/02_mvp2/bug_e2e_teardown_chain_node_delete_500/idea.md) | Bug | The E2E global-teardown deletes seeded rows in a fixed order (per `chore_e2e_test_rows_isolation` Story 1.2 cleanup registration). For auto-followup **chains**, the seeded nodes are `queued` studies c | — | Idea — tangential discovery during `feat_overnight_autopilot` (Story 4.2 E2E, PR forthcoming) | +| 8 | P2 | [bug_relyloop_spec_ubi_section_drift](planned_features/02_mvp2/bug_relyloop_spec_ubi_section_drift/idea.md) | Bug | [`docs/00_overview/relyloop-spec.md`](relyloop-spec.md) §"Click-derived judgments — OpenSearch UBI as the engine-neutral primary path" (line ~706) carries two staleness bugs from the 2026-05-27 releas | — | Idea — captured during `feat_ubi_judgments` preflight (2026-05-29) | +| 9 | P2 | [bug_reseed_failure_blocks_retry_arq_singleton_dedup](planned_features/02_mvp2/bug_reseed_failure_blocks_retry_arq_singleton_dedup/idea.md) | Bug | `run_demo_reseed` is enqueued with a fixed Arq job id `demo_reseed:singleton` (the singleton concurrency guard). When a run reaches a terminal state, Arq stores its **result** under `arq:result:demo_r | — | Idea — tangential discovery while verifying `fix(demo): add Solr (8983) to the reseed engine host-URL mapping` (branch `feat_demo_reseed_solr_and_steplog`) | +| 10 | P2 | [bug_seed_meaningful_demos_silent_bulk_errors](planned_features/02_mvp2/bug_seed_meaningful_demos_silent_bulk_errors/idea.md) | Bug | [`scripts/seed_meaningful_demos.py:917-935`](../../scripts/seed_meaningful_demos.py#L917-L935) bulk-indexes 1000 Amazon ESCI products into a dedicated index per demo scenario: | — | Idea — captured during `bug_smoke_seed_es_unavailable_shards_race` Phase 2.5 tangential sweep | +| 11 | P2 | [bug_studies_detail_vitest_intermittent_timeout](planned_features/02_mvp2/bug_studies_detail_vitest_intermittent_timeout/idea.md) | Bug | Under the full `pnpm test` run (`vitest run`, default worker pool), the Study-detail-page render test sometimes blocks past the 5 s `testTimeout` default — but the test itself is data-driven from mock | — | Idea — captured during `chore_template_library_expansion` post-impl tangential sweep | +| 12 | P2 | [bug_webhook_concurrent_merge_race_timing_sensitive](planned_features/02_mvp2/bug_webhook_concurrent_merge_race_timing_sensitive/idea.md) | Bug | Idea — surfaced during `bug_demo_clusters_unreachable_in_healthz` PR #236 CI. | — | Idea — surfaced during `bug_demo_clusters_unreachable_in_healthz` PR #236 CI. | +| 13 | Backlog | [feat_fts_rank_ordering](planned_features/02_mvp2/feat_fts_rank_ordering/idea.md) | Feature | `feat_data_table_primitive` shipped filter-only FTS — `?q=foo` matches rows where `search_vector @@ plainto_tsquery('english', 'foo')` is true but orders results by `created_at DESC, id DESC` (the def | — | Idea — deferred from `feat_data_table_primitive` (MVP1) per spec §16. | +| 14 | Backlog | [infra_arq_subprocess_test](planned_features/02_mvp2/infra_arq_subprocess_test/idea.md) | Infra | Idea (deferred from `feat_study_lifecycle` Phase 2 / PR #25 final GPT-5.5 review). Still applicable as of 2026-05-14: the three in-process tests cited below still cover the resume contract correctly; | — | Idea (deferred from `feat_study_lifecycle` Phase 2 / PR #25 final GPT-5.5 review). Still applicable as of 2026-05-14: the three in-process tests cited below still cover the resume contract correctly; a subprocess test would add a narrow Arq-version-regression guard. | +| 15 | Backlog | [chore_auto_followup_parent_advisory_lock](planned_features/02_mvp2/chore_auto_followup_parent_advisory_lock/idea.md) | Chore | The shipped `feat_auto_followup_studies` worker uses a two-layer idempotency scheme: | — | Idea — captured as a standalone file to resolve broken cross-references in `feat_auto_followup_studies` D-11 + plan F2 + `bug_auto_followup_completed_parent_stop_chain_race/idea.md`. The slug was coined 2026-05-24 in D-11 but only existed as descriptive prose across other documents until now. | +| 16 | Backlog | [bug_chat_long_conversation_truncation](planned_features/02_mvp2/bug_chat_long_conversation_truncation/idea.md) | Bug | [`backend/app/services/agent_chat.send_user_message`](../../backend/app/services/agent_chat.py) defensively caps the OpenAI history at the most recent `HISTORY_MAX_MESSAGES = 100` messages… | — | Held for MVP2 (decided 2026-05-13). Folder renamed with `_mvp2` suffix to make the deferral visible at-a-glance in `ls docs/00_overview/planned_features/`. Resume work when MVP2 starts — no technical dependency on MVP2 infra (audit_log is N/A; Langfuse is convenience only); the deferral is scope discipline + zero current impact (latent bug, no operator has hit the 100-message cap). | ## Dependency graph diff --git a/docs/00_overview/backlog_dashboard.html b/docs/00_overview/backlog_dashboard.html index 81c9e238..89b77fe0 100644 --- a/docs/00_overview/backlog_dashboard.html +++ b/docs/00_overview/backlog_dashboard.html @@ -369,7 +369,7 @@

RelyLoop BACKLOG Dashboard

- Reflects feature-folder state as of 2026-06-02 (latest mtime of any + Reflects feature-folder state as of 2026-06-03 (latest mtime of any docs/00_overview/planned_features/ or docs/00_overview/implemented_features/ file). See state.md for the active branch context, diff --git a/docs/00_overview/dashboard.html b/docs/00_overview/dashboard.html index 99ff3921..5219f5e9 100644 --- a/docs/00_overview/dashboard.html +++ b/docs/00_overview/dashboard.html @@ -368,7 +368,7 @@

RelyLoop — Release Roadmap

- Top-level index across MVP1 → GA v1+ as of 2026-06-02. Click a release name to + Top-level index across MVP1 → GA v1+ as of 2026-06-03. Click a release name to drill into the per-release dashboard. Theme labels sourced from tech-stack.md §"Canonical release matrix". See state.md for @@ -392,7 +392,7 @@

Releases

Three-Engine + Real Signals
-
13 / 24 scoped done · 24 remaining
+
13 / 24 scoped done · 25 remaining
In progress
diff --git a/docs/00_overview/mvp1_dashboard.html b/docs/00_overview/mvp1_dashboard.html index bd5b4bd6..9a168d7d 100644 --- a/docs/00_overview/mvp1_dashboard.html +++ b/docs/00_overview/mvp1_dashboard.html @@ -369,7 +369,7 @@

RelyLoop MVP1 Dashboard

- Reflects feature-folder state as of 2026-06-02 (latest mtime of any + Reflects feature-folder state as of 2026-06-03 (latest mtime of any docs/00_overview/planned_features/ or docs/00_overview/implemented_features/ file). See state.md for the active branch context, diff --git a/docs/00_overview/mvp2_dashboard.html b/docs/00_overview/mvp2_dashboard.html index c3c87ac4..e9ddfb51 100644 --- a/docs/00_overview/mvp2_dashboard.html +++ b/docs/00_overview/mvp2_dashboard.html @@ -369,7 +369,7 @@

RelyLoop MVP2 Dashboard

- Reflects feature-folder state as of 2026-06-02 (latest mtime of any + Reflects feature-folder state as of 2026-06-03 (latest mtime of any docs/00_overview/planned_features/ or docs/00_overview/implemented_features/ file). See state.md for the active branch context, @@ -398,12 +398,12 @@

MVP2 Progress

Specced features done
13 / 24
-
54% specced · 43 filed under MVP2
+
54% specced · 44 filed under MVP2
Pending work
-
28
+
29
every not-done feat/infra/chore/bug across all priorities
@@ -425,7 +425,7 @@

MVP2 Progress

P2 (default)
-
24
+
25
important to file, not blocking
@@ -435,7 +435,7 @@

MVP2 Progress

Legacy "Path to MVP2"
-
24
+
25
scoped not-done + bugs + chore-ideas only (excludes feat/infra ideas)
@@ -463,7 +463,7 @@

Pipeline

-

Idea 15

+

Idea 16

@@ -504,6 +504,19 @@

Idea 15

+
+ +
+ Chore + P2 + +
+
`.github/workflows/pr.yml` has a job named `backend (lint + typecheck + tests + coverage)` that runs four sequential things in one job: ruff/lint, mypy, the full pytest matrix (unit + integration + co
+ + +
+ +
diff --git a/docs/00_overview/planned_features/02_mvp2/chore_pr_yml_parallelize_backend_job/idea.md b/docs/00_overview/planned_features/02_mvp2/chore_pr_yml_parallelize_backend_job/idea.md new file mode 100644 index 00000000..3c134192 --- /dev/null +++ b/docs/00_overview/planned_features/02_mvp2/chore_pr_yml_parallelize_backend_job/idea.md @@ -0,0 +1,82 @@ +# chore_pr_yml_parallelize_backend_job — Split the 8m20s backend job into parallel lanes + +**Date:** 2026-06-02 +**Status:** Idea — captured during PR #426 CI watch +**Priority:** P2 — operator iteration cost, not a correctness gate +**Origin:** PR #426 CI watch. Operator noticed the `backend (lint + typecheck + tests + coverage)` job ran for **8m20s** while the rest of the suite finished in 2-3 min. Operator asked: "is it possible to run this quicker? Can we parallelize this?" The answer is yes — three orthogonal wins, all of which fit the existing `pr.yml` shape without architectural change. +**Depends on:** None. + +## Problem + +`.github/workflows/pr.yml` has a job named `backend (lint + typecheck + tests + coverage)` that runs four sequential things in one job: ruff/lint, mypy, the full pytest matrix (unit + integration + contract), and the coverage gate. It dominates the critical path of every PR check at ~8m20s. Meanwhile there's a separate `backend (unit tests — fast lane)` job that runs the unit subset in 38s — but it's a duplicate of part of the heavy lane's work, not a parallelization. + +Three concrete operator costs: + +1. **Round-trip latency on lint slips.** A one-line ruff or mypy error costs the full ~8 min before the failure surfaces. The fast-lane job catches unit-test failures faster but doesn't run lint/typecheck. +2. **Critical path blocks merge.** With the heavy job at 8m, the smoke job at 0s (skipped opt-in), and both docker buildxes at ~2.5 min, the wall-clock per PR is `max(2.5, 2.5, 8) ≈ 8 min`. Reducing the backend lane to ~3-4 min would let the docker buildxes dominate. +3. **Service-container boot waste.** The heavy job boots Postgres + Elasticsearch + OpenSearch service containers even for the lint/typecheck steps, which need none of them. + +## Proposed capabilities + +Three wins, ordered by effort/reward: + +### Win 1 — Split lint + typecheck into their own job (~30-40s) + +A new `backend-static-checks` job that runs `make lint && make typecheck` against a checkout-only container (no Postgres / ES / OpenSearch service containers, no `uv sync` for full pytest deps). The existing `backend (unit tests — fast lane)` already does ~38s with the dev-deps-only install; the lint+typecheck job can be even leaner. + +Today there's a `static-checks-backend` job near the top of `pr.yml` (line ~10 in run logs) — verify what it covers and either extend it (preferred — fewer jobs to manage) or add a sibling. If it already covers lint + typecheck, the heavy `backend (lint + typecheck + ...)` job can drop those steps entirely. + +**Effort:** ~30 min of YAML editing. **Expected wall-clock cut:** 3-4 min off the critical path (lint/typecheck failures surface in <1 min instead of ~5 min, and the heavy lane stops paying for the redundant work). + +### Win 2 — Use `pytest -n auto` on unit + contract layers + +Both `backend/tests/unit/` and `backend/tests/contract/` are hermetic (no DB, no Compose). On a 2-core GHA runner, `pytest -n auto` halves wall-clock for embarrassingly parallel suites. Local timing: `pytest backend/tests/unit/` at 2191 tests runs in ~8s already; on CI's slower runner it's currently 3-4 min serial. With `-n auto` it's likely <90s. + +**Effort:** add `pytest-xdist` to the dev deps (it's already widely used in the ecosystem; check `uv.lock`), update `make test-unit` and `make test-contract` to pass `-n auto`, verify no test-ordering assumptions break (`@pytest.mark.order` or `pytest-ordering` if any test relies on specific ordering). Per CLAUDE.md "Common Pitfalls" the codebase already gates flaky tests behind `pytest-randomly` randomization — that's a sibling concern, not a blocker. **Expected wall-clock cut:** 2-3 min off the unit + contract portion of the heavy job. + +### Win 3 — Split integration tests by service-container they need (~2-3 hours, deserves its own spec) + +Today all integration tests share one job that boots Postgres + Elasticsearch + OpenSearch. But: +- Most tests need only Postgres (the demo-data + repo-layer suites). +- A subset needs Elasticsearch (the adapter integration tests). +- A subset needs OpenSearch (also adapter integration tests). +- A tiny subset needs Solr (when reachable; mostly skip-gated today). + +Splitting into `integration:postgres` / `integration:elastic` / `integration:opensearch` (using pytest markers + targeted `-m` selection) lets each lane scale to what it actually needs and run concurrently. The Postgres-only lane gets a faster service-container boot since it doesn't wait for ES/OS to settle. + +**Effort:** larger — needs pytest marker pass + per-lane service-container config + makefile target split + verification that test ordering still works under independent runs. **Recommend escalating Win 3 to a separate `infra_pr_yml_split_integration_by_service` spec via `/pipeline` if pursued.** Not in scope for this chore. + +## Decisions (locked at idea time) + +- **D-1. Wins 1 + 2 only in this chore; Win 3 deferred to a separate `infra_` spec.** Rationale: Wins 1 + 2 are ~1 hour of YAML/Makefile editing each, fit cleanly under the `chore_` prefix, and capture ~5-6 min of the 8m20s. Win 3 is multi-file with test-ordering implications across the full integration suite — that's `infra_`-shaped work warranting `/pipeline` ceremony (a spec + plan, not a `bug_fix.md`). +- **D-2. NO change to the coverage gate.** Coverage runs last as an aggregation step that needs the full pytest output. Splitting it naively (e.g., coverage per shard with merge) is its own rabbit hole. Keep the coverage step in whatever job ends up running the full pytest matrix; only the lint/typecheck split affects it (those don't contribute to coverage). +- **D-3. NO change to the fast-lane job.** The 38s fast-lane stays as-is — it's the canary for unit-test correctness. If Win 1 lands and lint/typecheck split out, the fast-lane becomes literally the unit-test subset; if Win 2 adds `-n auto` to it too, fast-lane drops to ~10s. + +## Open questions for spec/impl-execute + +- **Confirm `static-checks-backend` already exists and what it covers.** If it covers ruff/format and mypy already, then Win 1 reduces to dropping the redundant steps from the heavy job (a 5-line YAML edit). If it covers only one of those, the chore extends it. Worth a 5-minute audit of `pr.yml` before starting. +- **`pytest-xdist` test-isolation audit.** A small subset of tests may rely on shared fixture state (rare in this codebase but worth a one-pass grep for `@pytest.mark.serial`, module-level `_GLOBAL = ...` patterns, or `conftest.py` autouse fixtures with side effects). If any are found, the chore either marks them `@pytest.mark.serial` (xdist supports a single-worker serial group) or rewrites them — depends on count. + +## Scope signals + +- **Backend:** none (no `backend/app/` source touched). +- **Frontend:** none. +- **Migration:** none. +- **Config:** `.github/workflows/pr.yml` (job split), `Makefile` (test target update), `pyproject.toml` (add `pytest-xdist` to dev deps if not already there). +- **Audit events:** N/A — CI workflow config. +- **Operator impact:** none on operator-path behavior. Affects CI feedback latency only. Operators flipping `SKIP_HEAVY_CI=true` see no change. + +## Relationship to other work + +- Sibling of `infra_smoke_reseed_runtime_budget` (shipped 2026-06-02 in PR #424) — that work made the `smoke` job feasible to opt-in via `SMOKE_TEST=true`. This chore reduces the *default* per-PR critical path (smoke is opt-in/off), so even when smoke runs it stops being the wall-clock bottleneck. +- Coordinates with — but does NOT block — Win 3 (`infra_pr_yml_split_integration_by_service`). Wins 1 + 2 are pure subtractions from the heavy job; Win 3 is a structural split. Either can ship first. +- Captured during `bug_llm_capability_cache_no_refresh` (PR #426) CI watch — the slow backend job made the operator wait through two cycles (initial push + post-Gemini-fixes push); each waited the same 8 min on the same backend lane. + +## Why filed instead of fixed inline during PR #426 + +Per CLAUDE.md "Tangential discoveries — fix inline by default": tested the inline-fix gate. This work is: +- **Cross-subsystem from the bug fix** (CI/workflow vs. LLM capability cache — different surfaces, no shared file). +- **>60 min of work** even for Wins 1 + 2 (need to audit existing `static-checks-backend`, verify `pytest-xdist` is in deps, test-isolation pass). +- **Would expand PR #426's review scope** from "Redis cache helper" to "Redis cache helper + CI parallelization" — confusing for the reviewer. + +So the rubric rows for "cross-subsystem + >60 min" + "expands PR review scope" both say defer. Filed as a separate chore for parallel execution.