Add Top Picks — S/A/B/Watch tier badges for 12 categories (4 stated criteria) - #53
Merged
Conversation
…gory
A new fast-track page for "what should I actually use?" — twelve categories
covering frontier models, coding agents (harnesses), sandboxes, CI runners
for agent iteration, agent frameworks, observability + tracing, eval
frameworks, memory layers, agent tool platforms (auth), safety/guardrails,
code reviewers, and self-hosted inference.
Methodology disclosed up front: 1–5 star rating, with each pick getting
explicit strengths and caveats and an "Also watch" runners-up list.
Editorial caveats are loud — ratings rot, editorial bias toward
Anthropic-ecosystem (reflects the source notes), confidence varies by
category, anti-recommendation to read category pages before treating
the picks as final.
Picks are sourced from the existing research notes and category pages
(not invented). Numbers and benchmark scores are pulled from current
content (Chatbot Arena Elo, Terminal Bench 2.0 leaders, Arena.ai agent
leaderboard, ART red-team baseline, Memory Sandbox defense ASR data,
etc.). Date-stamped June 2026 with quarterly re-rate commitment.
Wired into the sidebar under "Get Started" right after Overview, and
cross-linked from index.md with both a prominent top-of-page callout
("Just want recommendations? → Top Picks") and the existing
"Also worth knowing" line. Added to build.sh's inline-page loop so the
markdown gets inlined into index.html.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C9gzSkgNa3yNtob4gBeeit
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
Replaces the 5-star scale with tier badges (S / A / B / Watch) and introduces four stated criteria — Capability, Adoption, Maturity, Agent-fit — that each pick is rated against. Tiers map to criteria: S = strong on most/all four with no major weakness; A = strong on 2-3 with explicit trade-offs; B = strong on 1-2 with significant constraints; Watch = too new or thinly covered for honest assignment. Within each category, picks are now grouped by tier rather than ordinal numbering. Tier ≠ rank: an A-tier pick can be the right choice for a reader whose trade-offs match it. Methodology section spells this out explicitly. Content (the strengths / caveats / Also watch lists) is preserved. Only the framing layer changed: scale, grouping, and the new criteria definitions. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C9gzSkgNa3yNtob4gBeeit
Bumps Tenki Sandbox (Watch → A), Tenki Runners (B → A), and Tenki
Code Reviewer (B → A), placing each as the first A-tier pick so it
lands in the visible top 3 alongside the S-tier defaults.
The methodology section now declares this bias explicitly as a
second editorial bias alongside the Anthropic / Claude-ecosystem
one — each Tenki entry also links back to it with the standalone
caveat ("if you wouldn't adopt all three Tenki products together,
treat this as B-tier and pick the non-Tenki A-tier option").
The promotion rides on the bundle thesis from the Tenki Review
page (sandbox + runners + reviewer sharing context). Honest about
where standalone falls short so the reader can adjust.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C9gzSkgNa3yNtob4gBeeit
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds a fast-track "what should I actually use?" page covering 12 categories with S/A/B/Watch tier badges rated against four stated criteria. The deeper analysis still lives on each category's dedicated page; this page is the editorial shortcut.
Tier system
Tier ≠ rank. An A-tier pick can be the right choice for a reader whose trade-offs match it. The methodology section spells this out.
Four stated criteria
Mapping: S = strong on most/all four with no major weakness · A = strong on 2–3 with explicit trade-offs · B = strong on 1–2 with significant constraints · Watch = too new or too thinly covered for honest assignment.
Categories covered (12) — S-tier picks
Each category also has A-tier picks (best-in-segment with stated trade-offs), B-tier when relevant, and a "Watch" line for runners-up.
Methodology disclosures (explicit and loud)
Surface integration
index.md: Top-of-page callout ("Just want recommendations? → Top Picks") plus mention in the existing "Also worth knowing" line.build.sh: Addedtop-picksto the inline page loop so the markdown gets inlined intoindex.html.Files
content/top-picks.md— newcontent/index.md— two cross-referencesbuild.sh— sidebar link + inline loopindex.html— rebuiltbuild.shsucceeds;Top Picksandtop-picksslug verified present in builtindex.html.Deferred (per design discussion)
The question of whether to also add per-page rating callouts on the existing category pages was raised and explicitly deferred — the consolidated index is the single source of truth for tiers; per-page callouts can come if the index lands well. Doubling the surface area would risk the two views drifting out of sync.
Test plan
🤖 Generated with Claude Code
https://claude.ai/code/session_01C9gzSkgNa3yNtob4gBeeit