You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Env validation — scripts/validate_env.py checks vars, Ollama, Redis, data dirs
Roast Review V6 — Critical Fixes
Eval pipeline honest — EvalPipeline now uses LLMJudge when available, heuristic fallback. Reports judge_method in output. No more fake word-overlap CI gating
pickle → joblib + SHA-256 — Classifier and BM25 index use joblib.load/dump. Classifier saves .pkl.sha256 hash file, verifies integrity before loading
API async — All route handlers async def, orchestrator.run() wrapped in asyncio.to_thread() — non-blocking event loop under load
Clean repo — roast_review_fixed_V2.md untracked (stays local only)
2026 Production Features — Phase A: Retrieval Intelligence
Contextual Embeddings — Prepend 50-100 token context to each chunk before embedding (49% retrieval improvement). Uses ContextualEmbedder in ingestion pipeline. Updated vector_store.py and bm25_index.py to use contextual text when available
Skip Retrieval — Agent answers directly for known facts (what/who/when/where/define/explain) without retrieval. Saves 20-40% of queries from unnecessary retrieval
2026 Production Features — Phase B: Production Hardening
Rate Limiting — Sliding window per-tenant rate limiting (requests/minute, requests/hour, burst size). Configurable via API. Wired into all query endpoints
Structured Logging — JSON formatter for log aggregation (ELK/Datadog compatible). Request-scoped context (trace_id, tenant_id). Performance logging with duration_ms
Online Eval — Sample 5% of production queries for LLM-as-Judge evaluation. Daily stats, model distribution tracking. Wired into query endpoint
2026 Production Features — Phase C: Agent Intelligence
MCP Server — Exposes search, get_financials, compare_companies, list_companies as MCP-compatible tools with JSON schema. 2026 standard for tool connectivity
Multi-agent Architecture — ResearchAgent → AnalysisAgent → VerificationAgent pipeline. Specialized sub-agents with _extract_tickers helper. Accessible via /query/multi-agent endpoint
API Endpoints (50+)
Core Query
POST /query Sync query (full pipeline)
POST /query/stream SSE streaming response
POST /query/structured Structured output (Pydantic schema)
POST /query/multi-agent Multi-agent pipeline (Research→Analysis→Verification)
Ingestion & Upload
POST /upload Upload PDF for indexing
GET /upload/{doc_id}/status Check upload status
DELETE /upload/{doc_id} Delete document
GET /uploads List all uploads
Documents & Conversation
GET /documents List indexed documents
GET /conversation/history Chat history
POST /conversation/session Create/switch session
GET /conversation/sessions List sessions
DELETE /conversation/session/{id} Delete session
Feedback & Suggestions
POST /feedback Submit feedback
GET /feedback/stats Feedback statistics
GET /suggestions Query suggestions
GET /suggestions/related Related queries
Cost & Analytics
GET /cost/summary Cost summary
GET /cost/budget Budget check
GET /cost/savings Cost savings from routing
GET /cost/breakdown Cost by category/company
GET /cost/token-budgets Token usage by model
GET /analytics/models Model comparison
GET /analytics/routing Routing breakdown
GET /analytics/trend Cost trend
GET /analytics/tokens Token efficiency
GET /analytics/anomalies Anomaly detection
Latency
GET /latency/stats p50/p95/p99 per component
GET /latency/recent Recent latency entries
GET /latency/trend Time-bucketed averages
Cache
GET /cache/stats Semantic cache hit rates
Prompts
GET /prompts List all prompts + versions
GET /prompts/{name} Get prompt details
POST /prompts/{name}/rollback Rollback to version
POST /prompts/{name}/regression Run regression test
Tenants
POST /tenants Create tenant
GET /tenants List tenants
GET /tenants/{id} Get tenant details
PUT /tenants/{id} Update tenant
DELETE /tenants/{id} Delete tenant
GET /tenants/{id}/usage Usage stats
Export
GET /export/query Export query to PDF
GET /export/analytics Export analytics to PDF
GET /export/queries/csv Export queries to CSV
GET /export/list List exports
Knowledge Graph & Eval
GET /knowledge/stats Graph statistics
GET /knowledge/entity/{id} Query entity
POST /knowledge/extract Extract from text
POST /eval/run Run evaluation
GET /eval/history Eval history
GET /eval/averages Average scores
Compare Filings
POST /compare/years Compare company across years
POST /compare/companies Compare companies for a year
Admin (auth required)
POST /admin/login Login
POST /admin/logout Logout
GET /admin/users List users
POST /admin/users Create user
DELETE /admin/users/{id} Delete user
GET /admin/validate Validate session
POST /admin/clear-cache Clear cache
Health
GET /health System health
GET /health/metrics Health metrics
Failure Analysis
GET /failures Failure analysis from latest eval
GET /failures/trends Failure trends across runs
MCP Tools
GET /mcp/tools List MCP tool schemas
POST /mcp/call Call MCP tool by name
Human Escalation
GET /escalations List escalation tickets
GET /escalations/{id} Get specific ticket
POST /escalations/{id}/resolve Resolve ticket with notes
GET /escalations/stats Escalation statistics
Online Evaluation
GET /online-eval/stats Online eval statistics
POST /online-eval/evaluate Manually evaluate a query/answer
GET /online-eval/results Recent eval results
Rate Limiting
GET /rate-limits/stats Rate limit stats per tenant
POST /rate-limits/config Update rate limit config
Frontend Pages (10)
/ (Dashboard) Chat with streaming + feedback + suggestions
/app/upload Upload PDF documents
/app/documents SEC filing browser
/app/analytics Charts and metrics
/app/comparison Model comparison dashboard
/app/latency Latency dashboard (p50/p95/p99)
/app/cost-optimization Cost optimization dashboard
/app/failures Failure analysis dashboard
/app/admin Admin panel (auth, users, tenants, cache)
/ Landing page with quick links