Repository navigation
fix(app): stop idle polling from keeping the SQL warehouse running - #1569
OGordon100 wants to merge 3 commits into
Conversation
Codecov Report✅ All modified and coverable lines are covered by tests.
Additional details and impacted files@@ Coverage Diff @@
## main #1569 +/- ##
===========================================
- Coverage 93.04% 67.84% -25.20%
===========================================
Files 142 142
Lines 14210 14210
Branches 151 151
===========================================
- Hits 13221 9641 -3580
- Misses 920 4499 +3579
- Partials 69 70 +1
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
✅ 24/24 passed, 4 skipped, 2h2m33s total Running from acceptance #6136 |
|
✅ 1/1 passed, 25m25s total Running from mcp #885 |
|
✅ 195/195 passed, 1 skipped, 7h36m1s total Running from anomaly #2244 |
Code review findingsA review pass over this PR. The core approach is sound — the task runner writes a FAILED row and then 🟠 Staged-failure source-table lookup is uncached — partially defeats the "keep the warehouse idle" goal
For a run whose config was staged out of job params (oversized rule set), 🟠 Whole feature depends on
|
|
Fixed in 1802564:
Not changing:
This pull request and its description were written by Isaac. |
4fb8953 to
ae3b518
Compare
Two pollers queried the warehouse often enough that its 10-minute auto-stop never fired, even with nobody using the app: - The app-wide failure toast polled /dryrun and /profiler recent-failures every 60s per open tab, each reading the Delta run tables. Both endpoints now read failed task-runner job runs from the Jobs API (most recent 100 completed runs, count-bounded so hidden tabs still catch up), shared across requests by a 30s cache. The source table comes from the run's config_json parameter; only runs whose config was staged out of the job parameters fall back to a run-table lookup. Response contract unchanged. - The scheduler queried dq_validation_runs every 60s for each tracked run until a terminal row appeared, so a run that never wrote one kept the warehouse up for the 24h TTL. It now asks the Jobs API which tracked runs are still active and queries the warehouse only for runs whose job has finished, dropping a finished run with no terminal row after a second empty lookup. A Jobs API failure falls back to the previous behaviour. New services/task_runner_runs.py holds the shared Jobs API reads. api.ts regenerated (doc comments only for these endpoints; also picks up ProfilerConfig geospatial fields already present in the backend). Co-authored-by: Isaac <no-reply@databricks.com>
The staged-config stub now includes source_table_fqn, skip_history and run_type, so the Jobs API failure feed can name a staged run's table and exclude staged previews without a SQL lookup. This removes the last warehouse query from the recent-failures endpoints. Co-authored-by: Isaac <no-reply@databricks.com>
ae3b518 to
eff3b01
Compare
ghanse
left a comment
There was a problem hiding this comment.
A few scoped follow-ups on the Jobs-API switch — all about new code in this PR, not pre-existing behavior.
The scheduler matches swept run-set members against the Jobs API's active runs by the run_id job parameter. Pin that BindingRunService records the same id on the member as it submits to the job, so a divergence can't silently send still-running members to a Delta lookup every tick. Co-authored-by: Isaac <no-reply@databricks.com>
Summary
Two background polls kept the DQX Studio SQL warehouse awake even when no one was using it. The warehouse is billed for uptime and auto-stops after 10 idle minutes, but both polls ran more often than that, so it never stopped. Both now read run state from the Jobs API instead, so they don't use the warehouse.
1. App-wide failure toast (
/dryrun/runs/recent-failures,/profiler/runs/recent-failures)Every open tab polled these endpoints. Each call ran a reconcile plus a SELECT against the Delta run tables.
services/task_runner_runs.py, reads the most recent completed runs of the task-runner job fromjobs.list_runs. The look-back is a fixed number of runs (100) rather than a time window, so a tab that was hidden still catches up when it comes back. The app run id, task type andsource_table_fqnare all taken from the job parameters (therun_id,task_typeandconfig_jsonparameters).skip_history/run_type=preview) are excluded. A run counts as failed on the same rule as reconcile: it has reached a terminal state with a result other than SUCCESS or CANCELED.2. Scheduler's pending score-run poll
The scheduler queried Delta every tick to see whether pending score runs had finished.
job_idisn't set or the Jobs API call fails, it falls back to the previous behaviour (query all pending runs).Test plan
make app-test: 4,530 passed. New tests are intests/test_task_runner_runs.py;tests/test_recent_failures.pywas rewritten; new cases were added totests/test_job_service.pyandtests/test_scheduler_service.py.This pull request and its description were written by Isaac.