Skip to content

fix(app): stop idle polling from keeping the SQL warehouse running - #1569

Open
OGordon100 wants to merge 3 commits into
mainfrom
dqx-studio/warehouse-idle-polling
Open

OGordon100 wants to merge 3 commits into
mainfrom
dqx-studio/warehouse-idle-polling

Conversation

@OGordon100

Copy link
Copy Markdown
Contributor

Summary

Two background polls kept the DQX Studio SQL warehouse awake even when no one was using it. The warehouse is billed for uptime and auto-stops after 10 idle minutes, but both polls ran more often than that, so it never stopped. Both now read run state from the Jobs API instead, so they don't use the warehouse.

Stacked on #1564 (fix-run-config-table). Retarget to main once that merges.

1. App-wide failure toast (/dryrun/runs/recent-failures, /profiler/runs/recent-failures)

Every open tab polled these endpoints. Each call ran a reconcile plus a SELECT against the Delta run tables.

  • A new module, services/task_runner_runs.py, reads the most recent completed runs of the task-runner job from jobs.list_runs. The look-back is a fixed number of runs (100) rather than a time window, so a tab that was hidden still catches up when it comes back. The app run id, task type and source_table_fqn are all taken from the job parameters (the run_id, task_type and config_json parameters).
  • Results are cached in the app cache for 30s, so all tabs and users share a single Jobs API call.
  • Previews (skip_history / run_type=preview) are excluded. A run counts as failed on the same rule as reconcile: it has reached a terminal state with a result other than SUCCESS or CANCELED.
  • The validation feed still applies the caller's catalog filter.
  • One SQL fallback is left: when a run's config was too large and was staged out of the job parameters, its table is looked up in the run table. This happens only for those runs.

2. Scheduler's pending score-run poll

The scheduler queried Delta every tick to see whether pending score runs had finished.

  • It now asks the Jobs API which runs are still active. It queries Delta only for runs that have finished: runs it has seen active that are no longer active, or runs older than a 2-minute grace period.
  • If a finished run still has no results row after its second lookup, it is dropped, so a run that failed no longer has its results polled for indefinitely.
  • If job_id isn't set or the Jobs API call fails, it falls back to the previous behaviour (query all pending runs).

Test plan

  • make app-test: 4,530 passed. New tests are in tests/test_task_runner_runs.py; tests/test_recent_failures.py was rewritten; new cases were added to tests/test_job_service.py and tests/test_scheduler_service.py.
  • basedpyright: 0 errors. UI unit tests: 938 passed.
  • Deployed to a test workspace. The app started cleanly, the scheduler started, and the startup score refresh completed.
  • With the app idle and a tab open, confirm the SQL warehouse auto-stops after its idle timeout.
  • Trigger a failed run and confirm the failure toast still appears.

This pull request and its description were written by Isaac.

@OGordon100
OGordon100 requested a review from a team as a code owner October 2, 2026 14:56
@OGordon100
OGordon100 requested review from gergo-databricks and removed request for a team October 2, 2026 14:56
@OGordon100 OGordon100 added the DQX App Feature/Bug related to the DQX App label Oct 2, 2026
@codecov

codecov Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 67.84%. Comparing base (4c6c017) to head (72c289a).

❗ There is a different number of reports uploaded between BASE (4c6c017) and HEAD (72c289a). Click for more details.

HEAD has 4 uploads less than BASE
Flag BASE (4c6c017) HEAD (72c289a)
anomaly 1 0
anomaly-serverless 1 0
integration 1 0
integration-serverless 1 0
Additional details and impacted files
@@             Coverage Diff             @@
##             main    #1569       +/-   ##
===========================================
- Coverage   93.04%   67.84%   -25.20%     
===========================================
  Files         142      142               
  Lines       14210    14210               
  Branches      151      151               
===========================================
- Hits        13221     9641     -3580     
- Misses        920     4499     +3579     
- Partials       69       70        +1     
Flag Coverage Δ
anomaly ?
anomaly-serverless ?
integration ?
integration-serverless ?
mcp 80.44% <ø> (+0.09%) ⬆️
unit 66.85% <ø> (ø)

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actions

github-actions Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

✅ 24/24 passed, 4 skipped, 2h2m33s total

Running from acceptance #6136

@github-actions

github-actions Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

✅ 1/1 passed, 25m25s total

Running from mcp #885

@github-actions

github-actions Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

✅ 195/195 passed, 1 skipped, 7h36m1s total

Running from anomaly #2244

@mwojtyczka

Copy link
Copy Markdown
Contributor

Code review findings

A review pass over this PR. The core approach is sound — the task runner writes a FAILED row and then sys.exit(1) (app/tasks/src/dqx_task_runner/runner.py:1527-1547), so the Jobs-API result_state=FAILED reliably corresponds to an app-level FAILED validation, and substituting the Jobs API for the Delta read is a valid way to keep the warehouse idle. The findings below are edge cases plus one partial regression of the stated goal, ranked by impact. These are automated-review findings, not yet empirically verified — please treat them as items to confirm.

🟠 Staged-failure source-table lookup is uncached — partially defeats the "keep the warehouse idle" goal

app/src/databricks_labs_dqx_app/backend/routes/v1/dryrun.py:258

For a run whose config was staged out of job params (oversized rule set), run.source_table_fqn is None, so staged is non-empty and job_svc.lookup_source_tables() runs a SELECT … FROM dq_validation_runs (and dq_profiling_results for the profiler endpoint) on every request, uncached — unlike list_recent_failed_runs, which has a 30s cache. A failed staged run stays in the recent-100 completed-run window, so N tabs polling every 60s keep issuing warehouse SELECTs, which is the warehouse-wake behavior this PR set out to eliminate. Consider caching this lookup (same TTL as the recent-failures path) or resolving the source table without a warehouse round-trip.

🟠 Whole feature depends on list_runs() returning job_parameters — and that path is untested

app/src/databricks_labs_dqx_app/backend/services/task_runner_runs.py:68

parse_task_runner_run derives app_run_id solely from run.job_parameters, and list_recent_completed_runs / list_active_app_run_ids drop any run that parses to None. If the Jobs API list endpoint omits or empties job_parameters in list responses (as it does for some fields like tasks / cluster_spec without expand_tasks), every run is discarded: failure toasts never appear and the scheduler sees zero active runs (treating running jobs as finished). The PR's test plan leaves "Trigger a failed run and confirm the failure toast still appears" unchecked, so this end-to-end path was not verified against a live workspace. Worth confirming job_parameters is reliably populated in list responses (and adding the live check), or deriving app_run_id from a field guaranteed in list payloads.

🟡 Completed run dropped from score-refresh tracking on Delta visibility lag

app/src/databricks_labs_dqx_app/backend/services/scheduler_service.py:1429

When a finished run's terminal dq_validation_runs row isn't visible to the query for two consecutive ~60s ticks (Delta commit/visibility lag, or a lingering RUNNING placeholder not yet overwritten), the run is _forget_score_run'd and its cached quality score is never refreshed — a silently stale score for a run that actually completed successfully. Consider a longer grace window or a bounded retry before forgetting.

🟡 is_failed toasts SKIPPED / null-result terminal runs

app/src/databricks_labs_dqx_app/backend/services/task_runner_runs.py:56

is_failed treats any terminal lifecycle state with result_state ∉ {SUCCESS, CANCELED} as a failure, so a SKIPPED run (or TERMINATED with result_state=None) yields a spurious "run failed" error toast for a run the user never experienced as a failure. Consider matching result_state == FAILED explicitly rather than treating "not success/cancelled" as failure.

🟡 Staged preview runs escape preview-exclusion and can toast

app/src/databricks_labs_dqx_app/backend/services/task_runner_runs.py:87

is_preview reads skip_history / run_type from config_json, but for oversized (staged) configs that field holds only the manifest pointer, not those keys — so a failed staged preview evaluates is_preview == False, passes the not is_preview filter, and appears in the failure toast feed even though previews are meant to be excluded (no run-history row exists for it). Consider carrying the preview marker in a field that survives staging (e.g. a job parameter), not only in the staged config body.

@OGordon100

Copy link
Copy Markdown
Contributor Author

Fixed in 1802564:

  • Uncached staged lookup / staged previews: the staged-config stub now includes source_table_fqn, skip_history and run_type. The failure feed gets the table name and the preview flag from the Jobs API, lookup_source_tables is gone, and the recent-failures endpoints no longer query the warehouse.

Not changing:

  • job_parameters in list_runs(): checked on a live workspace. runs/list returns the full job_parameters (including run_id and config_json) for every run without needing expand_tasks.
  • is_failed on SKIPPED or null results: this matches reconcile_running_rows, which marks any terminal run that isn't SUCCESS or CANCELED as FAILED in Runs History. Narrowing it to FAILED would hide toasts for runs that history shows as failed.
  • Dropping score tracking on lag: the runner writes its final row before it exits, so the row is committed by the time the Jobs API reports the run finished. Two misses in a row means the runner crashed and reconcile will mark the run FAILED, so there's no score to refresh. Re-querying for 24h is the warehouse wake-up this PR removes.

This pull request and its description were written by Isaac.

OGordon100 and others added 2 commits October 8, 2026 16:34
Two pollers queried the warehouse often enough that its 10-minute
auto-stop never fired, even with nobody using the app:

- The app-wide failure toast polled /dryrun and /profiler
  recent-failures every 60s per open tab, each reading the Delta run
  tables. Both endpoints now read failed task-runner job runs from the
  Jobs API (most recent 100 completed runs, count-bounded so hidden tabs
  still catch up), shared across requests by a 30s cache. The source
  table comes from the run's config_json parameter; only runs whose
  config was staged out of the job parameters fall back to a run-table
  lookup. Response contract unchanged.

- The scheduler queried dq_validation_runs every 60s for each tracked
  run until a terminal row appeared, so a run that never wrote one kept
  the warehouse up for the 24h TTL. It now asks the Jobs API which
  tracked runs are still active and queries the warehouse only for runs
  whose job has finished, dropping a finished run with no terminal row
  after a second empty lookup. A Jobs API failure falls back to the
  previous behaviour.

New services/task_runner_runs.py holds the shared Jobs API reads.
api.ts regenerated (doc comments only for these endpoints; also picks up
ProfilerConfig geospatial fields already present in the backend).

Co-authored-by: Isaac <no-reply@databricks.com>
The staged-config stub now includes source_table_fqn, skip_history and
run_type, so the Jobs API failure feed can name a staged run's table and
exclude staged previews without a SQL lookup. This removes the last
warehouse query from the recent-failures endpoints.

Co-authored-by: Isaac <no-reply@databricks.com>

@ghanse ghanse left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A few scoped follow-ups on the Jobs-API switch — all about new code in this PR, not pre-existing behavior.

The scheduler matches swept run-set members against the Jobs API's active
runs by the run_id job parameter. Pin that BindingRunService records the
same id on the member as it submits to the job, so a divergence can't
silently send still-running members to a Delta lookup every tick.

Co-authored-by: Isaac <no-reply@databricks.com>

This branch is being deployed

1 in progress deployment
tool — 72c289aa Deployed Oct 9, 2026 by OGordon100 via e2e_serverless #6136
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Approved to Merge When PR is reviewed and approved. To be merged once all tests pass DQX App Feature/Bug related to the DQX App

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants