You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit fab9c3a
Browse filesBrowse the repository at this point in the historyBrowse files
Addresses the review comments on PR #1485:
- duplicateValues: replace the COUNT(*) OVER (...) window-function
indicator (nested inside SUM/AVG, which Spark rejects at apply time
for every non-mustBe:0 threshold) with a GROUP BY-based duplicate
count computed via the sql_query fallback.
- nullValues percent and missingValues forbidden list: stop embedding
live PySpark Column objects in generated rule dicts (broke
save_checks() serialization and silently disabled
ChecksSemanticValidator conflict detection on an unhashable Column).
- invalidValues: escape backslashes in RLIKE/IN literals and leave
numeric validValues unquoted so the aggregate path matches the
row-level is_in_list/regex_match path.
- Normalize mustBe zero-threshold detection so it only matches a
genuine numeric zero, not boolean False or the string "0".
- Disambiguate rowCount rule names when a schema carries more than one
rowCount entry.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Copy file name to clipboardExpand all lines: docs/dqx/docs/guide/data_contract_quality_rules_generation.mdx
+3-3Lines changed: 3 additions & 3 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -577,8 +577,9 @@ Every metric shares the same **threshold fields** (one of which must be set) and
577
577
Every library metric is evaluated against one of eight ODCS threshold fields, tried in this order — the first one present on the entry is used: `mustBe`, `mustNotBe`, `mustBeGreaterOrEqualTo`, `mustBeLessOrEqualTo`, `mustBeGreaterThan`, `mustBeLessThan`, `mustBeBetween`, `mustNotBeBetween`.
578
578
579
579
- **`mustBe: 0`** is special-cased for `nullValues`, `missingValues`, `invalidValues`, and `duplicateValues`: it maps onto a cheap **row-level** check (`is_not_null`, `is_not_in_list`, `is_in_list`/`regex_match`, or `is_unique`) that pinpoints the offending rows, rather than a dataset-level count.
580
-
- **`mustBe`, `mustNotBe`, `mustBeGreaterOrEqualTo`, `mustBeLessOrEqualTo`** (including `mustBe` with a non-zero value) map onto exact-fit **dataset-level aggregate checks** — `is_aggr_equal`, `is_aggr_not_equal`, `is_aggr_not_less_than`, `is_aggr_not_greater_than` — over the metric's count or percentage.
580
+
- **`mustBe`, `mustNotBe`, `mustBeGreaterOrEqualTo`, `mustBeLessOrEqualTo`** (including `mustBe` with a non-zero value) map onto exact-fit **dataset-level aggregate checks** — `is_aggr_equal`, `is_aggr_not_equal`, `is_aggr_not_less_than`, `is_aggr_not_greater_than` — over the metric's count or percentage, for `rowCount`, `nullValues`, `missingValues`, and `invalidValues`.
581
581
- **`mustBeGreaterThan`, `mustBeLessThan`, `mustBeBetween`, `mustNotBeBetween`** have no strict/exclusive-bound equivalent among DQX's aggregate checks, so they fall back to a dataset-level [`sql_query`](/docs/reference/quality_checks#using-sql-query) check with `condition_column: "condition"` (`true` means a violation). For `mustBeBetween`/`mustNotBeBetween`, **both bounds are exclusive**, per the ODCS specification — a value exactly equal to either bound does not count as being "between" them.
582
+
- **`duplicateValues` is the one exception**: every non-`mustBe: 0` threshold (not just the four strict/exclusive-bound ones above) falls back to `sql_query`. The duplicate count/percentage is computed via a `GROUP BY` subquery rather than a plain aggregate expression, since Spark rejects a window function (the `PARTITION BY` used to detect duplicates) nested inside an aggregate function (`SUM`/`AVG`).
An entry with none of the eight fields set is skipped with a warning naming all eight, so a contract author can spot a typo (e.g. `mustbe` instead of `mustBe`) without reading DQX source.
@@ -724,8 +725,7 @@ The single-column, argument-less form is a property-level entry; the composite-k
724
725
unit: percent
725
726
arguments:
726
727
properties: [order_id, line_number]
727
-
# → is_aggr_not_greater_than over the duplicate-row-percentage indicator,
728
-
# partitioned by (order_id, line_number)
728
+
# → sql_query over the duplicate-row-percentage, grouped by (order_id, line_number)
0 commit comments