Skip to content

Commit fab9c3a

Browse files
Fix ODCS type: library metric review findings
Addresses the review comments on PR #1485: - duplicateValues: replace the COUNT(*) OVER (...) window-function indicator (nested inside SUM/AVG, which Spark rejects at apply time for every non-mustBe:0 threshold) with a GROUP BY-based duplicate count computed via the sql_query fallback. - nullValues percent and missingValues forbidden list: stop embedding live PySpark Column objects in generated rule dicts (broke save_checks() serialization and silently disabled ChecksSemanticValidator conflict detection on an unhashable Column). - invalidValues: escape backslashes in RLIKE/IN literals and leave numeric validValues unquoted so the aggregate path matches the row-level is_in_list/regex_match path. - Normalize mustBe zero-threshold detection so it only matches a genuine numeric zero, not boolean False or the string "0". - Disambiguate rowCount rule names when a schema carries more than one rowCount entry. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
1 parent 974eea8 commit fab9c3a

4 files changed

Lines changed: 301 additions & 119 deletions

File tree

‎docs/dqx/docs/guide/data_contract_quality_rules_generation.mdx‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -577,8 +577,9 @@ Every metric shares the same **threshold fields** (one of which must be set) and
577577
Every library metric is evaluated against one of eight ODCS threshold fields, tried in this order — the first one present on the entry is used: `mustBe`, `mustNotBe`, `mustBeGreaterOrEqualTo`, `mustBeLessOrEqualTo`, `mustBeGreaterThan`, `mustBeLessThan`, `mustBeBetween`, `mustNotBeBetween`.
578578

579579
- **`mustBe: 0`** is special-cased for `nullValues`, `missingValues`, `invalidValues`, and `duplicateValues`: it maps onto a cheap **row-level** check (`is_not_null`, `is_not_in_list`, `is_in_list`/`regex_match`, or `is_unique`) that pinpoints the offending rows, rather than a dataset-level count.
580-
- **`mustBe`, `mustNotBe`, `mustBeGreaterOrEqualTo`, `mustBeLessOrEqualTo`** (including `mustBe` with a non-zero value) map onto exact-fit **dataset-level aggregate checks** — `is_aggr_equal`, `is_aggr_not_equal`, `is_aggr_not_less_than`, `is_aggr_not_greater_than` — over the metric's count or percentage.
580+
- **`mustBe`, `mustNotBe`, `mustBeGreaterOrEqualTo`, `mustBeLessOrEqualTo`** (including `mustBe` with a non-zero value) map onto exact-fit **dataset-level aggregate checks** — `is_aggr_equal`, `is_aggr_not_equal`, `is_aggr_not_less_than`, `is_aggr_not_greater_than` — over the metric's count or percentage, for `rowCount`, `nullValues`, `missingValues`, and `invalidValues`.
581581
- **`mustBeGreaterThan`, `mustBeLessThan`, `mustBeBetween`, `mustNotBeBetween`** have no strict/exclusive-bound equivalent among DQX's aggregate checks, so they fall back to a dataset-level [`sql_query`](/docs/reference/quality_checks#using-sql-query) check with `condition_column: "condition"` (`true` means a violation). For `mustBeBetween`/`mustNotBeBetween`, **both bounds are exclusive**, per the ODCS specification — a value exactly equal to either bound does not count as being "between" them.
582+
- **`duplicateValues` is the one exception**: every non-`mustBe: 0` threshold (not just the four strict/exclusive-bound ones above) falls back to `sql_query`. The duplicate count/percentage is computed via a `GROUP BY` subquery rather than a plain aggregate expression, since Spark rejects a window function (the `PARTITION BY` used to detect duplicates) nested inside an aggregate function (`SUM`/`AVG`).
582583

583584
<Admonition type="note" title="No recognized threshold field">
584585
An entry with none of the eight fields set is skipped with a warning naming all eight, so a contract author can spot a typo (e.g. `mustbe` instead of `mustBe`) without reading DQX source.
@@ -724,8 +725,7 @@ The single-column, argument-less form is a property-level entry; the composite-k
724725
unit: percent
725726
arguments:
726727
properties: [order_id, line_number]
727-
# → is_aggr_not_greater_than over the duplicate-row-percentage indicator,
728-
# partitioned by (order_id, line_number)
728+
# → sql_query over the duplicate-row-percentage, grouped by (order_id, line_number)
729729
```
730730
</TabItem>
731731
</Tabs>

0 commit comments

Comments
 (0)