Repository navigation
Commit 187ac14
Fix ODCS type: library metric review findings
Addresses the review comments on PR databrickslabs#1485:
- duplicateValues: replace the COUNT(*) OVER (...) window-function
indicator (nested inside SUM/AVG, which Spark rejects at apply time
for every non-mustBe:0 threshold) with a GROUP BY-based duplicate
count computed via the sql_query fallback.
- nullValues percent and missingValues forbidden list: stop embedding
live PySpark Column objects in generated rule dicts (broke
save_checks() serialization and silently disabled
ChecksSemanticValidator conflict detection on an unhashable Column).
- invalidValues: escape backslashes in RLIKE/IN literals and leave
numeric validValues unquoted so the aggregate path matches the
row-level is_in_list/regex_match path.
- Normalize mustBe zero-threshold detection so it only matches a
genuine numeric zero, not boolean False or the string "0".
- Disambiguate rowCount rule names when a schema carries more than one
rowCount entry.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>1 parent d2b6471 commit 187ac14
4 files changed
Lines changed: 301 additions & 119 deletions
File tree
- docs/dqx/docs/guide
- src/databricks/labs/dqx/datacontract
- tests
- integration
- unit
Lines changed: 3 additions & 3 deletions
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
577 | 577 | | |
578 | 578 | | |
579 | 579 | | |
580 | | - | |
| 580 | + | |
581 | 581 | | |
| 582 | + | |
582 | 583 | | |
583 | 584 | | |
584 | 585 | | |
| |||
724 | 725 | | |
725 | 726 | | |
726 | 727 | | |
727 | | - | |
728 | | - | |
| 728 | + | |
729 | 729 | | |
730 | 730 | | |
731 | 731 | | |
| |||
0 commit comments