Skip to content

Feat dataframe with special char cols - #1532

Draft
cornzyblack wants to merge 32 commits into
databrickslabs:mainfrom
cornzyblack:feat-dataframe-with-special-char-cols
Draft

cornzyblack wants to merge 32 commits into
databrickslabs:mainfrom
cornzyblack:feat-dataframe-with-special-char-cols

Conversation

@cornzyblack

Copy link
Copy Markdown
Contributor

Changes

Support unquoted column names with special characters in checks

Column names that require SQL identifier escaping, such as "Long Name" or "Col with $pecial character", failed in most built-in checks unless back-quoted. Validation accepted them (#1342), but at execution the checks parsed the string with F.expr, so Spark SQL read "Long Name" as column "Long" aliased to "Name". The check then failed, or silently validated a "Long" column if one existed.

The rule manager now resolves each check column against the input DataFrame once. A string that exactly matches a top-level column name is matched against the cached DataFrame schema, so the common case needs no extra Spark analysis round-trip. Other strings keep the existing expression-then-column-reference fallback.

Check functions opt in with the new register_for_column_name_resolution decorator. For those functions only, a column that resolves as a literal name but needs escaping is passed as F.col(name), back-quoted when the name contains a dot or a backtick. The decorator is applied to the 81 built-in checks that parse string columns as SQL.

Checks that already use column names (compare_datasets, has_valid_schema, is_unique, has_no_outliers, sql_expression) and all custom checks are not marked, so they keep receiving the columns exactly as defined and
existing custom checks are unaffected. Reported columns, messages and rule fingerprints are unchanged. Back-quoted names and SQL expressions work as before.

Scope is limited to the column and columns arguments. Other column arguments, such as ref_columns, group_by, column1 and column2, still require back-quoting.

Linked issues

Resolves #1202

Tests

  • manually tested
  • added unit tests
  • added integration tests
  • added end-to-end tests
  • added performance tests

Documentation and Demos

Updated the does_not_contain_pii custom function example in the pii_detection_funcs.py demo

  • added/updated demos
  • added/updated docs
  • added/updated agent skills

@mwojtyczka mwojtyczka added the under-review This PR is currently being reviewed by one of DQX maintainers. label Sep 25, 2026
@mwojtyczka mwojtyczka added the needs-changes Changes required after review label Sep 25, 2026
before. Any other name is returned as a column reference so that it is not parsed as SQL: back-quoted when it
contains a dot or a backtick, which Spark would otherwise treat as a nested field or a quote.
"""
if _UNQUOTED_IDENTIFIER_PATTERN.fullmatch(name):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reserved-keyword column names are returned unquoted, so they still mis-parse under ANSI mode (correctness).

_get_column_reference returns a name unchanged whenever it matches _UNQUOTED_IDENTIFIER_PATTERN ([A-Za-z_][A-Za-z0-9_]*). But reserved words like user, order, select, interval are syntactically valid unquoted identifiers, so they take the bare-string branch and are handed to the check, which evaluates F.expr("user").

Under Spark ANSI mode (default on recent Databricks runtimes) user parses as the USER() function, order/select/interval as keywords — not as the column. So a check on a real column named user evaluates the wrong thing or raises, which is exactly the failure class this PR set out to fix; the special-char path is covered but the reserved-word path is not. Consider back-quoting when the name is a reserved keyword (or, more simply, always returning F.col(name) for an exact schema match rather than the bare string).

if isinstance(column, str) and column in self._input_column_names:
return self._get_column_reference(column)

expression_error = self._get_analysis_error(F.expr(column) if isinstance(column, str) else column)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A malformed column string raises ParseException here and aborts the whole apply instead of skipping one check (robustness).

F.expr(column) is evaluated eagerly to build the argument to _get_analysis_error, and it parses the SQL string immediately. _get_analysis_error only catches AnalysisException, and ParseException is a sibling of AnalysisException (both extend CapturedException), not a subclass — and in any case the parse happens at this call site, outside that try. So a column string that is neither an exact DataFrame column nor parseable SQL (e.g. column="a b(") raises ParseException out of _resolve_column, crashing apply_checks for the entire rule set rather than marking just that check skipped.

(This is pre-existing behaviour from the replaced _is_invalid_column, but it's re-exposed in this rewritten path.) Catching ParseException alongside AnalysisException (or guarding the F.expr call) would degrade to a skipped check as intended.

]

@cached_property
def check_columns(self) -> list[str | Column]:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A primary column supplied positionally via check_func_args bypasses resolution (edge, but a coverage gap).

check_columns reads only check.column / check.columns. If the primary column is instead supplied positionally through check_func_args (and not promoted to column/columns), check_columns returns [], resolved_check returns the check unchanged, and the special-character handling this PR adds is silently skipped for that path — e.g. DQRowRule(check_func=is_not_null, check_func_args=["Long Name"]) on a real "Long Name" column runs F.expr("Long Name"), which Spark reads as column Long aliased Name, validating the wrong/absent column with no skip message. Worth either resolving columns passed via args too, or documenting that special-char resolution applies only to the column/columns fields.

for column, resolved in zip(self.check_columns, self.resolved_check_columns)
]
if self.check.column is not None:
return self.check.replace(column=resolved_columns[0])

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

resolved_check rebuilds the whole rule via replace() just to swap in the resolved column (efficiency — minor/optional).

For a marked check whose column needs escaping, this calls check.replace(column=...), and replace() (by its own contract) "rebuilds the instance through the constructor so validation re-runs" and cached derived state is recomputed — including get_check_condition, which was already built once at original construction. So the check-condition Column is rebuilt a second time per marked check per apply, for a column value already known to be valid. A targeted field swap (or reusing the already-built condition) would avoid re-running validation here. Minor, but it's on the hot apply path.

When you define checks **declaratively** (YAML, JSON, or list of dicts), check arguments must contain every required parameter of the check function, matching its Python signature (optional parameters may be omitted). If you use `for_each_column`, DQX merges `column` or `columns` into the arguments for each expansion. Use `DQEngine.validate_checks` to catch unknown keys, type mismatches, and missing required parameters before apply. See [Quality checks definition](/docs/guide/quality_checks_definition) for the declarative format and validation section.

<Admonition type="info" title="Column names with spaces or special characters">
The `column` and `columns` arguments of built-in checks accept column names that contain spaces or special characters exactly as they appear in the DataFrame, for example `Long Name` or `Col with $pecial character`, without back-quoting.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The admonition overpromises: reserved-keyword names (and positional check_func_args) aren't handled (docs).

This states that column/columns accept names with spaces or special characters "without back-quoting," but a column named after a SQL reserved word (user, order, select, interval, …) is still returned unquoted by _get_column_reference and mis-parses under ANSI mode (see the manager.py thread). A user who reads this and names a column order will hit an execution failure that contradicts the documented guarantee. Worth adding a caveat that reserved-keyword column names — and columns passed positionally via check_func_args — are not covered, until those paths are handled.

@cornzyblack

Copy link
Copy Markdown
Contributor Author

Thanks for the feedback, I'll have a look into it

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

needs-changes Changes required after review under-review This PR is currently being reviewed by one of DQX maintainers.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[FEATURE]: data frame with special character columns

2 participants