Skip to content

Classifier Rules

ALTR Native classification matches data using classifier rules. A classifier has one or more rules, combined with AND/OR. A rule is a container that holds one or more conditions: a column name lookup, a regular expression match against row data, an ML-based classifier hit, and so on. Conditions within a single rule can be nested (combined with AND/OR) to express more specific matches, and any individual condition can be negated with NOT.

A single-rule classifier might hold just one Regex Pattern Match condition. A multi-rule classifier combines several conditions, nested within one or more rules, to target data more precisely — for example, a column whose values match a credit-card pattern and whose name contains card, but not in the staging schema.

Column Metadata covers condition types that match a column’s metadata, without sampling row data:

Built-In covers condition types that evaluate sampled column values directly:

External covers condition types that match using an external classification service:

Combining Rules and Conditions with AND, OR, and NOT

Section titled “Combining Rules and Conditions with AND, OR, and NOT”

A classifier’s rules combine with AND or OR. Within a single rule, conditions can be nested (grouped into AND/OR subgroups to express multi-factor logic) and any individual condition can be negated with NOT. This nesting among conditions inside a rule is what “nested conditions” (or “sub-conditions”) refers to; rules themselves don’t nest inside one another.

This example shows one rule matching a column whose name contains ssn and whose sampled values match an SSN pattern:

{
"operator": "AND",
"conditions": [
{ "target": "METADATA", "comparator": "contains", "pattern": "ssn" },
{ "target": "ROW_DATA", "comparator": "matches", "pattern": "\\d{3}-\\d{2}-\\d{4}" }
]
}
LimitValue
Maximum nesting depth5
Maximum conditions per group10
Maximum leaf conditions in a classifier20

A classifier that exceeds these limits is rejected when you save it.

The classifier editor groups condition types into 3 categories: Column Metadata, Built-In, and External.

These conditions match a column’s metadata and don’t require sampling row data.

These conditions evaluate the sampled column values directly.

  • Regex Conditions — match a column’s sampled values against a regular expression.
  • Column Content Conditions — match the dominant data format detected in a column’s sampled values, such as JSON or PDF.
  • Data Length Conditions — match the length, always measured in UTF-8 bytes, of a column’s sampled values against a numeric cutoff.

These conditions match using an external classification service.

  • Snowflake Conditions — match using the Snowflake Native condition: the privacy and semantic categories Snowflake’s own classification produces.
  • Google DLP Conditions — match a Google DLP infoType detection at or above a likelihood threshold.
  • Amazon Comprehend Conditions — match PII entities Amazon Comprehend detects in sampled values.

Several condition types share the same comparator options:

  • Pattern-matching conditions (Regex, Column Name) use matches (evaluates the pattern as an RE2 regex), contains, equals, starts_with, or ends_with. RE2 patterns are case-sensitive by default; use an inline flag such as (?i) at the start of the pattern for case-insensitive matching.
  • Numeric comparison conditions (Column Size, Data Length) use greater_than, greater_than_or_equal, less_than, less_than_or_equal, or equals.

Every condition’s JSON includes a target field identifying its condition type: METADATA (Column Name), COLUMN_LOCATION (Data Location), COLUMN_SIZE, ROW_DATA (Regex), CONTENT_TYPE (Column Content), DATA_LENGTH, SNOWFLAKE_NATIVE, GDLP (Google DLP), or AMAZON_COMPREHEND.

Availability by Data Source and Processing Location

Section titled “Availability by Data Source and Processing Location”

Which conditions a classification job evaluates varies by data source and, on Snowflake, by processing location. A condition that can’t be evaluated for a given job is treated as not evaluated rather than failing the job: vacuously satisfied under AND logic, non-contributing under OR logic. Each condition type’s own page states where it evaluates; Classification Processing Location covers the full picture by location.