Entropy Conditions
A Data Entropy condition matches a column whose sampled values (a sampled subset; see Sampling) look like machine-generated key material: API keys, access tokens, private keys, and other secrets. It complements a Regex condition. A regular expression recognizes a value’s shape. A Data Entropy condition measures whether the value is random, so it can tell a real key apart from an order ID, a UUID, or a hash that shares the same characters.
How Data Entropy Works
Section titled “How Data Entropy Works”ALTR scores each sampled value by how close it is to truly random over its own alphabet:
- Each value is split into segments. Separators such as whitespace,
=,?,&,@, and:split a value, so a secret embedded asAPI_KEY=<SECRET>,?token=<SECRET>, oruser:<PASSWORD>@hostis scored on its own. Segments shorter than 16 characters aren’t scored individually. - Each segment’s alphabet is inferred from 12 fixed classes, such as hexadecimal, Base64, or alphanumeric, so hex and Base64 values are each judged against their own alphabet.
- Each segment is scored in bits of evidence. Random key material scores positive. Prose, JSON, dates, UUID values, and identifiers have structure and score strongly negative.
- Base64 values are checked for structure. A segment that decodes as Base64 to a structured payload, such as Base64-encoded JSON, is scored on its decoded content and isn’t mistaken for a key. A Base64-encoded key keeps its full evidence.
Evidence is measured in bits, and each bit doubles the odds. With 1 bit of evidence, the sampled values are 2 times as likely to come from random key material as from structured text. 3.3 bits makes them about 10 times as likely, and 13.3 bits about 10,000 times. A negative value points the other way: at -6.6 bits, the values are about 100 times as likely to come from structured text as from a key. For scale, a single random 64-character hex API key typically scores about 12 bits on its own, and individual keys vary widely around that, while one such key among 40 rows of prose adds 3.44 bits to that column.
The column’s evidence is pooled into a small statistic of evidence totals. The statistic contains no sampled value and no column name. The condition passes when the column’s evidence reaches the condition’s minimum.
Fields
Section titled “Fields”- Minimum evidence – the bits of evidence the column must reach for the condition to pass. Defaults to 13.3 bits, which corresponds to roughly a 1-in-10,000 chance of flagging a column that holds no secrets. Must be greater than 0 and at most 1024.
- Minimum key size (optional) – the minimum apparent key size, in bits, of the column’s random values. The apparent key size is how many bits of randomness a value holds: its length times the bits each character carries, never more than its alphabet allows. A random 32-character hex key has an apparent key size of 128 bits, and a random 8-character alphanumeric ID about 48 bits. Use it to skip short random identifiers. For example, 112 bits excludes most random IDs while admitting real API keys. Must be a whole number from 0 to 4096; 0 or blank turns the check off.
- Only these alphabets (optional) – limits which inferred alphabets can satisfy the condition. Choose from Digits, Hex (lowercase), Hex (uppercase), Lowercase letters, Uppercase letters, Lowercase + digits, Uppercase + digits, Letters, Alphanumeric, Base64, Base64 URL-safe, and Any printable. Blank means any alphabet passes.
- Accumulate evidence across scans – carries the column’s evidence forward from one classification job to the next. See Cumulative Evidence. Off by default.
- Decode Base64-encoded values before scoring – on by default. Turn it off to score Base64 values as their encoded text.
A Data Entropy condition has no comparator, pattern, or Match Threshold. The minimum evidence is the condition’s only threshold.
Example
Section titled “Example”This condition matches a column whose sampled values carry at least 13.3 bits of evidence of random key material, with an apparent key size of at least 112 bits, over hexadecimal or Base64 alphabets, accumulating evidence across jobs:
{ "target": "ENTROPY", "minimum_evidence_bits": 13.3, "minimum_key_bits": 112, "alphabet_classes": ["hex", "HEX", "base64", "base64url"], "cumulative": true}The alphabet_classes values are digits, hex, HEX, lower, upper, lower+digits, upper+digits, alpha, alnum, base64, base64url, and printable. Set decode_base64 to false to turn off Base64 decoding.
Cumulative Evidence
Section titled “Cumulative Evidence”A secret that appears in only a few rows of a column can be missed by any single job. For example, a support_notes column of mostly prose where someone occasionally pastes a live API key produces too little evidence in one sample to flag.
When Accumulate evidence across scans is on, ALTR keeps a running evidence total for each column and adds each new job’s evidence to it. Older evidence fades by 10% per job, so one unusual sample can’t flag a column forever, while a secret that keeps appearing keeps adding up. In one measured example, a column of 40 prose rows holding a single 64-character hex key reached 3.44 bits on the first job, 10.13 bits on the second, and 16.14 bits on the third, crossing the 13.3-bit minimum on job three.
- Only evidence totals are kept. ALTR stores a small set of numbers per column, keyed by a one-way hash of the column’s fully qualified name. No sampled value and no column name is stored.
- A retried job counts once. A job whose results are reprocessed doesn’t add its evidence a second time.
- Evidence can also fall. A column whose recent samples look structured accumulates negative evidence, which can keep a single unusual sample from flagging it.
Cumulative evidence is available only for Data Entropy conditions. It applies to columns in connected data sources. Each API payload is always decided on its own.
Behavior
Section titled “Behavior”- A Data Entropy condition needs sampled row data. A metadata-only job doesn’t evaluate it, and the decision lineage marks it as not evaluated.
- An all-null or empty column is evaluated and doesn’t pass.
- Some random-looking values that aren’t secrets also score high, such as UUID values without dashes, MongoDB ObjectIds, and SHA-256 digests. Combine a Data Entropy condition with a negated Column Name condition, such as
(id|uuid|guid|hash|digest|checksum)$, the Minimum key size field, or a Regex or Data Length condition to narrow the match. - Human-chosen passwords aren’t random enough to score high. Pair a Data Entropy condition with a Column Name condition such as
passwto find password columns. - In Match Confidence scoring, a Data Entropy match is outranked by Google DLP, Snowflake Native, Amazon Comprehend, and Regex matches, and outranks a Column Content, Column Name, Data Location, Data Length, or Column Size match when both reach the same tier.
- A Data Entropy match earns its own Match Confidence tier from how far its evidence clears the condition’s minimum: High at 1.5 times the minimum or more, Medium from 1.2 times up to 1.5 times, and Low below 1.2 times. A classifier’s Match Confidence is the highest tier any of its matched conditions earns, so a classifier whose Data Entropy match rates Low and whose Column Name match rates Medium shows Medium. See Match Confidence.
Decision Lineage
Section titled “Decision Lineage”An evaluated Data Entropy condition’s decision lineage shows the evidence the decision used and the column’s evidence state:
- Secret-like – the evidence is at least 13.3 bits.
- Accumulating – the evidence is between -6.6 and 13.3 bits.
- Not secret – the evidence is -6.6 bits or lower. This state is advisory: it doesn’t change whether the condition matches, and it doesn’t stop a later job from accumulating evidence again.
The state measures the evidence against these fixed bars, so it reads the same for every classifier. Whether the condition matches depends on its own minimum evidence and other fields. With a minimum of 20 bits, a column at 16 bits shows Secret-like but doesn’t match. With a minimum of 5 bits, a column at 8 bits shows Accumulating and matches. Use the condition’s result to decide whether a column holds a secret, and the state to see how close the evidence is to the default bar.
A cumulative condition’s lineage shows both the accumulated evidence and the current job’s own evidence, with the number of rows and jobs that contributed.
ALTR Managed Secrets & Credentials Collection
Section titled “ALTR Managed Secrets & Credentials Collection”The ALTR Managed - Secrets & Credentials collection pairs a Data Entropy condition with a structural check in every classifier, so a column matches only when it both looks like a credential and carries key-like randomness. It covers generic API keys and tokens, private key material, credentials in connection strings, cloud provider access keys, and MFA or TOTP seeds. See Import the ALTR Managed Collection for how to import and sync ALTR Managed collections.
Availability
Section titled “Availability”Data Entropy conditions evaluate on ALTR Hosted, In-Warehouse, and OLTP Classification Agent jobs, and cumulative evidence works at every processing location. OLTP jobs require OLTP Classification Agent 1.23.0 or later. See Classification Processing Location for the full picture.