File Classification
File classification finds sensitive data in files that don’t live in a database table: exports in an Amazon S3 bucket, a file share, or a document your own application holds. ALTR reads a CSV, JSON, or XML file the way it reads a table, one field per CSV header, JSON key, or XML element, reads a plain-text file line by line, and evaluates the classifiers in a collection against the values it finds. The collections you build for a database work unchanged against files.
ALTR classifies files two ways:
- The OLTP Classification Agent reads an S3 bucket or a file system directory registered as a data source, and produces a classification report. See File Sources.
- The Classification API classifies one file your application uploads, synchronously, and returns where each sensitive value sits. See Classifying a Single File with the API.
Supported Formats
Section titled “Supported Formats”ALTR reads the values in these formats and classifies each one:
- CSV – one field per header column.
- JSON – one field per leaf key, with array indexes collapsed so every element of a list is the same field.
- XML – one field per element or attribute. HTML is read as XML.
- Plain text – one field per file, with each non-blank line as a value. YAML and other text files are read as plain text.
ALTR identifies these formats and reports the file with its format, as a single file-level entry:
- Documents – PDF, DOCX, XLSX, PPTX, legacy Microsoft Office, RTF
- Archives – ZIP, GZIP
- Images – PNG, JPEG, GIF, TIFF
- Columnar – Parquet
ALTR detects a file’s format from its content rather than its extension, so a .txt file that holds JSON is read as JSON.
File Sources
Section titled “File Sources”The OLTP Classification Agent classifies two kinds of file data source, each registered as a repository the same way a database is. Contact ALTR to enable file classification.
The two file data sources are:
- Amazon S3 – a bucket, optionally narrowed to a key prefix. The agent discovers objects by listing the bucket, or by reading the storage locations of the tables in an AWS Glue Data Catalog database.
- File System – a directory tree mounted into the agent container.
In ALTRNet, you add either one under Data Configuration > Data Sources as Amazon S3 or File System, and classify it from Classify Data like any other data source. A new job you run from Classify Data has Needs Human Review selected by default, so you approve or reject its findings in Human-in-the-Loop Review before they reach the report. Registering a file data source, supplying AWS credentials, and running a job are covered in File Classification with Agents.
To find documents stored as values in a database column, such as a column of PDF attachments, use the Column Content condition on a database job instead.
What a Job Reads
Section titled “What a Job Reads”A file job reads each file’s values and classifies them. With Classify file contents cleared, it instead inventories every file with its identified format and produces no findings. Which paths a job reads, what it excludes, how much of each file it reads, and whether it samples or scans every value are set when you run the job – see Run a File Classification Job on File Classification with Agents.
Field Names in the Report
Section titled “Field Names in the Report”A file job discovers structural names – CSV headers, JSON keys, XML element and attribute names – and each becomes a field in the report. What the report keeps of those names is a per-job choice, made under Field Names when you run the job and set for every format at once or per detected format:
- Hashed (the default) replaces each name with a keyed hash, so the same name groups together across files and jobs that use the same hash key. The report keeps the shape of the path, array markers, and attribute markers.
The key comes from
METADATA_MAP_HASH_KEYon the agent container; set it so hashes stay stable if the agent’s signing key is rotated. Without it, the agent derives the key from its signing key, and if it can’t read that either, a hashed job keeps no field names and logs a warning. - Plaintext keeps the names as written.
- Not retained collapses each file to a single file-level field. Values are still read and classified; only the record of which field they came from is dropped.
Column Name and Data Location conditions match the real field names under Hashed, the same as under Plaintext. Choose plaintext when a reviewer needs to read the structure in the report itself. Per-format settings let one job hash CSV headers, which may be a data row misread as a header, while keeping JSON keys in plaintext.
Where a Value Sits
Section titled “Where a Value Sits”A finding on a file carries the byte offsets, in the file’s original bytes, of the values that matched. The offsets appear in the downloaded report and in the classification map returned by the single-file API, so a downstream process can tokenize or encrypt exactly that region.
A classifier’s Google DLP or Amazon Comprehend condition inspects the file’s text as a document, in overlapping chunks for large files, rather than one field at a time, so an entity split across two CSV cells or two lines is found as one value and attached to every field it spans. Text found outside any field, such as prose between records or an XML comment, is reported on the file itself.
File Reports in ALTRNet
Section titled “File Reports in ALTRNet”A file job’s classification report is organized by file rather than by table. Its levels are Bucket, Prefix, Object, and Field for Amazon S3, and Root, Directory, File, and Field for a file system, and the report header counts Fields Scanned where a database report counts columns. Selecting a file shows its size, how much of it was read, and each field’s classifier matches with their Match Confidence.
A file that was read and matched no classifier is marked as checked and clean, and a file in an identify-only format is listed with its format and no fields or findings. A structured file appears as a single CONTENT field when the job ran with Classify file contents cleared, when its field names were set to Not retained, or when the file was empty or didn’t parse within the per-file read limit.
A file report offers Run New Scan, which opens Classify Data with the previous job’s scope, exclusions, and options already selected.
Classifying a Single File with the API
Section titled “Classifying a Single File with the API”The Classification API classifies one file your application uploads, synchronously, against a collection you name. The file is at most 128 KiB, sent in the request body, and ALTR classifies it in memory: the bytes are never stored, and the record ALTR keeps holds the offsets, classifier names, and tags, never the matched values. Regular expression, Google DLP, and Amazon Comprehend conditions inspect the file’s text as a document rather than one token at a time, and the resulting matches are placed onto the tokens they cover.
Unlike a file job, a single-file request keeps field names in plaintext unless it says otherwise.
The Classification API is listed with the other ALTR APIs on the API page, and its reference documents the request, the classification map, and the supporting endpoints. Calls authenticate with an API key.