File Classification with Agents
File classification finds sensitive values in an Amazon S3 bucket or a file share without moving the files out of your environment. The OLTP Classification Agent reads the files where they are and sends ALTR the findings. You register the bucket or directory as a data source, register a service user for it, and run a job from Classify Data the same way you classify a database. For which formats ALTR reads and what a file report contains, see File Classification.
File classification runs on the OLTP Classification Agent, which is enabled per organization. Contact ALTR to enable file classification.
Prerequisites
Section titled “Prerequisites”-
The classification agent is deployed and registered – see Data Classification with Agents.
-
For Amazon S3, a credential with
s3:ListBucketon the bucket ands3:GetObjecton its objects, provided in one of two ways:- An AWS Identity and Access Management (IAM) role the agent assumes. The agent’s own AWS identity needs
sts:AssumeRoleon the role, and the role’s trust policy must allow that identity. - An access-key pair stored in AWS Secrets Manager in the bucket’s region. The agent’s own AWS identity needs
secretsmanager:GetSecretValueon the secret.
A job that reads the AWS Glue Data Catalog also needs
glue:GetTables, plusglue:GetCrawlerandglue:StartCrawlerif it starts a crawler. - An AWS Identity and Access Management (IAM) role the agent assumes. The agent’s own AWS identity needs
-
For a file system, the directory is mounted into the agent container, read-only where possible. Its path as the container sees it, not the host path, is the Root Path you enter when registering the data source.
Register a File Data Source
Section titled “Register a File Data Source”A file data source is registered as a repository, like a database, but with a bucket or directory in place of a hostname and port.
To register an Amazon S3 bucket or a file system directory:
- Select Data Configuration > Data Sources in the navigation menu.
- Click Add Data Source.
- On Select Connection Type, locate Amazon S3 or File System and click Select.
- Enter a Repository Name. The name must be unique, use only lowercase letters, numbers, and underscores, and be at most 32 characters.
- Optionally, enter a Description.
- For Amazon S3, enter the Bucket and its Region, such as
us-east-1. Optionally, enter:- Prefix – a key prefix that bounds every job on this data source. Paths in the report are relative to it. A prefix that doesn’t exist lists as empty rather than failing, so a job on it finishes with zero files.
- Glue Database – the AWS Glue Data Catalog database a job can read to find the bucket’s objects. In ALTRNet, the Glue Data Catalog scan mode is available only when this is set.
- For File System, enter the Root Path – the absolute path of the directory as the agent container sees it, such as
/app/scan/hr. - Click Save. ALTRNet registers the data source and lists it under Repositories in the Type dropdown on the Data Sources page.
Register the Service User
Section titled “Register the Service User”A job on a file data source runs as a repository service user, the same as a job on a database. For Amazon S3, the service user names the AWS credential the agent uses. For a file system, the agent reads the root path with its own permissions, so the service user stores no credential; it exists only because every job requires one.
To register the service user:
- Select Data Configuration > Data Sources in the navigation menu.
- Select Repositories from the Type dropdown.
- Click the file data source.
- Click the Service Users tab.
- Click Register Service User.
- Enter a Username.
- For Amazon S3, enter the credential under AWS Secrets Manager, the only Secret Source offered. If you enter both an IAM Role and an Amazon Resource Name (ARN), the agent uses the secret’s access keys and doesn’t assume the role. Enter one of the following:
- Amazon Resource Name (ARN) – a secret holding an access-key pair, as a JSON object with
aws_access_key_id,aws_secret_access_key, and optionallyaws_session_token. - IAM Role – a role for the agent to assume.
- Amazon Resource Name (ARN) – a secret holding an access-key pair, as a JSON object with
- For File System, skip the credential. ALTRNet shows no Secret Source and stores no credential.
- Click Register User. ALTRNet adds the user to the Service Users tab.
Run a File Classification Job
Section titled “Run a File Classification Job”A job on a file data source uses the same Classify Data dialog as a database job, with steps for choosing which paths to read and how to read them.
To classify a file data source:
- Select Data Classification > Classification Reports in the navigation menu.
- Click Classify Data. ALTR displays a dialog to configure the job.
- On Select a data source, select Amazon S3 or File System as the Connection Type, then select the Data Source.
- Under Scan Scope, keep Entire bucket or Entire root to read the whole data source, or select Scope to a prefix or Scope to a directory to read only the paths you name. Scoping adds a step to the dialog.
- If you scoped the job, add each prefix or directory to read, relative to the data source’s prefix or root path, and press Enter or click Add. A job can scope to up to 50 paths, and anything outside them is never listed or read. A path covers everything beneath it and stops at the folder boundary, so
finance/doesn’t coverfinance-archive/. - On Configure data access, select the Service User and the Agent, then set the scan options.
- On Exclude Paths, optionally enter each prefix or directory to skip and press Enter. A job can exclude up to 100 paths, and excluded paths are never listed or read.
- On Configure classification, select a Collection. Processing Location is always Agent for a file data source.
- Leave Needs Human Review selected to approve or reject findings before ALTR generates the report, or clear it to generate the report directly. See Human-in-the-Loop Review.
- Under Field Names, select Hashed (the default), Plaintext, or Not retained for the field names the job discovers. To choose differently for one format, open Per-format overrides. See Field Names in the Report.
- On Per-File Read Limit, select how much of each file to read. 10 MB (default) reads files up to 10 MiB (10485760 bytes) in full and the first 10 MiB of anything larger. Sizes in the dialog are binary, so each option is that many MiB or GiB. The report marks a file read only up to the limit as partially read.
- On Review details, check the job’s scope, options, and exclusions, then click Classify Data. ALTR sends the job to the agent.
Scan Options
Section titled “Scan Options”The Configure data access step sets how the agent reads the data source:
- Classify file contents – on by default. Reads each file, expands its structure into fields, and classifies the values. When cleared, the job inventories every file with its identified format and produces no findings.
- Scan Mode (Amazon S3 data sources only) – how the agent finds objects in the bucket. Crawl (list & sniff), the default, lists the bucket. Glue Data Catalog uses the AWS Glue Data Catalog, AWS’s index of the tables stored in S3, and reads the storage location of each table in the data source’s Glue database that lies inside the bucket and prefix. It lists the bucket as a fallback when the database has no tables or can’t be read.
- Glue Crawler (Amazon S3 data sources only, with Glue Data Catalog selected) – optionally, the AWS Glue crawler that catalogs the bucket. Select Start the Glue crawler when the catalog is empty to have the agent start it when the catalog has no tables, wait up to about 10 minutes for it to finish, then read the catalog as it stands.
- Sniff file bytes to detect content type (Amazon S3 data sources only) – on by default. Reads the leading bytes of each object to identify its format. When cleared on a job with Classify file contents off, the agent identifies each object from its Content-Type and key extension and reads no object bytes.
Under Advanced Options:
- Scan full file – classifies every value in every field instead of a sample. The per-file read limit still applies.
- Include Glob – reads only paths matching the glob, such as
finance/**/*.csv. - Exclude Glob – skips paths matching the glob, such as
**/tmp/**. Applied after Include Glob. - Max Files – stops discovery after this many files.
Manage File Sources with the API
Section titled “Manage File Sources with the API”Everything in the procedures above is also available through the ALTR APIs listed on the API page. Repositories and service users are managed through the Sidecar Repo Config API, and jobs and their reports through the Classification API; calls authenticate with an API key.