Data activity results
Every data activity produces a results ZIP file, downloadable from the job details screen. The contents depend on which activity was run. This page describes what you can expect to find in each type.
Data masking
Masking configuration (JSON)
The masking configuration file records the configuration and audit detail for the masking job. It is an ETL job configuration that tells the masking engine three things:
Source - where the data was read from, including the connection, table and the specific fields selected for masking.
Transform - the masking rule applied to each field, such as date shifting, alphanumeric replacement or lookup-based substitution.
Sink - where the output was written, and the files generated before and after masking.
Use this file to confirm exactly how sensitive data was protected, verify that the intended masking rules were applied, and evidence the transformation from original to masked values. Because it documents both the data sources and the masking logic, it is the primary artifact for data governance, validation and regulatory review.
Generated masking scripts
Where the activity generates high-performance masking scripts, the results folder also contains a conf.json job configuration.
This defines:
The source connection and table the data is read from.
The masking transformations applied to each sensitive field.
The target the masked values are written to.
Environment settings and job-level parameters.
Data discovery (scan)
A scan produces a parent job and one or more child jobs. Each has its own results file, selectable from the drop-down on the job details screen. Both contain the same set of files.
Before and after pairs
Scan outputs are produced in matching pairs. The Before file captures the state prior to any masking or remediation; the After file captures the same information once processing has completed. Compare the two to confirm that processing produced the expected changes.
File | Contains |
|---|---|
| Database columns and the distinct or sample values found in each. Use the pair to validate masking outcomes by seeing how values changed. |
| Database columns and the classification tags assigned to them by data discovery for example PII, Email, Phone, FirstName, Address along with whether each tag is active. These tags drive the masking and protection rules applied afterwards. |
| The same enumerated column values, recorded at definition level. |
| The same column classifications, recorded at definition level. |
Definition-level files carry a numeric suffix, for example DefinitionTagsBefore_909_1270.
Detection rules
These files define how the scan identifies sensitive data. They are included in the results so the classification outcome can be traced back to the rules that produced it.
File | Contains |
|---|---|
NameRules | Pattern-matching rules that identify sensitive data from column names, linking common naming conventions to classifications such as PII, contact details, financial information and account data. |
RegexRules | Regular expression rules that identify sensitive data from the content of a field - email addresses, National Insurance numbers, credit card numbers, postal codes, phone numbers, IP addresses and other PII. |
SeedlistRules | Rules that classify values against predefined seed lists, including the category to assign, the seed list to reference and the validation threshold required for a match. |
ScanConfig | The configuration the scan ran with: database connection reference, scan identifiers, profiling thresholds, enumeration rules, processing limits and output report locations. |
Reports and Outputs
File | Contains |
|---|---|
ScanReport.xlsx | A summary of the scan results, listing the columns identified and the sensitivity tags assigned to each. Use it to review and validate the outcome of the scan. |
ScanReport.txt | A detailed log of the scan, recording the detection rules, categories and data patterns used. This is the audit trail of how each sensitive element was detected and tagged. |
ScanResults.xlsx | The full scan results: identified columns with their classification categories and the representative values found during profiling. This is the evidence behind each classification decision. |
TableSizes.xlsx | An inventory of the tables scanned - schema name, table name, table identifier and row count. Use it to assess the scope of the dataset, plan scans and confirm that the expected tables were included. |
UpdateDatabaseStatistics.sql | A script that refreshes optimiser statistics across the tables in the schema. Run it after significant data changes, such as a data load or a masking run, so the query planner has accurate row and distribution statistics. |
Data generation
The results ZIP contains a CSV file for each synthetic data table requested on the submit form, along with a report folder recording the activity version, run status and report information.
Pipelines
The variables listed in the pipeline submit form results are a snapshot of every value the pipeline resolved when the job was submitted. Because a pipeline can be built from a wide range of activities and expression functions, the variables it produces vary considerably from one design to the next - the exact names, number and purpose will always be specific to the pipeline that was run.
The types below are examples of what this section commonly contains, rather than a complete list:
Derived values - values calculated from the submit form inputs, often using conditional logic.
Constructed queries - database statements assembled as text from parameters and other variables before being passed to a database activity.
Loop or row context - values that change on each iteration as the pipeline works through a collection of records.
Read-back queries - statements that retrieve the records the job has just created or amended, usually keyed on the job ID, so the results can be verified or exported.
File paths - output locations built from a base directory parameter and a file name.
A given pipeline may include some, none or several of these, alongside other values entirely. Variables prefixed par are parameters exposed on the submit form or supplied by the platform; the remainder are working variables defined inside the pipeline itself.
Use this section to confirm that each value resolved as expected. A variable holding an empty string, a zero, or an unresolved placeholder is usually the first sign of a mis-configured parameter.