Machine Learning Further Information

Prev Next

Article

What it covers

Analyze

Checking your data and getting suggested rules before training

Training Output Files

The files training produces

Post-Processing Rules Reference

All 25 rules and their settings


Analyze

Analyze reads your training data before you train. It checks the data can be read, and suggests:

  • Post-processing rules that keep generated values inside the limits of your real data

  • Relationships between parent and child tables that the model should keep when it generates related tables

When to run Analyze

Run Analyze after you've set which columns are active, and before you train. It's optional, but quick and worth running every time:

  • It's an early check. If Analyze can't read your data, training won't be able to either. You find out in minutes instead of partway through a long run.

  • It gives you a starting set of post-processing rules, based on your real data.

  • It finds relationships across tables when you're training more than one related table.

Run Analyze

  1. Open your data activity.

  2. From the Actions panel, choose Analyze.

  3. Enter a threshold between 0 and 1, or leave the default of 0.3.

  4. Start the job and open the job log. When Analyze finishes, the log lists its suggestions, and the reports are saved on your host server.

The threshold

The threshold only affects relationship suggestions. It doesn't affect post-processing rule suggestions.

It sets how strong the evidence for a relationship must be before Analyze reports it:

Threshold

Effect

Lower (e.g. 0.1)

More relationships reported, including weaker ones that may not matter

Default (0.3)

A balanced starting point

Higher (e.g. 0.6)

Fewer relationships reported, only the clearest ones

Start with the default. If Analyze reports relationships you know aren't real, raise it. If it misses one you know is there, lower it.

Post-processing rule suggestions

Analyze looks at each column and records limits the real data always keeps to, such as the range of a numeric column or whether it only holds whole numbers. It suggests a post-processing rule for each.

Each suggestion appears in the log with a severity, the rule, the table and column(s), and the reason:

[MEDIUM] range_clipping on 'CUSTOMER'.credit_score - source values fall in
[473.0, 850.0] and are non-negative
[HIGH] integer_column_consistency on 'CUSTOMER'.tenure_months, credit_score,
credit_limit - these columns are whole numbers in the source data

The suggestions are also saved on your host server as post_generation_suggestions.json.

Relationship suggestions

When your rule set includes related tables, Analyze follows each foreign key from a parent table to its child table. It looks for parent columns that affect what appears in the child rows. For example, customers in a higher income band tending to take out larger loans.

These matter because each table is generated separately. Without them, a generated loan could be for any amount, whoever the customer is. With them, child rows stay consistent with their parent.

Relationships found appear in the log and in Column Correlations for the connection in the Data Dictionary.

If you run Analyze on one table, the log reports Conditional analysis complete: 0 candidate(s). This is expected. See Why did Analyze find no relationships?

Using the results

Post-processing rule suggestions. You have two choices:

  • Use them. Add them to your generation rule set, so generated data stays within the ranges and formats of the real data from the first run.

  • Leave them out for the first run. Generate once without them to see what the model produces on its own, then add rules to fix what you find. This tells you more about how well the model learned your data.

Either way, check each suggestion makes sense. A range is only as good as the data it came from. If your training data doesn't cover the full range of real values, the suggested range will be too narrow.

Relationship suggestions. Review each one before you train:

  • Keep relationships that reflect how your data really works.

  • Delete relationships that are coincidence, or that you don't want the generated data to follow.

Why did Analyze find no relationships?

  • The rule set has only one table. Analyze only looks for relationships between parent and child tables. Relationships between columns in the same table are left for the model to learn during training. A result of 0 means Analyze didn't look, not that there aren't any.

  • The child table isn't in the rule set. Both sides of the foreign key must be included.

  • The foreign key isn't defined. Check the connection in the Data Dictionary under Foreign Keys. If the key isn't declared in the database, define it as a soft key.

  • The threshold is too high. Run again with a lower threshold.

Why did Analyze only suggest ranges and whole numbers?

line_item_consistency and parent_child_aggregate only apply to data with line numbers or parent totals. For most single tables, range and whole-number rules are all Analyze can suggest. Add any other rules your data needs by hand.

When to run Analyze again

Run it again if you:

  • Add or remove tables from the rule set

  • Make columns active or inactive

  • Change or reload the source data

You don't need to run it again if you only change engine hyperparameters.


Training Output Files

When training finishes, these files are available to download from the data activity:

File

What it contains

training.json

training.json records what the training job was told to do, so it works as the job's spec. It holds the settings the job ran with: the engine chosen for each table, its hyperparameters (such as n_iter, batch_size and data_encoder_max_clusters) and the rest of the run configuration.

metadata.json

metadata.json records what the model was trained on. It lists the tables and active columns, each column's type (integer, float or categorical, as in the log's type_metadata_mapping), and any added __missing_indicator columns.


Post-Processing Rules Reference

A generative model produces realistic rows, but it has no concept of a rule. It might produce a credit score of 912.7, an invoice whose lines don't add up to its total, or a status that jumps from DRAFT to CLOSED without passing through OPEN.

Post-processing rules fix these problems. They're added to the generation rule set and applied to the generated data after the model has run.

How rules run

  • Rules run in the order they're listed. Put a rule that recalculates a value after any rules that fix its inputs. For example, clip a quantity to its range before recalculating a total from it.

  • Inactive rules are skipped. Use the Active toggle to switch a rule off without deleting it.

  • You don't need to retrain to add, change or remove a rule. Generate again to apply it.

Fields every rule has

Every rule takes a Table* (the table it applies to) and an optional Note, as well as the fields listed below. The exception is sql_alignment, which works across all tables and has no Table field.

Required fields are marked *.

Value constraints

Keep individual values inside legal limits.

range_clipping

Clips a numeric column to a minimum and maximum. Values below the minimum are raised to it; values above the maximum are lowered to it.

Field

Description

Target Column

The column to clip

Min

Lowest allowed value

Max

Highest allowed value

Use when: generated values fall outside the real range, such as a negative age.

integer_column_consistency

Rounds decimal values back to whole numbers.

Field

Description

Target Columns

One or more columns to round

Use when: whole-number columns, such as counts or months, come back with decimals.

sign_consistency

Makes the sign of an amount (positive or negative) match a category column.

Field

Description

Amount Column

The numeric column

Flag Column

The category column that decides the sign

Sign Map

Each flag value and its sign, +1 or −1

Use when: for example, debits must be negative and credits positive.

resource_limit

Caps a usage column at a capacity taken from another column, or at a fixed limit.

Field

Description

Target Column

The usage column

Capacity Column

A column holding each row's limit

Hard Limit

A fixed limit applied to every row

Use when: for example, a balance must never exceed the row's credit limit.


Calculations

Make one column equal a calculation from others.

product_rule

Sets a column to one column multiplied by another.

Field

Description

Factor A Column

The first value

Factor B Column

The second value

Target Column

Set to Factor A × Factor B

Use when: for example, line total = quantity × unit price.

conditional_ratio

Sets a column to a base column multiplied by a rate. The rate can be fixed, or chosen by the value of another column.

Field

Description

Target Column

The column to calculate

Base Column

The value the rate is applied to

Condition Column

The column whose value chooses the rate

Fixed Rate

A single rate, used when there's no rate map

Rate Map

Each condition value and its rate

Use when: for example, tax = amount × a rate that depends on the region.

zero_sum

Adjusts a numeric column so it adds up to zero within each group.

Field

Description

Target Column

The column to balance

Partition By Columns

The columns that define each group

Enforcement Strategy

How the adjustment is made: proportional (default), plug_last, random_one or round_trip

Rounding (decimals)

Decimal places, default 2

Use when: for example, the debits and credits in each journal entry must balance.

delta_smoothing

Limits how much a value can change from one row to the next in a series.

Field

Description

Target Column

The value being smoothed

Partition By Columns

The columns that define each series

Max Allowed Change

The largest change allowed between consecutive rows

Use when: a time series has unrealistic spikes.


Dates and periods

Keep dates and periods in a sensible order.

date_sequence

Makes sure a later date is at least a set number of days after an earlier one.

Field

Description

Early Date Column

The date that should come first

Late Date Column

The date that should come second

Offset (days)

The minimum gap, default 0

Strategy

push_forward (move the later date, default) or swap (swap the two)

Use when: for example, an end date falls before its start date.

sequential_event_gate

Keeps a series of date columns in order.

Field

Description

Date Columns (in order)

The date columns, listed in the order they must happen

Fix Strategy

nudge (move dates into order, default) or nullify (empty out-of-order dates)

Use when: three or more dates must happen in order, such as ordered, shipped and delivered.

holiday_date_shifter

Moves dates that fall on a weekend or public holiday to the nearest working day.

Field

Description

Date Column

The date to move

Country Code

The holiday calendar to use, default US

Shift Direction

forward (default) or backward

Use when: dates such as payment or trading dates can only fall on working days.

temporal_continuity

Links periods together so each period's opening value matches the previous period's closing value.

Field

Description

Group By Columns

The columns that define each chain of periods

Time/Period Column

The column that orders the periods

Value Column

The value carried forward

Target Value Column

The column set to the previous period's value

Recalculate Ending

Recalculate the closing value after the change

Movement Column

The change between opening and closing

Use when: for example, monthly figures must follow on from each other.

balance_carry_forward

Sets each period's opening balance to the previous period's closing balance.

Field

Description

Partition By Columns

The columns that define each account or entity

Period Column

The column that orders the periods

Opening Balance Column

Set from the previous closing balance

Closing Balance Column

The balance carried forward

Use when: account balances must carry over from one statement to the next. A simpler version of temporal_continuity for balances.

temporal_completeness

Makes sure every entity has a row for every period, and adds rows to fill any gaps.

Field

Description

Master Table*

The table listing every entity that must appear

Group By Columns

The columns that define each entity

Time/Period Column

The period column

Periods

The periods required, comma-separated, e.g. 1..12

Child Join Key Columns

Key columns in this table

Master Join Key Columns

Matching key columns in the master table

Fill Columns

Columns filled on added rows

Type Column

The column that classifies an added row

Probability Profile

Each type and its probability (0–1)

Sign Logic

Each type and its sign, +1 or −1

Metadata Columns

Other columns copied onto added rows

Use when: for example, every account needs a row for each of 12 months.


Relationships between tables

Keep parent and child tables consistent.

relational_alignment

Makes sure every child row points to a parent that exists. Can remove child rows that don't.

Field

Description

Parent Table*

The parent table

Child Key Columns

The foreign key columns in this table

Parent Key Columns

The matching key columns in the parent

Drop Orphans

Remove child rows with no matching parent

foreign_key_sync

Changes foreign keys that point to a parent that doesn't exist, so they point to a real parent instead.

Field

Description

Parent Table*

The parent table

Foreign Key Column

The column to fix

foreign_key_sync or relational_alignment? foreign_key_sync keeps the row and fixes its link. relational_alignment can remove the row instead.

parent_child_aggregate

Either recalculates a parent value from its child rows, or adjusts the child rows to match the parent.

Field

Description

Parent Table*

The parent table

Mode

update_parent (recalculate the parent, default) or align_child (adjust the children)

Child Key Columns

The foreign key columns in this table

Parent Key Columns

The matching key columns in the parent

Calculation

The calculation, e.g. sum(amount)

Target Column (parent)

The parent column to set

Use when: for example, an order total must equal the sum of its lines.

cross_table_sync

Sets a parent total to the sum of its child line amounts.

Field

Description

Parent Table*

The parent table

Group Id Column

The column linking each line to its parent

Line Amount Column

The child value to add up

Target Column (parent)

The parent total to set

cross_table_sync or parent_child_aggregate? cross_table_sync only sums lines into a parent total. parent_child_aggregate takes any calculation, and can adjust the children instead of the parent.


Structure and grouping

Control how rows are organised.

line_item_consistency

Numbers line items 1, 2, 3 … with no gaps within each group.

Field

Description

Target Column

The line number column

Group By Columns

The columns that define each document

Order By Columns

The order of lines within a document

First Number

The starting number, default 1

stochastic_grouping

Groups rows into documents, each with a random number of lines.

Field

Description

Document Id Column

The column that holds each row's document ID

Min Group Size

The fewest lines per document, default 5

Max Group Size

The most lines per document, default 100

Start Id

The first document ID

Sort By Columns

The order of rows before grouping

Use when: for example, generated invoice lines need to be grouped into invoices.

state_machine

Makes sure each entity's status only changes in an allowed order.

Field

Description

State Column

The status column

Partition By Columns

The columns that define each entity

Valid States (in order)

The statuses in their allowed order, comma-separated

Sort By Column

The column that orders each entity's rows

Use when: for example, an order must go DRAFT → OPEN → CLOSED.


Testing and privacy

Rules that deliberately change data quality, or check it.

outlier_injection_profile

Adds deliberate bad values, so you can test how other systems handle them.

Field

Description

Target Column

The column to add bad values to

Probability (0-1)

The share of rows affected, default 0.01

Method

multiplier (default), fixed or extreme_sample

Factor

The multiplier, default 10

Fixed Value

The value used by the fixed method, default 9999

This is the only rule that makes data worse on purpose.

privacy_similarity_score

Finds synthetic rows that are too similar to real rows, and flags or removes them.

Field

Description

Critical Threshold

How similar a row must be to be flagged, default 0.1

Log Risks To (path)

The file flagged rows are written to

Drop Risky Records

Remove flagged rows instead of only listing them

Use when: you're using synthetic data for privacy or compliance. A model can sometimes reproduce real rows, and this rule checks for it.


Custom logic

For anything the built-in rules don't cover.

custom_python

Runs Python code against the table's data, held as a DataFrame called df.

Field

Description

Python Code

Code that works on df

The code runs in a restricted environment. df is the only object available. You can use methods on df and its columns, but the pandas module itself (pd.) isn't available and you can't import other modules.

sql_alignment

Runs SQL (DuckDB) across all the generated tables.

Field

Description

SQL

A DuckDB SQL statement

Example:

UPDATE JET SET debitCreditIndicator = IF(amountLC >= 0, 'D', 'C');

This is the only rule with no Table field. It can read and change every generated table.


Quick index

Rule

What it does

balance_carry_forward

Opening balance = previous closing balance

conditional_ratio

Target = base × rate, rate chosen by a condition

cross_table_sync

Parent total = sum of child lines

custom_python

Python on the table's data (df)

date_sequence

Later date at least N days after the earlier date

delta_smoothing

Limits the change between consecutive rows

foreign_key_sync

Repoints broken foreign keys at real parents

holiday_date_shifter

Moves dates off weekends and holidays

integer_column_consistency

Rounds decimals to whole numbers

line_item_consistency

Numbers lines 1..n in each group

outlier_injection_profile

Adds deliberate bad values

parent_child_aggregate

Parent from children, or children to fit the parent

privacy_similarity_score

Flags rows too similar to real ones

product_rule

Target = A × B

range_clipping

Clips values to a minimum and maximum

relational_alignment

Child rows must have a parent; can remove those that don't

resource_limit

Caps usage at a capacity or fixed limit

sequential_event_gate

Keeps several dates in order

sign_consistency

Sign matches a category column

sql_alignment

SQL across all generated tables

state_machine

Statuses change only in an allowed order

stochastic_grouping

Groups rows into documents

temporal_completeness

A row for every entity and period

temporal_continuity

Opening value matches previous closing value

zero_sum

Values add up to zero in each group