Article | What it covers |
|---|---|
Checking your data and getting suggested rules before training | |
The files training produces | |
All 25 rules and their settings |
Analyze
Analyze reads your training data before you train. It checks the data can be read, and suggests:
Post-processing rules that keep generated values inside the limits of your real data
Relationships between parent and child tables that the model should keep when it generates related tables
When to run Analyze
Run Analyze after you've set which columns are active, and before you train. It's optional, but quick and worth running every time:
It's an early check. If Analyze can't read your data, training won't be able to either. You find out in minutes instead of partway through a long run.
It gives you a starting set of post-processing rules, based on your real data.
It finds relationships across tables when you're training more than one related table.
Run Analyze
Open your data activity.
From the Actions panel, choose Analyze.
Enter a threshold between 0 and 1, or leave the default of 0.3.
Start the job and open the job log. When Analyze finishes, the log lists its suggestions, and the reports are saved on your host server.
The threshold
The threshold only affects relationship suggestions. It doesn't affect post-processing rule suggestions.
It sets how strong the evidence for a relationship must be before Analyze reports it:
Threshold | Effect |
|---|---|
Lower (e.g. 0.1) | More relationships reported, including weaker ones that may not matter |
Default (0.3) | A balanced starting point |
Higher (e.g. 0.6) | Fewer relationships reported, only the clearest ones |
Start with the default. If Analyze reports relationships you know aren't real, raise it. If it misses one you know is there, lower it.
Post-processing rule suggestions
Analyze looks at each column and records limits the real data always keeps to, such as the range of a numeric column or whether it only holds whole numbers. It suggests a post-processing rule for each.
Each suggestion appears in the log with a severity, the rule, the table and column(s), and the reason:
[MEDIUM] range_clipping on 'CUSTOMER'.credit_score - source values fall in
[473.0, 850.0] and are non-negative
[HIGH] integer_column_consistency on 'CUSTOMER'.tenure_months, credit_score,
credit_limit - these columns are whole numbers in the source data
The suggestions are also saved on your host server as post_generation_suggestions.json.
Relationship suggestions
When your rule set includes related tables, Analyze follows each foreign key from a parent table to its child table. It looks for parent columns that affect what appears in the child rows. For example, customers in a higher income band tending to take out larger loans.
These matter because each table is generated separately. Without them, a generated loan could be for any amount, whoever the customer is. With them, child rows stay consistent with their parent.
Relationships found appear in the log and in Column Correlations for the connection in the Data Dictionary.
If you run Analyze on one table, the log reports Conditional analysis complete: 0 candidate(s). This is expected. See Why did Analyze find no relationships?
Using the results
Post-processing rule suggestions. You have two choices:
Use them. Add them to your generation rule set, so generated data stays within the ranges and formats of the real data from the first run.
Leave them out for the first run. Generate once without them to see what the model produces on its own, then add rules to fix what you find. This tells you more about how well the model learned your data.
Either way, check each suggestion makes sense. A range is only as good as the data it came from. If your training data doesn't cover the full range of real values, the suggested range will be too narrow.
Relationship suggestions. Review each one before you train:
Keep relationships that reflect how your data really works.
Delete relationships that are coincidence, or that you don't want the generated data to follow.
Why did Analyze find no relationships?
The rule set has only one table. Analyze only looks for relationships between parent and child tables. Relationships between columns in the same table are left for the model to learn during training. A result of 0 means Analyze didn't look, not that there aren't any.
The child table isn't in the rule set. Both sides of the foreign key must be included.
The foreign key isn't defined. Check the connection in the Data Dictionary under Foreign Keys. If the key isn't declared in the database, define it as a soft key.
The threshold is too high. Run again with a lower threshold.
Why did Analyze only suggest ranges and whole numbers?
line_item_consistency and parent_child_aggregate only apply to data with line numbers or parent totals. For most single tables, range and whole-number rules are all Analyze can suggest. Add any other rules your data needs by hand.
When to run Analyze again
Run it again if you:
Add or remove tables from the rule set
Make columns active or inactive
Change or reload the source data
You don't need to run it again if you only change engine hyperparameters.
Training Output Files
When training finishes, these files are available to download from the data activity:
File | What it contains |
|---|---|
| training.json records what the training job was told to do, so it works as the job's spec. It holds the settings the job ran with: the engine chosen for each table, its hyperparameters (such as |
| metadata.json records what the model was trained on. It lists the tables and active columns, each column's type (integer, float or categorical, as in the log's |
Post-Processing Rules Reference
A generative model produces realistic rows, but it has no concept of a rule. It might produce a credit score of 912.7, an invoice whose lines don't add up to its total, or a status that jumps from DRAFT to CLOSED without passing through OPEN.
Post-processing rules fix these problems. They're added to the generation rule set and applied to the generated data after the model has run.
How rules run
Rules run in the order they're listed. Put a rule that recalculates a value after any rules that fix its inputs. For example, clip a quantity to its range before recalculating a total from it.
Inactive rules are skipped. Use the Active toggle to switch a rule off without deleting it.
You don't need to retrain to add, change or remove a rule. Generate again to apply it.
Fields every rule has
Every rule takes a Table* (the table it applies to) and an optional Note, as well as the fields listed below. The exception is sql_alignment, which works across all tables and has no Table field.
Required fields are marked *.
Value constraints
Keep individual values inside legal limits.
range_clipping
Clips a numeric column to a minimum and maximum. Values below the minimum are raised to it; values above the maximum are lowered to it.
Field | Description |
|---|---|
Target Column | The column to clip |
Min | Lowest allowed value |
Max | Highest allowed value |
Use when: generated values fall outside the real range, such as a negative age.
integer_column_consistency
Rounds decimal values back to whole numbers.
Field | Description |
|---|---|
Target Columns | One or more columns to round |
Use when: whole-number columns, such as counts or months, come back with decimals.
sign_consistency
Makes the sign of an amount (positive or negative) match a category column.
Field | Description |
|---|---|
Amount Column | The numeric column |
Flag Column | The category column that decides the sign |
Sign Map | Each flag value and its sign, +1 or −1 |
Use when: for example, debits must be negative and credits positive.
resource_limit
Caps a usage column at a capacity taken from another column, or at a fixed limit.
Field | Description |
|---|---|
Target Column | The usage column |
Capacity Column | A column holding each row's limit |
Hard Limit | A fixed limit applied to every row |
Use when: for example, a balance must never exceed the row's credit limit.
Calculations
Make one column equal a calculation from others.
product_rule
Sets a column to one column multiplied by another.
Field | Description |
|---|---|
Factor A Column | The first value |
Factor B Column | The second value |
Target Column | Set to Factor A × Factor B |
Use when: for example, line total = quantity × unit price.
conditional_ratio
Sets a column to a base column multiplied by a rate. The rate can be fixed, or chosen by the value of another column.
Field | Description |
|---|---|
Target Column | The column to calculate |
Base Column | The value the rate is applied to |
Condition Column | The column whose value chooses the rate |
Fixed Rate | A single rate, used when there's no rate map |
Rate Map | Each condition value and its rate |
Use when: for example, tax = amount × a rate that depends on the region.
zero_sum
Adjusts a numeric column so it adds up to zero within each group.
Field | Description |
|---|---|
Target Column | The column to balance |
Partition By Columns | The columns that define each group |
Enforcement Strategy | How the adjustment is made: |
Rounding (decimals) | Decimal places, default 2 |
Use when: for example, the debits and credits in each journal entry must balance.
delta_smoothing
Limits how much a value can change from one row to the next in a series.
Field | Description |
|---|---|
Target Column | The value being smoothed |
Partition By Columns | The columns that define each series |
Max Allowed Change | The largest change allowed between consecutive rows |
Use when: a time series has unrealistic spikes.
Dates and periods
Keep dates and periods in a sensible order.
date_sequence
Makes sure a later date is at least a set number of days after an earlier one.
Field | Description |
|---|---|
Early Date Column | The date that should come first |
Late Date Column | The date that should come second |
Offset (days) | The minimum gap, default 0 |
Strategy |
|
Use when: for example, an end date falls before its start date.
sequential_event_gate
Keeps a series of date columns in order.
Field | Description |
|---|---|
Date Columns (in order) | The date columns, listed in the order they must happen |
Fix Strategy |
|
Use when: three or more dates must happen in order, such as ordered, shipped and delivered.
holiday_date_shifter
Moves dates that fall on a weekend or public holiday to the nearest working day.
Field | Description |
|---|---|
Date Column | The date to move |
Country Code | The holiday calendar to use, default |
Shift Direction |
|
Use when: dates such as payment or trading dates can only fall on working days.
temporal_continuity
Links periods together so each period's opening value matches the previous period's closing value.
Field | Description |
|---|---|
Group By Columns | The columns that define each chain of periods |
Time/Period Column | The column that orders the periods |
Value Column | The value carried forward |
Target Value Column | The column set to the previous period's value |
Recalculate Ending | Recalculate the closing value after the change |
Movement Column | The change between opening and closing |
Use when: for example, monthly figures must follow on from each other.
balance_carry_forward
Sets each period's opening balance to the previous period's closing balance.
Field | Description |
|---|---|
Partition By Columns | The columns that define each account or entity |
Period Column | The column that orders the periods |
Opening Balance Column | Set from the previous closing balance |
Closing Balance Column | The balance carried forward |
Use when: account balances must carry over from one statement to the next. A simpler version of temporal_continuity for balances.
temporal_completeness
Makes sure every entity has a row for every period, and adds rows to fill any gaps.
Field | Description |
|---|---|
Master Table* | The table listing every entity that must appear |
Group By Columns | The columns that define each entity |
Time/Period Column | The period column |
Periods | The periods required, comma-separated, e.g. |
Child Join Key Columns | Key columns in this table |
Master Join Key Columns | Matching key columns in the master table |
Fill Columns | Columns filled on added rows |
Type Column | The column that classifies an added row |
Probability Profile | Each type and its probability (0–1) |
Sign Logic | Each type and its sign, +1 or −1 |
Metadata Columns | Other columns copied onto added rows |
Use when: for example, every account needs a row for each of 12 months.
Relationships between tables
Keep parent and child tables consistent.
relational_alignment
Makes sure every child row points to a parent that exists. Can remove child rows that don't.
Field | Description |
|---|---|
Parent Table* | The parent table |
Child Key Columns | The foreign key columns in this table |
Parent Key Columns | The matching key columns in the parent |
Drop Orphans | Remove child rows with no matching parent |
foreign_key_sync
Changes foreign keys that point to a parent that doesn't exist, so they point to a real parent instead.
Field | Description |
|---|---|
Parent Table* | The parent table |
Foreign Key Column | The column to fix |
foreign_key_sync or relational_alignment? foreign_key_sync keeps the row and fixes its link. relational_alignment can remove the row instead.
parent_child_aggregate
Either recalculates a parent value from its child rows, or adjusts the child rows to match the parent.
Field | Description |
|---|---|
Parent Table* | The parent table |
Mode |
|
Child Key Columns | The foreign key columns in this table |
Parent Key Columns | The matching key columns in the parent |
Calculation | The calculation, e.g. |
Target Column (parent) | The parent column to set |
Use when: for example, an order total must equal the sum of its lines.
cross_table_sync
Sets a parent total to the sum of its child line amounts.
Field | Description |
|---|---|
Parent Table* | The parent table |
Group Id Column | The column linking each line to its parent |
Line Amount Column | The child value to add up |
Target Column (parent) | The parent total to set |
cross_table_sync or parent_child_aggregate? cross_table_sync only sums lines into a parent total. parent_child_aggregate takes any calculation, and can adjust the children instead of the parent.
Structure and grouping
Control how rows are organised.
line_item_consistency
Numbers line items 1, 2, 3 … with no gaps within each group.
Field | Description |
|---|---|
Target Column | The line number column |
Group By Columns | The columns that define each document |
Order By Columns | The order of lines within a document |
First Number | The starting number, default 1 |
stochastic_grouping
Groups rows into documents, each with a random number of lines.
Field | Description |
|---|---|
Document Id Column | The column that holds each row's document ID |
Min Group Size | The fewest lines per document, default 5 |
Max Group Size | The most lines per document, default 100 |
Start Id | The first document ID |
Sort By Columns | The order of rows before grouping |
Use when: for example, generated invoice lines need to be grouped into invoices.
state_machine
Makes sure each entity's status only changes in an allowed order.
Field | Description |
|---|---|
State Column | The status column |
Partition By Columns | The columns that define each entity |
Valid States (in order) | The statuses in their allowed order, comma-separated |
Sort By Column | The column that orders each entity's rows |
Use when: for example, an order must go DRAFT → OPEN → CLOSED.
Testing and privacy
Rules that deliberately change data quality, or check it.
outlier_injection_profile
Adds deliberate bad values, so you can test how other systems handle them.
Field | Description |
|---|---|
Target Column | The column to add bad values to |
Probability (0-1) | The share of rows affected, default 0.01 |
Method |
|
Factor | The multiplier, default 10 |
Fixed Value | The value used by the |
This is the only rule that makes data worse on purpose.
privacy_similarity_score
Finds synthetic rows that are too similar to real rows, and flags or removes them.
Field | Description |
|---|---|
Critical Threshold | How similar a row must be to be flagged, default 0.1 |
Log Risks To (path) | The file flagged rows are written to |
Drop Risky Records | Remove flagged rows instead of only listing them |
Use when: you're using synthetic data for privacy or compliance. A model can sometimes reproduce real rows, and this rule checks for it.
Custom logic
For anything the built-in rules don't cover.
custom_python
Runs Python code against the table's data, held as a DataFrame called df.
Field | Description |
|---|---|
Python Code | Code that works on |
The code runs in a restricted environment. df is the only object available. You can use methods on df and its columns, but the pandas module itself (pd.) isn't available and you can't import other modules.
sql_alignment
Runs SQL (DuckDB) across all the generated tables.
Field | Description |
|---|---|
SQL | A DuckDB SQL statement |
Example:
UPDATE JET SET debitCreditIndicator = IF(amountLC >= 0, 'D', 'C');
This is the only rule with no Table field. It can read and change every generated table.
Quick index
Rule | What it does |
|---|---|
| Opening balance = previous closing balance |
| Target = base × rate, rate chosen by a condition |
| Parent total = sum of child lines |
| Python on the table's data ( |
| Later date at least N days after the earlier date |
| Limits the change between consecutive rows |
| Repoints broken foreign keys at real parents |
| Moves dates off weekends and holidays |
| Rounds decimals to whole numbers |
| Numbers lines 1..n in each group |
| Adds deliberate bad values |
| Parent from children, or children to fit the parent |
| Flags rows too similar to real ones |
| Target = A × B |
| Clips values to a minimum and maximum |
| Child rows must have a parent; can remove those that don't |
| Caps usage at a capacity or fixed limit |
| Keeps several dates in order |
| Sign matches a category column |
| SQL across all generated tables |
| Statuses change only in an allowed order |
| Groups rows into documents |
| A row for every entity and period |
| Opening value matches previous closing value |
| Values add up to zero in each group |