TVAE Training Guide

Prev Next

This guide covers everything you need to train with TVAE (Tabular Variational Autoencoder) to produce a model which can be used to generate synthetic data with the same characteristics as your source data.

Section

What it covers

1. Training

Step-by-step: create the activity and train a model

2. Engine

What TVAE is, how it works and when to choose it

3. Hyperparameters

Activity settings and every TVAE parameter, with tuning advice

4. Model

What training produces, how to read the log and when to retrain

5. Column Inclusion / Exclusion Guide

Which columns to train on and which to leave out

6. Generate and Validate

Generating data, checking it against the source and fixing issues


The example dataset

New to machine learning? Load our ML Sample Dataset - Lending into your database first then work through this guide with the CUSTOMER table on its own first, then add the other tables once you're comfortable.

The examples in this guide use the CUSTOMER table from a synthetic consumer-lending dataset. It has 12,000 rows, one per customer, and 11 columns:

Column

Type

What it holds

customer_id

Text

Unique identifier, e.g. C000001

open_date

Date

When the customer relationship started

tenure_months

Whole number

Months since open_date (0–127)

age_band

Category

Six bands, 18-24 to 65+

income_band

Category

Five bands. About 7% of values are missing

employment_cd

Category

FULLTIME, PARTTIME, SELFEMP, CONTRACT, RETIRED

region_cd

Category

Eight regions across Ireland and the UK

channel_cd

Category

BRANCH, ONLINE or BROKER

credit_score

Whole number

473–850

credit_limit

Decimal

1,000–40,000, in steps of 500

dti_ratio

Decimal

Debt-to-income ratio, 0.05–0.65

It mixes categories, whole numbers and decimals, includes missing values, and has strong relationships between columns. For example, customers with higher credit scores have higher credit limits. Those are the things you'll check the synthetic data for.


1. Training

Before you start, you need:

  • A connection to your source database, scanned into the Data Dictionary

  • A Database Definition built from that connection

Create the data activity

1. Create a Machine Learning data activity. Go to Data Activities and create a new activity of type Machine Learning.

2. Add a name and description. Use a name you'll recognise later, such as Customer – TVAE. In the description, note the table and engine you're training.

3. Select a Database Definition and rule set. Select your Database Definition, then choose New Rule Set. If you already have a training rule set for this data, choose Existing Rule Set instead.

Configure the activity

4. Select the source connection. On the Configuration tab, select the Source Connection the training data is read from.

5. Leave Find Rules off. It adds significant training time, and on a first run you want to confirm that everything completes. See Activity settings.

6. Set Imputation Strategy to Random Sample. It keeps each column's real proportions and works for every column type. See Activity settings.

7. Set Generative Engine to tvae. See 2. Engine for why you'd choose it.

8. Set the engine parameters. For a first run on the example dataset, use the starting values in Suggested starting values.

Choose tables and columns

9. Select your tables. On the Tables tab, check the table(s) you want and click Add Selected. Start with one table. It's the quickest way to confirm everything works before you add related tables.

10. Finish the wizard. Click Next, then Finish. Your new data activity opens.

11. Run Analyze (optional). From the Actions panel, choose Analyze. It confirms the data can be read and suggests post-processing rules, such as the valid range of each numeric column. Relationship suggestions only appear when your rule set includes related tables. See Analyze.

12. Set which columns are active. Open your training rule set and expand your table. Make identifiers and repeated information inactive. In the example, customer_id and open_date are inactive. See 5. Column Inclusion / Exclusion Guide.

13. Set the engine at table level. In the same rule set, set the table's generative engine to tvae. A table-level engine overrides the one on the Configuration tab for that table.

Train

14. Train the model. Go back to the dashboard and select Train Model.

15. Check the log. Within the first minute, open the job log and find the line beginning Fit internal model. Confirm it says tvae and shows the parameters you set. If it doesn't, stop the job and check steps 7 and 13. See Reading the training log.

When training finishes, the model is attached to the activity and its output files are available to download. See 4. Model.


2. Engine

What TVAE is

TVAE is a neural network made of two halves:

  • The encoder takes a row of your data and compresses it into a short list of numbers, called the latent representation.

  • The decoder takes that list of numbers and rebuilds the row.

During training, the model learns to rebuild rows accurately. It also learns to keep the compressed space smooth and evenly spread out. Because the space is smooth, you can pick a random point in it, pass it through the decoder, and get a new row that looks like your real data but isn't a copy of any real row. That's how generation works.

Before training, each column is converted into a form the network can use:

  • Categories are split into one yes/no value per category.

  • Numbers are grouped into clusters, and each value is stored as its cluster plus its position within that cluster. This helps the model handle columns whose values bunch into several groups. The number of clusters is set by data_encoder_max_clusters. See 3. Hyperparameters.

How TVAE compares with the other engines

tvae

ctgan

bayesian_network

Type

Neural network (autoencoder)

Neural network (two competing networks)

Probability graph

Best for

Small to medium tables with mixed categories and numbers

Imbalanced data with many categories

Small to medium tables where you want to see the relationships

Training stability

Good

Can be unstable

Good

Speed (platform rating)

7/10

6/10

8/10

Can you see what it learned?

No

No

Yes, as a graph of dependencies

When to choose TVAE

Choose TVAE when:

  • Your table mixes categories and numbers, and relationships between columns matter

  • You want a neural model but ctgan is unstable or slow on your data

  • You don't need to see which columns depend on which

Choose something else when:

  • You need to explain or inspect the relationships the model found. Use bayesian_network.

  • One or more category columns are very imbalanced (a rare value in under 1% of rows) and must still appear. Try ctgan.

  • The table is very small, with a few hundred rows or fewer. Neural models are more likely to copy real rows. Use bayesian_network, or check the output with privacy_similarity_score.

Global engine vs table-level engine

You can set the engine in two places:

  • Configuration tab: the default engine for the activity.

  • Table level, in the training rule set: the engine for that table.

The table-level setting wins. Set both to tvae and confirm in the log (see Reading the training log).

CPU vs GPU

TVAE runs on the CPU unless the server has a GPU. On CPU it trains noticeably slower than bayesian_network. If the log includes CUDA libraries not found, you're on CPU. Reduce n_iter for a first run so you get a result sooner.

Do I need a target column?

Usually not. TVAE learns how all the columns relate to each other, so no single column is treated as the answer. A target column is only needed when one column should drive the rest of the generation.

The example CUSTOMER table has no outcome column in any case. The default flag in this dataset is held on the loan table.


3. Hyperparameters

A hyperparameter is a setting you choose before training. It controls how the model learns, as opposed to what it learns from the data. Changing a hyperparameter means you need to retrain.

Activity settings

These two settings are on the Configuration tab and apply to the whole activity, whichever engine you use.

Find Rules. Searches for mathematical formulas linking your columns, such as one column being another divided by a third. It adds significant time to training. Leave it off for a first run, and turn it on later if you suspect columns are calculated from each other.

Imputation Strategy. Decides how missing values are filled before training.

Option

How it fills a gap

Use for

Random Sample

A value drawn from the column's existing values

Any column. Recommended default

Mode

The most common value

Categories

Least Frequent

The least common value

Rarely useful

Mean

The average

Numbers only

Median

The middle value

Numbers only

Zero

0

Numbers where missing genuinely means zero

Random Sample keeps each column's proportions closest to the real data. Whichever option you choose, the model also records which values were originally missing, in an extra column called <column>__missing_indicator. This lets it learn the pattern of what's missing as well as the values.

TVAE parameters

Defaults shown are the engine's standard values. The form on your platform version may show different defaults or a subset of these parameters.

Training length and pace

Parameter

Default

What it does

n_iter

1000

Number of training rounds (epochs). Each round passes through all the training data once. More rounds give the model longer to learn, and take longer to run

batch_size

500

Rows the model looks at before each update. Larger is faster per round but can learn less detail. Keep it well below your row count

lr

0.001

Learning rate: how big a change the model makes at each update. Too high and training is erratic; too low and it's slow to improve

weight_decay

0.00001

Small penalty that discourages the model from relying on extreme internal values. Helps prevent it memorising the training data

Network size

Parameter

Default

What it does

n_units_embedding

500

Size of the compressed representation each row is squeezed into

encoder_n_layers_hidden

3

Number of layers in the encoder

encoder_n_units_hidden

500

Size of each encoder layer

decoder_n_layers_hidden

3

Number of layers in the decoder

decoder_n_units_hidden

500

Size of each decoder layer

encoder_nonlin, decoder_nonlin

leaky_relu

Mathematical function applied between layers. Rarely needs changing

More layers and larger layers let the model capture more complex patterns. They also take longer to train, use more memory, and make it easier for the model to memorise rows on a small table.

Regularisation

Parameter

Default

What it does

encoder_dropout

0.1

Share of the encoder's connections randomly switched off during each update. Stops the model depending on any single connection, which reduces memorising

decoder_dropout

0

The same, for the decoder

loss_factor

1

How much weight the model gives to rebuilding rows accurately, compared with keeping the compressed space smooth. Higher values favour accuracy; lower values favour variety

clipping_value

1

Caps how large a single update can be, to stop training becoming unstable

Data encoding and generation

Parameter

Default

What it does

data_encoder_max_clusters

10

Maximum number of clusters each numeric column is split into. More clusters capture more detail, such as a sharp threshold, at the cost of more memory and training time

compress_dataset

Off

Reduces the data before training. Leave off unless the table is very large

sampling_patience

500

Number of attempts the engine makes to produce a valid row before giving up

Suggested starting values

For a first run on a table like the example (around 10,000 rows, a mix of 5–10 category and numeric columns, CPU only):

Parameter

Value

Why

n_iter

300

Enough to learn the main patterns and finish in a reasonable time on CPU. Increase once the process works

batch_size

500

Default. Around 24 updates per round at 12,000 rows

lr

0.001

Default

n_units_embedding

128

A smaller compressed space suits a table with around 10 columns and trains faster

Network layers and units

Defaults

Change only if the output shows the model missing relationships

encoder_dropout

0.1

Default

loss_factor

1

Default

data_encoder_max_clusters

10

Default. Raise it if numeric columns come back too smooth

Tuning: which setting to change

Change one setting at a time and compare each run against the source.

What you see in the output

Try

Distributions or relationships are weak across most columns

Increase n_iter

Numeric columns look too smooth, or a sharp step in the data is lost

Increase data_encoder_max_clusters

The model still misses relationships after longer training

Increase n_units_embedding or the layer sizes

Rows look too similar to each other, with little variety

Lower loss_factor

Synthetic rows are copies, or near-copies, of real rows

Increase encoder_dropout or weight_decay, lower n_iter, or use a smaller network

Training is erratic, or fails partway through

Lower lr

Training is too slow or runs out of memory

Lower n_iter, the layer sizes or data_encoder_max_clusters; check for identifier columns (see section 5)


4. Model

What a trained model is

A trained model is the result of a training job, attached to your data activity. It contains everything the engine learned from your data. It takes a snapshot of the training rule set when the job starts, so changes you make to the rule set afterwards don't affect a model that's already trained.

You use the model by selecting it in a generation rule set (see 6. Generate and Validate).

Output files

When training finishes, these files are available to download:

training.json

metadata.json

See Training output files for more detail.

Reading the training log

Open the job log shortly after training starts and check these lines.

Engine and parameters:

Fit internal model 'tvae' for table '<table>' with params: {...}

Confirm the engine is tvae and the parameters match what you set.

Columns and types:

type_metadata_mapping: {'CUSTOMER': {'integer_cols': [...],
'float_cols': [...], 'categorical_cols': [...]}}

Check that inactive columns are missing from this list, and that each column is classed as you'd expect. A number stored as text, for example, would appear under categorical_cols.

Category encoding:

VERIFY [CUSTOMER][channel_cd]: [('BRANCH', 0), ('BROKER', 1), ('ONLINE', 2)]

Each category is listed with the number it's stored as. Check for unexpected values, such as a misspelling or a stray blank.

Missing values: columns with gaps get an extra <column>__missing_indicator column. In the example, you'd see income_band__missing_indicator.

Hardware: CUDA libraries not found means training is on CPU and will take longer.

Checking what the model learned

You can't look inside a TVAE model to see what it learned. The only way to check it is to generate data and compare it against the source. See Compare with the source.

When to retrain

Situation

Retrain?

You changed an engine parameter

Yes

You made columns active or inactive

Yes

The source data has changed

Yes

The model missed a relationship or distribution

Yes, with adjusted settings

Values break a fixed rule (range, whole numbers, date order)

No, add a post-processing rule

You want more or fewer rows

No, set the row count at generation


5. Column Inclusion / Exclusion Guide

Every active column is something the model has to learn. Columns that carry no useful information slow training down, use more memory, and can make the output worse. Columns that give too much away can teach the model the wrong thing.

Exclude these

Unique identifiers. Primary keys, reference numbers and any other column with a different value in every row. The model can't learn a pattern from them. Worse, each value is treated as a separate category, which is the most common reason a run is slow or runs out of memory. Generate these with a generator function instead. Example: customer_id.

Columns calculated from another column. If one column is worked out from another, keep one of them. The model would otherwise have to learn the calculation, and will get it slightly wrong. Example: tenure_months is calculated from open_date, so open_date is inactive. If you need both, recalculate the second one with a post-processing rule.

Columns that give away another column's value. A column set directly from another column's value. For example, a status column that is set to DEFAULT whenever a default flag is 1. It adds nothing, and it can mask the real relationships.

Free text. Notes, comments and descriptions. TVAE treats each distinct text as a separate category, so it can't produce new text and may repeat real text. Mask or generate these separately.

Personal details. Names, emails, phone numbers and addresses. Excluding them keeps real personal details out of the model. Generate them with generator functions.

Include these

Columns with the patterns you want to keep. Any column involved in a relationship you care about must be active, on both sides of that relationship. Example: credit_score and credit_limit both stay active, so the model can learn that they rise together.

Columns with missing values. Keep them. The model learns where values are missing as well as what they are, and the pattern often carries information. Example: income_band is missing more often for BROKER customers. Keeping the column keeps that pattern.

Low-cardinality categories. Codes, bands and flags with a manageable number of values are what TVAE learns best.

Think before including these

Column type

Guidance

Dates

TVAE learns dates, but not rules between them. Add date_sequence or sequential_event_gate for date order

Two columns that say nearly the same thing

For example, a balance and the same balance as a percentage. Keep one, and recalculate the other with a post-processing rule

Categories with hundreds of values

Each value adds to memory and training time. Group rare values into an OTHER category in the source, or exclude the column

Foreign keys

Needed when you train related tables together, so child rows link to parents. Leave active in multi-table rule sets

Columns filled by a generator function

Exclude from training. They're filled at generation time

Quick checklist

Column

Active?

Why

Primary key / unique ID

✗

No pattern to learn; slows training

Foreign key (multi-table)

✓

Keeps parent–child links

Category, code or band

✓

What TVAE learns best

Number

✓

Unless calculated from another column

Date

✓

Add date-order rules afterwards

Calculated from another column

✗

Recalculate with a post-processing rule

Gives away another column's value

✗

Adds nothing, masks real relationships

Free text

✗

Can't be generated; may repeat real text

Personal details

✗

Use generator functions

Column with missing values

✓

The pattern of what's missing is information


6. Generate and Validate

Generate data

1. Create a generation rule set. Select your trained model in its configuration.

2. Review the rule set. For a first run you don't need to change anything. The table and column settings are copied from training. Leave post-processing rules out for now, so you can see what the model produces on its own.

3. Generate a small sample. Select Generate and enter a small row count, such as 500. Check that the job completes, the right columns are filled, and data types and formats are what you expect.

4. Generate a full run. Generate again with the same number of rows as your source table (12,000 in the example). Matching the row count makes the comparison fair. The job log shows the row count and where each output file was saved.

Compare with the source

Download the generated data from your host server and compare it with the source data. This shows whether the model has reproduced the patterns in the real data closely enough for your purpose, and where it falls short.

Compare on:

  • Distributions. Categories should appear in similar proportions. If 43% of source customers came through BRANCH, expect about 43%. Numeric columns should have similar ranges, averages and shape.

  • Rare values. Rare categories and extreme values should still appear, in roughly the same amounts.

  • Relationships. Columns that move together in the source should still move together. In the example, credit limit should still rise with credit score.

  • Missing values. Each column should have a similar share of missing values. About 7% of income_band values should be empty.

  • Valid values. Every value should be possible: known categories, sensible ranges, and business rules that still hold.

  • Links between tables. If you generated more than one table, every child row should point to a real parent row, and each parent should have a similar number of children.

  • Similarity to real records. No synthetic row should be a copy of a real one.

Leave inactive columns and generator-filled columns out of the comparison.

What to expect: a close match, not an exact one. Proportions and missing-value rates are usually within a few percent of the source. Relationships between columns often come back a little weaker. Business rules are the most likely thing to break, because the model has no concept of a rule.

A small difference across many columns is normal. If one column or relationship is far off, improve the output as described below.

Improve the output

There are two ways to improve the output. You can use either or both.

Retrain the model if the model missed something, such as a weak relationship or a distorted distribution. See Tuning: which setting to change. Also check that the relevant columns were active.

Add post-processing rules to fix things the model can't be expected to learn, such as fixed ranges, whole numbers or date order. Rules are applied after the model has run, so you can add them and generate again without retraining. Rules run in the order they're listed, so put a rule that recalculates a value after any rules that fix its inputs.

Problem in the output

Rule

Values outside the real range

range_clipping

Decimals in whole-number columns

integer_column_consistency

Dates on weekends or holidays

holiday_date_shifter

Dates in the wrong order

date_sequence, sequential_event_gate

Child rows pointing to parents that don't exist

foreign_key_sync, relational_alignment

Parent totals that don't match their child rows

cross_table_sync, parent_child_aggregate

Values that should be calculated from other columns

product_rule, custom_python, sql_alignment

Synthetic rows too similar to real ones

privacy_similarity_score

For the example, Analyze suggests range_clipping on credit_score, credit_limit, tenure_months and dti_ratio, and integer_column_consistency on credit_score, credit_limit and tenure_months. These are a good starting set.

See Post-processing rules reference for all 25 rules and their settings.

Repeat until you're happy

Retrain and regenerate until the output meets your needs. Change one thing at a time and compare each run against the source, so you can see which change made the difference.