Machine Learning Introduction

Prev Next

Before you start

Read this section before setting up a Machine Learning data activity. The choices you make here shape your setup from the start.

It helps to understand your data before you begin: how values are spread across each column, which columns are related, and whether any values are rare or missing. This page covers the basics and should help you profile your data and pick sensible starting parameters for your first training run.

Expect the first run to be a starting point rather than the finished result. Once you've generated synthetic data, compare it with your source data, adjust the parameters based on what you find, and train again.

What is machine learning?

In the Curiosity Machine Learning data activity, machine learning means training to produce a model on your existing data so that it learns the data's statistical patterns:

  • How values are distributed in each column

  • How columns relate to each other

  • How related tables connect

Once trained, the model generates new synthetic records that follow those same patterns or characteristics. The records are newly generated rather than copied from the source, so you get realistic, representative test data without exposing the real records it learned from.

Generative Engines

The generative engine is the machine learning algorithm that learns the patterns in your training data. Training produces a model, which you then use to generate new rows with the same characteristics. Three engines are available currently.

Pick one based on what your data looks like:

Engine

Type

Data size

Use it when

bayesian_network

Statistical

Small (up to around 25,000 rows)

You want speed, and to see which columns depend on which

tvae

Neuro (variational autoencoder)

Medium to large (over 25,000 rows)

Your data is mostly numeric, smooth distributions and complex numeric relationships

ctgan

Neural (GAN)

Medium to large (over 25,000 rows)

Your data is heavily categorical, has many distinct labels, or has very imbalanced categories

bayesian_network models your data as a graph of column dependencies. It's fast, and it's the only engine whose learned relationships you can inspect directly. It can struggle with very wide tables.

tvae compresses each row and learns to rebuild it. It doesn't use adversarial training, so it's usually more stable than ctgan. It tends to produce average values, so outliers can get smoothed away.

ctgan trains two networks against each other. The generator makes rows and the discriminator tries to tell them from real ones. It specifically targets under-represented categories, but it's the slowest engine and the hardest to tune.

All engines

Parameter

What it does

When to change it

encoder_max_clusters (data_encoder_max_clusters on tvae)

How finely continuous columns are divided before training

Raise it (for example from 10 to 20) if values come out too coarse, or if less common values go missing

sampling_patience

How many attempts the engine gets to produce a valid row before it fails

Raise it if generation fails with "failed to meet constraints"

bayesian_network - Training Guide (coming soon)

Parameter

What it does

When to change it

struct_learning_search_method

How it searches for column relationships

Use tabu. tree_search is faster but gives each column only one parent

struct_max_indegree

Maximum number of columns any one column can depend on

Higher is more accurate but much slower. For tables with more than 20 columns, keep it at 2–3

struct_learning_n_iter

Maximum number of search steps

Raise it only if the learned graph looks incomplete

tvae - Training Guide

Parameter

What is does

When to change it

n_iter

Number of training passes

Raise it for a better fit. Lower it if training takes too long

weight_decay

Discourages the model from memorising individual rows

Increase it if the output is noisy. Reduce it if outliers disappear

lr

Learning rate: how big a step each update takes

Lower it if training fails with NaN errors

batch_size

Rows processed per update

Don't set it too small on small data, because training can become unstable

ctgan - Training Guide (coming soon)

Parameter

What it does

When to change it

n_iter

Number of training passes. The default is 2000, which is heavy on a CPU

Lower it for a first run

generator_n_layers_hidden

Depth of the generator network

2-3 is usually enough

discriminator_n_iter

How many times the discriminator updates for each generator update

Raise it if the generator's output is poor and doesn't improve

discriminator_dropout

Randomly switches off part of the discriminator during training

Increase it if the discriminator is too strong for the generator to learn

lr / batch_size

Learning rate and rows per update

If ctgan keeps producing the same values, lower lr