Before you start
Read this section before setting up a Machine Learning data activity. The choices you make here shape your setup from the start.
It helps to understand your data before you begin: how values are spread across each column, which columns are related, and whether any values are rare or missing. This page covers the basics and should help you profile your data and pick sensible starting parameters for your first training run.
Expect the first run to be a starting point rather than the finished result. Once you've generated synthetic data, compare it with your source data, adjust the parameters based on what you find, and train again.
What is machine learning?
In the Curiosity Machine Learning data activity, machine learning means training to produce a model on your existing data so that it learns the data's statistical patterns:
How values are distributed in each column
How columns relate to each other
How related tables connect
Once trained, the model generates new synthetic records that follow those same patterns or characteristics. The records are newly generated rather than copied from the source, so you get realistic, representative test data without exposing the real records it learned from.
Generative Engines
The generative engine is the machine learning algorithm that learns the patterns in your training data. Training produces a model, which you then use to generate new rows with the same characteristics. Three engines are available currently.
Pick one based on what your data looks like:
Engine | Type | Data size | Use it when |
|---|---|---|---|
Statistical | Small (up to around 25,000 rows) | You want speed, and to see which columns depend on which | |
Neuro (variational autoencoder) | Medium to large (over 25,000 rows) | Your data is mostly numeric, smooth distributions and complex numeric relationships | |
Neural (GAN) | Medium to large (over 25,000 rows) | Your data is heavily categorical, has many distinct labels, or has very imbalanced categories |
bayesian_network models your data as a graph of column dependencies. It's fast, and it's the only engine whose learned relationships you can inspect directly. It can struggle with very wide tables.
tvae compresses each row and learns to rebuild it. It doesn't use adversarial training, so it's usually more stable than ctgan. It tends to produce average values, so outliers can get smoothed away.
ctgan trains two networks against each other. The generator makes rows and the discriminator tries to tell them from real ones. It specifically targets under-represented categories, but it's the slowest engine and the hardest to tune.
All engines
Parameter | What it does | When to change it |
|---|---|---|
encoder_max_clusters (data_encoder_max_clusters on tvae) | How finely continuous columns are divided before training | Raise it (for example from 10 to 20) if values come out too coarse, or if less common values go missing |
sampling_patience | How many attempts the engine gets to produce a valid row before it fails | Raise it if generation fails with "failed to meet constraints" |
bayesian_network - Training Guide (coming soon)
Parameter | What it does | When to change it |
|---|---|---|
struct_learning_search_method | How it searches for column relationships | Use tabu. tree_search is faster but gives each column only one parent |
struct_max_indegree | Maximum number of columns any one column can depend on | Higher is more accurate but much slower. For tables with more than 20 columns, keep it at 2–3 |
struct_learning_n_iter | Maximum number of search steps | Raise it only if the learned graph looks incomplete |
tvae - Training Guide
Parameter | What is does | When to change it |
|---|---|---|
n_iter | Number of training passes | Raise it for a better fit. Lower it if training takes too long |
weight_decay | Discourages the model from memorising individual rows | Increase it if the output is noisy. Reduce it if outliers disappear |
lr | Learning rate: how big a step each update takes | Lower it if training fails with NaN errors |
batch_size | Rows processed per update | Don't set it too small on small data, because training can become unstable |
ctgan - Training Guide (coming soon)
Parameter | What it does | When to change it |
|---|---|---|
n_iter | Number of training passes. The default is 2000, which is heavy on a CPU | Lower it for a first run |
generator_n_layers_hidden | Depth of the generator network | 2-3 is usually enough |
discriminator_n_iter | How many times the discriminator updates for each generator update | Raise it if the generator's output is poor and doesn't improve |
discriminator_dropout | Randomly switches off part of the discriminator during training | Increase it if the discriminator is too strong for the generator to learn |
lr / batch_size | Learning rate and rows per update | If ctgan keeps producing the same values, lower lr |