The following is a sample dataset for you to test out the Machine Learning data activity with. It is a synthetic consumer-lending dataset consisting of a small bank’s customers, loans taken out and how the loans are being repaid. The zip folder contains 3 related table CSVs which can be loaded into your database. Once the lending tables are loaded, make them available to the platform as you would for any other data activity: create a Database Connection to the database, create a Data Definition from it, and run a Scan Database activity if this database hasn't been scanned already.
New to machine learning? Start with the TVAE training guide and the CUSTOMER table on its own. A single table has no relationships to keep in sync, so you'll get a working result sooner. Add the other tables as you get more confident.
The example dataset
A customer can have up to three loans, and some have none. Each loan has one or two behaviour snapshots taken on different dates.
CUSTOMER
12,000 rows, one per customer.
Column | Type | What it holds |
|---|---|---|
| Text | Unique identifier, e.g. |
| Date | When the customer relationship started |
| Whole number | Months since |
| Category | Six bands, |
| Category | Five bands. About 7% of values are missing |
| Category |
|
| Category | Eight regions across Ireland and the UK |
| Category |
|
| Whole number | 473–850 |
| Decimal | 1,000–40,000, in steps of 500 |
| Decimal | Debt-to-income ratio, 0.05–0.65 |
LOAN_ACCOUNT
18,000 rows, up to three per customer. About 3% of customers have no loans.
Column | Type | What it holds |
|---|---|---|
| Text | Unique identifier, e.g. |
| Text | Links the loan to a customer |
| Category |
|
| Decimal | Amount borrowed, 500–33,000. Never more than the customer's credit limit |
| Whole number | Length of the loan: 24, 36, 48, 60, 72 or 84 months |
| Decimal | Annual rate, 4.27–17.23% |
| Date | When the loan started. Never before the customer's |
| Category |
|
| Whole number | 1 if the loan defaulted. About 10% of loans |
ACCOUNT_BEHAVIOUR
20,000 rows, one or two per loan.
Column | Type | What it holds |
|---|---|---|
| Text | Unique identifier, e.g. |
| Text | Links the snapshot to a loan |
| Date | The date this snapshot describes |
| Decimal | How much is still owed |
| Decimal | Outstanding balance as a share of the credit limit, 0–1 |
| Whole number | Payments made so far |
| Decimal | The most recent payment |
| Whole number | Days behind on payments, 0–180 |
| Whole number | Number of missed payments, 0–6 |
Why it's a useful training set
It mixes column types. Categories, whole numbers, decimals and dates, all in the same table.
It has missing values. About 7% of
income_bandis empty, and the gaps aren't spread evenly — they're far more common for customers who came throughBROKER.Columns are related to each other. Customers with higher credit scores have higher credit limits and are offered lower interest rates. These are the relationships you'll check the synthetic data for in Part 4.
The tables are linked. Every loan belongs to a real customer and every snapshot to a real loan, so you can see whether those links survive generation once you train on more than one table.
ACCOUNT_BEHAVIOUR also contains around 100 rows with deliberately impossible values, such as a negative balance or a payment larger than the amount owed. They're there for testing data validation, so don't be surprised to find them.