ML Sample Dataset - Lending

Prev Next

The following is a sample dataset for you to test out the Machine Learning data activity with. It is a synthetic consumer-lending dataset consisting of a small bank’s customers, loans taken out and how the loans are being repaid. The zip folder contains 3 related table CSVs which can be loaded into your database. Once the lending tables are loaded, make them available to the platform as you would for any other data activity: create a Database Connection to the database, create a Data Definition from it, and run a Scan Database activity if this database hasn't been scanned already.

New to machine learning? Start with the TVAE training guide and the CUSTOMER table on its own. A single table has no relationships to keep in sync, so you'll get a working result sooner. Add the other tables as you get more confident.


Machine_Learning_Sample_Dataset
771.13 KB

The example dataset

A customer can have up to three loans, and some have none. Each loan has one or two behaviour snapshots taken on different dates.

CUSTOMER

12,000 rows, one per customer.

Column

Type

What it holds

customer_id

Text

Unique identifier, e.g. C000001

open_date

Date

When the customer relationship started

tenure_months

Whole number

Months since open_date (0–127)

age_band

Category

Six bands, 18-24 to 65+

income_band

Category

Five bands. About 7% of values are missing

employment_cd

Category

FULLTIME, PARTTIME, SELFEMP, CONTRACT, RETIRED

region_cd

Category

Eight regions across Ireland and the UK

channel_cd

Category

BRANCH, ONLINE or BROKER

credit_score

Whole number

473–850

credit_limit

Decimal

1,000–40,000, in steps of 500

dti_ratio

Decimal

Debt-to-income ratio, 0.05–0.65

LOAN_ACCOUNT

18,000 rows, up to three per customer. About 3% of customers have no loans.

Column

Type

What it holds

account_id

Text

Unique identifier, e.g. A000001

customer_id

Text

Links the loan to a customer

product_cd

Category

PERSONAL, AUTO, HOMEIMP, CONSOLID

principal_amt

Decimal

Amount borrowed, 500–33,000. Never more than the customer's credit limit

term_months

Whole number

Length of the loan: 24, 36, 48, 60, 72 or 84 months

interest_rate

Decimal

Annual rate, 4.27–17.23%

origination_dt

Date

When the loan started. Never before the customer's open_date

status_cd

Category

ACTIVE, CLOSED, DEFAULT or DELINQ

default_flag

Whole number

1 if the loan defaulted. About 10% of loans

ACCOUNT_BEHAVIOUR

20,000 rows, one or two per loan.

Column

Type

What it holds

behaviour_id

Text

Unique identifier, e.g. B0000001

account_id

Text

Links the snapshot to a loan

snapshot_dt

Date

The date this snapshot describes

outstanding_bal

Decimal

How much is still owed

utilisation_ratio

Decimal

Outstanding balance as a share of the credit limit, 0–1

payments_made_cnt

Whole number

Payments made so far

last_payment_amt

Decimal

The most recent payment

days_delinquent

Whole number

Days behind on payments, 0–180

arrears_cnt

Whole number

Number of missed payments, 0–6

Why it's a useful training set

  • It mixes column types. Categories, whole numbers, decimals and dates, all in the same table.

  • It has missing values. About 7% of income_band is empty, and the gaps aren't spread evenly — they're far more common for customers who came through BROKER.

  • Columns are related to each other. Customers with higher credit scores have higher credit limits and are offered lower interest rates. These are the relationships you'll check the synthetic data for in Part 4.

  • The tables are linked. Every loan belongs to a real customer and every snapshot to a real loan, so you can see whether those links survive generation once you train on more than one table.

ACCOUNT_BEHAVIOUR also contains around 100 rows with deliberately impossible values, such as a negative balance or a payment larger than the amount owed. They're there for testing data validation, so don't be surprised to find them.