Week 10

Classification

Instructors
Maclean Gaulin

What is Classification

  • Predicting outcome “groups” not “levels”
  • Regression: what is bankruptcy risk
  • Classification: is firm bankrupt

What’s different from Regression?

  • Regression Approach
  • Classification Approach
What’s different from Regression?
What’s different from Regression?

What’s different from Regression?

  • Regression Approach
  • Classification Approach
What’s different from Regression?
What’s different from Regression?

What’s different from Regression?

  • Regression Approach
  • Classification Approach
What’s different from Regression?
What’s different from Regression?

Most basic: Logistic Regression

Most basic: Logistic Regression

Most basic: Logistic Regression
Most basic: Logistic Regression

Most basic: Logistic Regression

Classifier Models

  • Many different classifiers
  • Differ in assumptions
  • Differ in treatment of data
  • Differ in treatment of “error”
  • Differ in what the analyst must specify
Classifier Models

Other Regression Classifiers

Other Regression Classifiers

Improving on Regression Classifiers

  • Like OLS, we can add non-linearity in the x variables
  • We can also add more x variables
  • Deal with “unbalanced labels”
  • Bankruptcy is rare, so mostly 0s, a few 1s
  • This is the intent of C-log-log
  • Down-sampling & upweighting the “majority” class

FICO Score

  • Payment History (35%): Timeliness of past payments
  • Amounts Owed (30%): Total debt and credit utilization ratio
  • Length of Credit History (15%): How long accounts have been open
  • New Credit (10%): Number of recent inquiries and new accounts
  • Credit Mix (10%): Variety of credit types (credit cards, loans)

Non-regression Classifiers

  • Regression classifiers try to predict classes as if they were some continuous variable (e.g., probability)
  • Other classifiers just try to learn the best separation
  • Examples: SVM, perceptron

Classifying without Probability (SVM)

Classifying without Probability (SVM)

Selecting Parametric Classifiers

Selecting Parametric Classifiers

Parametric vs Non-parametric

  • Parametric classifiers specify some structure with which to approach splitting the groups or labels
  • Non-parametric classifiers just learn unconstrained
  • The “parameters” grow with the data
  • Non-parametric classifiers often better at handling text than parametric
  • Mostly because the “functional form” of text is unknown

K Nearest Neighbors

  • Assign average neighbor’s label (instead of value)
K Nearest Neighbors

Tree-based models

Tree-based models

Neural Networks

Neural Networks

Neural Networks

  • Inputs: features believed to be useful for prediction
  • Outputs: Each class has its own “neuron”
  • Output is 1 or 0 to signal that it’s class is “True”
  • In the middle: “hidden” layers learn the complex relationship between features & classes
  • Connections between neurons, and calculations in each neuron define the NN architecture

Neural Networks

Neural Networks

Cautions and Caveats

  • Overfitting is still an issue, solved the same way
  • Interpretability is harder with more complex models
  • Measure of “accuracy” is even more important for classifiers
  • Especially as it pertains to real business costs

Duration Analysis

  • Many classifiers are “one shot”
  • What about factors that build to a tipping point?
  • E.g., bankruptcy, M&A completion
  • Duration analysis, survival analysis, hazard models
  • Two components:
  • Discrete: did it happen
  • Continuous: when did it happen

Example: Loan Default

  • Syndicated loan is instantiated
  • Borrower must adhere to covenants, tested regularly
  • Default on violated covenant or missed payments
  • Over life of loan, outcome either default or term ends
  • Prediction: will borrower default and when?

Duration Analysis

  • Identify what correlates with “survival”
  • Whether a variable relates to the outcome occurring
  • How much that variable effects outcome probability
  • Measure effect on probability (more likely ever)
  • Measure effect on time to failure (more likely sooner)

Hazard Function

  • How likely outcome is to occur at time t
Hazard Function
  • Credit: https://www.linkedin.com/pulse/using-survival-model-credit-risk-scoring-loan-pricing-tirabassi-66i5c/

Survival Function

  • Probability that outcome hasn’t happened at time t
Survival Function
  • Credit: https://www.linkedin.com/pulse/using-survival-model-credit-risk-scoring-loan-pricing-tirabassi-66i5c/

Hazard & Survival

Hazard & Survival
Hazard & Survival
  • Credit: https://www.linkedin.com/pulse/using-survival-model-credit-risk-scoring-loan-pricing-tirabassi-66i5c/

Censoring

  • Right censoring is when the outcome is not observed
  • E.g., loan term ends before default
  • Left censoring is when outcome has already occurred
  • E.g., firm already in default when observation starts
  • Interval censoring is when outcome occurs in some range, or window
  • E.g., default only measured quarterly

Measuring Classifier Performance

  • Accuracy has many definitions
  • How many predicted customers actually default?
  • How many defaulting customers were predicted so?
  • Different measures used for different goals

Confusion Matrix

Outcomes
DefaultedPaid
PredictionsPredict defaultCorrectly predict defaultPaying customer called defaulter
Predict payingDefaulter predicted payingCorrectly predict paying

Confusion Matrix

Outcomes
Is PositiveIs Negative
PredictionsPredict PositiveTrue PositiveFalse Positive
Predict NegativeFalse NegativeTrue Negative
  • Positive and negative don’t mean anything other than the two outcomes
  • You could reverse them, nothing would change

Accuracy Calculations

Outcomes
PredictionsTPFP
FNTN

Accuracy Calculations

Outcomes
PredictionsTPFP
FNTN

Precision and Recall Shortcomings

Outcomes
➕➖
Predictions➕991
➖00

Classifier asymmetric Costs

  • TP / TN – predicted correctly
  • False Positive – expect default, customer pays
  • Loss of customer, or waste of resources to mitigate risk
  • False Negative – expect paying customer, defaults instead
  • Up to complete loss
Outcomes
➕➖
Predictions➕TPFP
➖FNTN

Receiver Operating Curve

Receiver Operating Curve

Receiver Operating Curve

Receiver Operating Curve

Receiver Operating Curve

Receiver Operating Curve

Cost Functions and Training

  • Training minimizes cost
  • Specifying different cost functions will result in different classifiers
  • They will weight different errors (FP/FN) differently
  • Understanding the business decision is imperative to designing a good cost function and classifier

Coming Up

  • Project 4 Proposal – Due Sunday
  • Lab/Homework 10 – Due Sunday
  • Project 3 – Due next Sunday
  • Lab next week will be a work-session for Project 3
Classification