Week 4.2 — ML Important Practice & Interview Prep
🎯 TL;DR — This page is your end-to-end ML interview arsenal: bias-variance, regularisation, cross-validation, feature engineering, model selection, and full pipeline patterns — the exact knowledge gaps that separate junior from senior ML candidates.
🧠 Mental Model
ML model quality is like tuning a guitar: too tight (high variance/overfitting) and the string snaps on new songs; too loose (high bias/underfitting) and it plays the wrong notes. Cross-validation is your tuning fork — it tells you where you are on the dial. Pipelines are your guitar case: everything (scaler + model) travels together safely, preventing data leakage (your most expensive mistake).
📋 Core Concepts — Quick Reference Table
| Concept | What It Is | Why It Matters |
|---|---|---|
| Bias | Error from wrong assumptions (model too simple) | Causes underfitting — misses real patterns |
| Variance | Error from sensitivity to training noise (model too complex) | Causes overfitting — fails on new data |
| Cross-Validation | Rotate held-out fold K times, average scores | Reliable estimate without wasting data |
| Stratified K-Fold | CV that preserves class ratios per fold | Required for imbalanced classification |
| L1 (Lasso) | Regularisation adding sum of absolute weights | Drives weak features to exactly 0 (feature selection) |
| L2 (Ridge) | Regularisation adding sum of squared weights | Shrinks all weights; handles correlated features |
| ElasticNet | L1 + L2 combined | Best of both: selection + shrinkage |
| Data Leakage | Test info contaminating training (e.g., scaling before split) | Gives falsely optimistic metrics; model fails in production |
| SMOTE | Synthetic minority oversampling inside pipeline | Safe way to fix class imbalance without leakage |
| GridSearchCV | Exhaustive hyperparameter search with CV | Reproducible best-param selection |
| RandomizedSearchCV | Random sample from param distributions | Faster; often comparable to grid search |
| Feature Importance | Tree-model weight per feature | Identify which features drive predictions |
| Pipeline | Chain of preprocessing + model steps | Prevents leakage; enables clean CV and serialisation |
| Learning Curve | Train vs val score vs training set size | Diagnoses overfitting/underfitting visually |
| Ensemble Bagging | Train models on bootstrap samples, average | Reduces variance (Random Forest) |
| Ensemble Boosting | Sequentially correct previous model errors | Reduces bias (XGBoost, LightGBM) |
🔢 Key Steps / Process
- EDA — distributions, nulls, class balance, correlations; decide metrics upfront
- Split —
train_test_split(stratify=y)before any preprocessing - Baseline —
DummyClassifier(strategy='most_frequent')to establish floor - Preprocess inside Pipeline — imputer → scaler → encoder → model (no leakage)
- Model selection — try LR, RF, GBM; evaluate with Stratified K-Fold CV
- Handle imbalance — SMOTE inside pipeline or
class_weight='balanced' - Hyperparameter tuning — RandomizedSearchCV first, then GridSearchCV to narrow
- Diagnose bias/variance — plot learning curves; overfitting → regularise; underfitting → more features/complexity
- Final evaluation — test set used exactly once; report F1, AUC, confusion matrix
- Serialise & document —
joblib.dump(best_pipeline, 'model.pkl'); log feature list, metrics, training date
💻 Code Cheatsheet
# ============================================================
# WEEK 4.2 — ML INTERVIEW PREP: COMPLETE CODE CHEATSHEET
# All code is copy-paste runnable. Requires:
# pip install scikit-learn imbalanced-learn shap scipy matplotlib seaborn
# ============================================================
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
from sklearn.datasets import make_classification
from sklearn.model_selection import (train_test_split, StratifiedKFold,
cross_val_score, cross_validate,
GridSearchCV, RandomizedSearchCV,
learning_curve)
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.impute import SimpleImputer, KNNImputer
from sklearn.metrics import (classification_report, confusion_matrix,
roc_auc_score, roc_curve,
precision_recall_curve, average_precision_score,
f1_score, precision_score, recall_score)
from sklearn.ensemble import RandomForestClassifier, GradientBoostingClassifier
from sklearn.linear_model import LogisticRegression, Lasso, Ridge, ElasticNet
# ─── 0. Generate Imbalanced Sample Data ──────────────────────
X, y = make_classification(n_samples=1000, n_features=20, n_informative=10,
n_classes=2, weights=[0.7, 0.3], random_state=42)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, stratify=y, random_state=42
)
# ─── 1. Stratified K-Fold Cross-Validation ───────────────────
skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
model = RandomForestClassifier(n_estimators=100, random_state=42)
# Single metric
scores = cross_val_score(model, X_train, y_train, cv=skf, scoring='f1')
print(f"F1 (5-fold): {scores.mean():.4f} +/- {scores.std():.4f}")
# Multiple metrics at once
results = cross_validate(model, X_train, y_train, cv=skf,
scoring=['f1', 'roc_auc', 'precision', 'recall'])
for metric, vals in results.items():
if metric.startswith('test_'):
print(f" {metric}: {vals.mean():.4f} +/- {vals.std():.4f}")
# ─── 2. Regularisation — L1, L2, ElasticNet ──────────────────
from sklearn.datasets import make_regression
X_reg, y_reg = make_regression(n_samples=500, n_features=30, noise=0.1, random_state=42)
X_tr, X_te, y_tr, y_te = train_test_split(X_reg, y_reg, test_size=0.2, random_state=42)
# L1 (Lasso) — drives weak coefficients to exactly 0
lasso = Lasso(alpha=0.1)
lasso.fit(X_tr, y_tr)
print(f"Lasso zero coefficients: {(lasso.coef_ == 0).sum()}")
# L2 (Ridge) — shrinks all coefficients
ridge = Ridge(alpha=1.0)
ridge.fit(X_tr, y_tr)
# ElasticNet — combines L1 + L2
enet = ElasticNet(alpha=0.1, l1_ratio=0.5) # l1_ratio=1 → pure Lasso; =0 → pure Ridge
enet.fit(X_tr, y_tr)
# Logistic Regression with regularisation (classification)
lr_l1 = LogisticRegression(penalty='l1', C=0.1, solver='saga', max_iter=500)
lr_l2 = LogisticRegression(penalty='l2', C=0.1, solver='lbfgs', max_iter=500)
# Note: C = 1/lambda; smaller C = stronger regularisation
# ─── 3. Feature Scaling — Which Scaler When ──────────────────
from sklearn.preprocessing import StandardScaler, MinMaxScaler, RobustScaler
# StandardScaler: (x-mean)/std — for SVM, LR, KNN, PCA
sc = StandardScaler()
X_std = sc.fit_transform(X_train) # fit ONLY on train
X_tst = sc.transform(X_test) # transform test
# MinMaxScaler: (x-min)/(max-min) → [0,1] — for neural networks
mm = MinMaxScaler()
# RobustScaler: (x-median)/IQR — when outliers present
rb = RobustScaler()
# Tree-based models (RF, XGBoost) do NOT need scaling
# ─── 4. Missing Value Imputation ─────────────────────────────
from sklearn.impute import SimpleImputer, KNNImputer
# Median imputation (robust to outliers)
imputer = SimpleImputer(strategy='median')
X_imp = imputer.fit_transform(X_train)
# KNN imputation (uses feature relationships)
knn_imp = KNNImputer(n_neighbors=5)
X_knn = knn_imp.fit_transform(X_train)
# ─── 5. Class Imbalance — SMOTE + class_weight ───────────────
from imblearn.over_sampling import SMOTE
from imblearn.pipeline import Pipeline as ImbPipeline
print(f"Before SMOTE: {np.bincount(y_train)}")
sm = SMOTE(sampling_strategy='minority', random_state=42)
X_res, y_res = sm.fit_resample(X_train, y_train)
print(f"After SMOTE : {np.bincount(y_res)}")
# SMOTE inside CV-safe pipeline (correct approach)
imb_pipe = ImbPipeline([
('smote', SMOTE(random_state=42)),
('sc', StandardScaler()),
('clf', LogisticRegression(max_iter=500))
])
imb_pipe.fit(X_res, y_res)
print(classification_report(y_test, imb_pipe.predict(X_test)))
# ─── 6. All Classification Metrics ───────────────────────────
model.fit(X_train, y_train)
y_pred = model.predict(X_test)
y_proba = model.predict_proba(X_test)[:, 1]
print(classification_report(y_test, y_pred))
# Confusion matrix heatmap
cm = confusion_matrix(y_test, y_pred)
sns.heatmap(cm, annot=True, fmt='d', cmap='Blues',
xticklabels=['Neg','Pos'], yticklabels=['Neg','Pos'])
plt.ylabel('Actual'); plt.xlabel('Predicted'); plt.title('Confusion Matrix')
plt.tight_layout(); plt.show()
# ROC Curve
fpr, tpr, _ = roc_curve(y_test, y_proba)
auc = roc_auc_score(y_test, y_proba)
plt.figure(figsize=(6,5))
plt.plot(fpr, tpr, label=f'AUC = {auc:.4f}')
plt.plot([0,1],[0,1],'--', color='grey')
plt.xlabel('FPR'); plt.ylabel('TPR'); plt.title('ROC Curve'); plt.legend(); plt.show()
# Precision-Recall Curve (better than ROC for imbalanced data)
prec, rec, _ = precision_recall_curve(y_test, y_proba)
ap = average_precision_score(y_test, y_proba)
plt.figure(figsize=(6,5))
plt.plot(rec, prec, label=f'AP = {ap:.4f}')
plt.xlabel('Recall'); plt.ylabel('Precision'); plt.title('PR Curve'); plt.legend(); plt.show()
# ─── 7. GridSearchCV + RandomizedSearchCV ────────────────────
from scipy.stats import randint, uniform
# GridSearchCV — exhaustive
param_grid = {'n_estimators': [100, 200], 'max_depth': [3, 5, None],
'min_samples_leaf': [1, 5]}
grid = GridSearchCV(RandomForestClassifier(random_state=42),
param_grid, cv=skf, scoring='f1', n_jobs=-1, verbose=1)
grid.fit(X_train, y_train)
print(f"Grid best: {grid.best_params_} F1={grid.best_score_:.4f}")
# RandomizedSearchCV — faster, comparable quality
param_dist = {'n_estimators': randint(50, 500), 'max_depth': randint(2, 20),
'min_samples_leaf': randint(1, 20), 'max_features': uniform(0.3, 0.7)}
rand = RandomizedSearchCV(RandomForestClassifier(random_state=42),
param_dist, n_iter=50, cv=skf, scoring='f1',
n_jobs=-1, random_state=42)
rand.fit(X_train, y_train)
print(f"Rand best: {rand.best_params_} F1={rand.best_score_:.4f}")
# ─── 8. Feature Importance + SHAP ────────────────────────────
importances = model.feature_importances_
feature_names = [f'feat_{i}' for i in range(X.shape[1])]
indices = np.argsort(importances)[::-1]
plt.figure(figsize=(10, 5))
plt.bar(range(15), importances[indices[:15]])
plt.xticks(range(15), [feature_names[i] for i in indices[:15]], rotation=45, ha='right')
plt.title('Top 15 Feature Importances'); plt.tight_layout(); plt.show()
# SHAP (pip install shap)
import shap
explainer = shap.TreeExplainer(model)
shap_vals = explainer.shap_values(X_test)
shap.summary_plot(shap_vals[1], X_test, feature_names=feature_names)
# ─── 9. Learning Curves — Diagnose Bias vs Variance ──────────
def plot_learning_curve(estimator, X, y, cv=5, scoring='f1'):
sizes, tr_sc, va_sc = learning_curve(
estimator, X, y, cv=cv, scoring=scoring,
train_sizes=np.linspace(0.1, 1.0, 10), n_jobs=-1
)
tr_m, tr_s = tr_sc.mean(1), tr_sc.std(1)
va_m, va_s = va_sc.mean(1), va_sc.std(1)
plt.figure(figsize=(8,5))
plt.plot(sizes, tr_m, 'o-', color='royalblue', label='Train')
plt.fill_between(sizes, tr_m-tr_s, tr_m+tr_s, alpha=0.15, color='royalblue')
plt.plot(sizes, va_m, 'o-', color='tomato', label='Validation')
plt.fill_between(sizes, va_m-va_s, va_m+va_s, alpha=0.15, color='tomato')
plt.xlabel('Training Size'); plt.ylabel(scoring); plt.title('Learning Curve')
plt.legend(); plt.tight_layout(); plt.show()
# Interpretation:
# Large gap (train >> val): OVERFITTING → regularise, add data, reduce complexity
# Both low: UNDERFITTING → more features, more complex model
plot_learning_curve(RandomForestClassifier(n_estimators=100, random_state=42), X, y)
# ─── 10. Full Production ML Pipeline ─────────────────────────
from sklearn.compose import ColumnTransformer
numeric_features = [0, 1, 2, 3, 4] # column indices or names
categorical_features = [] # none in this example
numeric_transformer = Pipeline([
('imputer', SimpleImputer(strategy='median')),
('scaler', StandardScaler())
])
preprocessor = ColumnTransformer([
('num', numeric_transformer, numeric_features)
])
full_pipeline = Pipeline([
('preprocessor', preprocessor),
('classifier', GradientBoostingClassifier(random_state=42))
])
param_grid_full = {
'classifier__n_estimators': [100, 200],
'classifier__max_depth': [3, 5],
'classifier__learning_rate': [0.05, 0.1]
}
cv_full = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
grid_full = GridSearchCV(full_pipeline, param_grid_full,
cv=cv_full, scoring='roc_auc', n_jobs=-1)
grid_full.fit(X_train, y_train)
print(f"Best params : {grid_full.best_params_}")
print(f"Best AUC-ROC: {grid_full.best_score_:.4f}")
print(classification_report(y_test, grid_full.predict(X_test)))
# Save and load
import joblib
joblib.dump(grid_full.best_estimator_, 'best_pipeline.pkl')
loaded_model = joblib.load('best_pipeline.pkl')
preds = loaded_model.predict(X_test)
# ─── 11. Categorical Encoding ────────────────────────────────
import pandas as pd
df = pd.DataFrame({'city': ['NY','LA','NY','Chicago','LA'],
'size': ['S','M','L','M','XL'],
'target': [1,0,1,0,1]})
# One-Hot Encoding (nominal, low cardinality)
df_ohe = pd.get_dummies(df, columns=['city'], drop_first=True)
# Label Encoding (ordinal)
from sklearn.preprocessing import OrdinalEncoder
oe = OrdinalEncoder(categories=[['S','M','L','XL']])
df[['size_encoded']] = oe.fit_transform(df[['size']])
# Target Encoding (high cardinality — use inside CV to avoid leakage)
# pip install category_encoders
from category_encoders import TargetEncoder
te = TargetEncoder(cols=['city'])
df_te = te.fit_transform(df[['city']], df['target'])
⚙️ Key Parameters / Hyperparameters
| Parameter | What It Does | Typical Values |
|---|---|---|
LR: C | Inverse regularisation (lower = stronger) | 0.001, 0.01, 0.1, 1, 10 |
Lasso/Ridge: alpha | Regularisation strength (higher = stronger) | 0.001 to 100 (log scale) |
ElasticNet: l1_ratio | Mix of L1 vs L2 | 0 (Ridge) to 1 (Lasso) |
RF/GBM: n_estimators | Number of trees | 100, 200, 500 |
GBM: learning_rate | Step size per tree | 0.01, 0.05, 0.1 (smaller = better with more trees) |
GBM: max_depth | Tree depth per iteration | 3–6 |
GridSearchCV: scoring | Metric to optimise | 'f1', 'roc_auc', 'f1_weighted' |
GridSearchCV: cv | CV strategy | StratifiedKFold(5) |
RandomizedSearchCV: n_iter | How many param combos to try | 50–200 |
SMOTE: sampling_strategy | Ratio to resample to | 'minority', 0.5, 1.0 |
SimpleImputer: strategy | How to fill nulls | 'mean', 'median', 'most_frequent' |
🎤 Top Interview Q&A
Q1: Explain the bias-variance tradeoff. How does regularisation help?
A1: Bias = error from oversimplifying the pattern; variance = error from sensitivity to training data noise. A simple model (e.g., linear) has high bias, low variance — it underfits. A complex model (e.g., deep tree) has low bias, high variance — it overfits. Regularisation (L1/L2) adds a penalty on model complexity, reducing variance at the cost of slight bias increase.
Q2: What is data leakage and how do you prevent it?
A2: Leakage occurs when information from outside the training set (including the test set) influences the model. Examples: fitting the scaler on all data before splitting, applying SMOTE before the train/test split, or using future features. Prevention: always split first, then fit ALL preprocessing inside a Pipeline using only training data.
Q3: Why use Stratified K-Fold instead of regular K-Fold for classification?
A3: Regular K-Fold splits randomly, so some folds might have very few or zero samples of the minority class. Stratified K-Fold ensures each fold has the same class proportion as the full dataset — critical for imbalanced problems.
Q4: When would you use RandomizedSearchCV over GridSearchCV?
A4: When the hyperparameter space is large or when some parameters are continuous. RandomizedSearchCV samples a fixed number of random combinations (n_iter=50), making it far faster while often finding equally good results. Use GridSearchCV to fine-tune once you've narrowed the range.
Q5: How do you detect and fix overfitting?
A5: Detection: large gap between train and validation scores in learning curves or CV. Fixes: add regularisation (L1/L2), get more training data, reduce model complexity (fewer layers/trees/depth), use dropout (neural nets), or apply early stopping (boosting). Cross-validation must confirm improvement.
Q6: What is the difference between Bagging and Boosting?
A6: Bagging trains independent models on bootstrap samples and averages predictions — reduces variance (Random Forest). Boosting trains models sequentially, where each focuses on errors of the previous — reduces bias (XGBoost, LightGBM). Bagging is more robust to noise; Boosting typically achieves higher accuracy with proper tuning.
⚠️ Common Mistakes
- Scaling before train/test split — scaler learns test statistics, inflating performance; always split first, scale inside Pipeline.
- Evaluating on validation set and calling it "test set" — the true test set must be used exactly once at the very end.
- SMOTE on full data before CV — synthetic points from the same neighbourhood appear in both train and validation folds, causing leakage.
- Using accuracy for imbalanced data — 99% accuracy possible by predicting all majority class; use F1, AUC-ROC, or PR-AUC.
- Skipping baseline model — without
DummyClassifier, you don't know if your complex model is actually better than a trivial rule. - Feature selection outside CV — selecting features on all training data then cross-validating introduces target leakage from feature selection step.
- Not saving the full pipeline — if you save only the model weights, you lose scaler parameters and predictions on new data will be wrong.
🚀 Quick Reference — When to Use What
| Situation | Use This |
|---|---|
| Fast interpretable baseline | Logistic Regression with Pipeline |
| Max accuracy on tabular data | GradientBoostingClassifier or XGBoost |
| Need to detect/fix overfitting | Learning curve + regularisation (L1/L2) |
| Imbalanced classes | SMOTE inside Pipeline + class_weight='balanced' |
| Large hyperparameter space | RandomizedSearchCV first, then GridSearchCV |
| Need probability calibration | Logistic Regression or Platt scaling on SVM |
| Feature selection needed | L1 (Lasso) regularisation |
| Correlated features | L2 (Ridge) regularisation |
| Production model serving | joblib.dump(full_pipeline, 'model.pkl') |
| Understanding model decisions | SHAP TreeExplainer + summary_plot |
📋 Completion Checklist
- Explain bias-variance tradeoff; draw the U-shaped test error curve from memory
- Implement Stratified K-Fold CV with multiple metrics (
cross_validate) - Apply L1, L2, ElasticNet regularisation and explain what each does to coefficients
- Build a full
Pipelinewith imputer + scaler + model; confirm no data leakage - Handle class imbalance using SMOTE inside a pipeline-safe manner
- Run
GridSearchCVANDRandomizedSearchCV; compare results - Plot learning curves and correctly diagnose overfitting vs underfitting
- Plot ROC curve, PR curve, and confusion matrix heatmap; choose the right one per scenario