"The First Model" — scikit-learn, Pipelines and Honest Metrics
Sentinel's first fraud model. The estimator API, stratified train/test splits, Pipeline and ColumnTransformer, logistic regression vs random forests, why accuracy lies on imbalanced data, precision-recall and threshold tuning, cross-validation, data leakage, and persisting models.
Story Opening
Arjun trained his first model in an afternoon. He split the data, fitted a classifier, and proudly posted the result in the team channel: 96.9% accuracy.
Priya replied with a single line of code:
# A "model" that always answers "not fraud" — on data with a 3.1% fraud rate:fraud_rate = 0.031always_legit_accuracy = 1 - fraud_rateprint(f"{always_legit_accuracy:.1%}") # -> 96.9%His model was exactly as accurate as a function that returns False. It had learned to never flag anything — and on a dataset where 97% of transactions are legitimate, that’s the easiest way to look good.
This part is about building a model honestly: the scikit-learn API, pipelines that prevent subtle bugs, metrics that tell the truth on imbalanced data, and the leakage traps that make models look brilliant offline and fail in production.
Java → Python: The Mental Map
| Concept | Java-world analogy | scikit-learn |
|---|---|---|
| Model contract | An interface every model implements | The estimator API: fit, predict, predict_proba, transform |
| Learned state | Fields populated by an init() call | Attributes with a trailing underscore: coef_, mean_, classes_ |
| Processing chain | A chain of Functions / a builder | Pipeline |
| Per-column processing | A routing table of handlers | ColumnTransformer |
| Config | Constructor parameters | Hyperparameters, set in __init__, read with get_params() |
| Serialised artefact | A JAR / serialised object | joblib.dump → .joblib file |
| Unit test vs integration test | — | Validation set vs held-out test set |
scikit-learn’s power comes from that uniform interface — duck typing (Part 4) at ecosystem scale. Any object with fit and predict plugs into pipelines, grid search and cross-validation, including models from XGBoost, LightGBM or your own classes.
The Dataset
Every block in this part regenerates the same synthetic transactions with a small helper, so each one runs on its own. The fraud label depends on the features in a known way (unusual amounts for the customer, foreign cards, night-time, risky merchant categories), so there is real signal to learn — about 3% of transactions are fraudulent.
import numpy as npimport pandas as pd
def make_data(n=20_000, seed=7): """Synthetic card transactions with a ~3% fraud rate and a learnable signal.""" rng = np.random.default_rng(seed) df = pd.DataFrame({ "amount": rng.lognormal(4.0, 1.0, n).round(2), "hour": rng.integers(0, 24, n), "merchant": rng.choice(["grocery", "fuel", "electronics", "travel", "gaming"], n, p=[0.35, 0.25, 0.15, 0.15, 0.10]), "amount_vs_history": rng.lognormal(0.0, 0.6, n).round(3), # amount / customer's usual "is_foreign": (rng.random(n) < 0.08).astype(int), }) risk = {"grocery": 0.0, "fuel": 0.0, "electronics": 1.0, "travel": 0.6, "gaming": 1.2} logit = (-6.0 + 2.5 * np.log(df["amount_vs_history"]) + 2.0 * df["is_foreign"] + 2.2 * df["hour"].isin([0, 1, 2, 3, 4]).astype(int) + df["merchant"].map(risk) + 0.4 * (np.log(df["amount"]) - 4.0)) df["is_fraud"] = (rng.random(n) < 1 / (1 + np.exp(-logit))).astype(int) return df
df = make_data()print(df.shape) # -> (20000, 6)print(round(df["is_fraud"].mean(), 4)) # -> 0.0308print(df.head(3).to_string(index=False))The Estimator API in Five Lines
Before the full pipeline, look at the contract itself. Every scikit-learn object follows the same pattern: configure in the constructor, learn with fit, then transform (for preprocessors) or predict (for models).
import numpy as npfrom sklearn.preprocessing import StandardScalerfrom sklearn.linear_model import LogisticRegression
X = np.array([[100.0, 1], [200.0, 0], [300.0, 1], [400.0, 0]]) # 4 samples, 2 featuresy = np.array([0, 0, 1, 1])
scaler = StandardScaler() # 1. configure (hyperparameters only — no data yet)scaler.fit(X) # 2. LEARN state from dataprint(scaler.mean_) # -> [250. 0.5] (learned attributes end in '_')X_scaled = scaler.transform(X) # 3. APPLY the learned stateprint(X_scaled[:, 0].round(2)) # -> [-1.34 -0.45 0.45 1.34]
model = LogisticRegression().fit(X_scaled, y) # fit returns self, so calls chainprint(model.predict(X_scaled)) # -> [0 0 1 1]print(model.predict_proba(X_scaled).shape) # -> (4, 2) (one column per class)print(model.get_params()["C"]) # -> 1.0Tip —
fitvstransformis the most important distinction in this part.fitlearns from data (means, vocabulary, category lists, model weights).transform/predictapplies what was learned. You fit on training data only and apply the same fitted objects to everything else. Breaking this rule is the most common source of leakage.
Deep Dive: Splits, Pipelines and ColumnTransformer
Why a held-out test set
A model’s job is to perform on data it has never seen. So you hold back a test set before doing anything else, and look at it once, at the end — like a production smoke test you’re not allowed to tune against. For imbalanced classes, stratify the split so both halves have the same fraud rate.
Why a Pipeline
Real features need different preprocessing: numbers get log-transformed and scaled, categories get one-hot encoded. You could do that by hand with pandas — and then you’d have to repeat exactly the same steps, with exactly the same fitted parameters, at scoring time in production. A Pipeline bundles preprocessing and model into one estimator: one fit, one predict, one artefact to deploy. It also guarantees that preprocessing is fitted only on training data, even inside cross-validation.
import numpy as npimport pandas as pdfrom sklearn.compose import ColumnTransformerfrom sklearn.linear_model import LogisticRegressionfrom sklearn.metrics import average_precision_score, classification_report, confusion_matrix, roc_auc_scorefrom sklearn.model_selection import train_test_splitfrom sklearn.pipeline import Pipelinefrom sklearn.preprocessing import FunctionTransformer, OneHotEncoder, StandardScaler
def make_data(n=20_000, seed=7): rng = np.random.default_rng(seed) df = pd.DataFrame({ "amount": rng.lognormal(4.0, 1.0, n).round(2), "hour": rng.integers(0, 24, n), "merchant": rng.choice(["grocery", "fuel", "electronics", "travel", "gaming"], n, p=[0.35, 0.25, 0.15, 0.15, 0.10]), "amount_vs_history": rng.lognormal(0.0, 0.6, n).round(3), "is_foreign": (rng.random(n) < 0.08).astype(int), }) risk = {"grocery": 0.0, "fuel": 0.0, "electronics": 1.0, "travel": 0.6, "gaming": 1.2} logit = (-6.0 + 2.5 * np.log(df["amount_vs_history"]) + 2.0 * df["is_foreign"] + 2.2 * df["hour"].isin([0, 1, 2, 3, 4]).astype(int) + df["merchant"].map(risk) + 0.4 * (np.log(df["amount"]) - 4.0)) df["is_fraud"] = (rng.random(n) < 1 / (1 + np.exp(-logit))).astype(int) return df
df = make_data()X = df.drop(columns="is_fraud") # features: a DataFrame (scikit-learn accepts pandas)y = df["is_fraud"] # target: a Series
# 1. Hold out 25% for the final test, preserving the fraud rate in both halves.X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.25, stratify=y, random_state=42)print(round(y_train.mean(), 3), round(y_test.mean(), 3)) # -> 0.031 0.031
# 2. Route columns to the right preprocessing.preprocess = ColumnTransformer([ ("num", Pipeline([ ("log", FunctionTransformer(np.log1p)), # tame the long right tail ("scale", StandardScaler()), # mean 0, std 1 (Part 7's broadcasting) ]), ["amount", "amount_vs_history"]), ("cat", OneHotEncoder(handle_unknown="ignore"), ["merchant"]), # unseen category -> all zeros ("raw", "passthrough", ["is_foreign", "hour"]),])
# 3. One estimator: preprocessing + model.clf = Pipeline([ ("prep", preprocess), ("model", LogisticRegression(max_iter=1_000, class_weight="balanced")),])
clf.fit(X_train, y_train) # fits the scaler, encoder AND model — on train only
# 4. Evaluate on the untouched test set.proba = clf.predict_proba(X_test)[:, 1] # probability of class 1 (fraud)pred = clf.predict(X_test) # hard labels at the default 0.5 threshold
print(f"ROC AUC: {roc_auc_score(y_test, proba):.3f}") # -> ROC AUC: 0.902print(f"PR AUC: {average_precision_score(y_test, proba):.3f}") # -> PR AUC: 0.333print(confusion_matrix(y_test, pred))# [[3931 915]# [ 26 128]]print(classification_report(y_test, pred, digits=3))# ...class 1 (fraud): precision 0.123, recall 0.831Read the confusion matrix row by row: of 154 frauds in the test set, the model caught 128 (recall 83%) — but it also flagged 915 legitimate transactions, so only about one flag in eight is real fraud (precision 12%). Whether that’s acceptable is not a modelling question; it’s a business one, and we’ll answer it with a threshold below.
Tip —
class_weight="balanced"tells the model that each fraud example counts as much as ~32 legitimate ones (the inverse of their frequency). Without it, the easiest way to minimise loss on 3% positives is to predict “legit” almost always — Arjun’s first model.
Deep Dive: Metrics That Don’t Lie
With 97% negatives, accuracy is dominated by the easy majority. Fraud work lives in the confusion matrix:
| Predicted legit | Predicted fraud | |
|---|---|---|
| Actually legit | True negative | False positive — a customer annoyed, a review cost |
| Actually fraud | False negative — money lost | True positive — fraud caught |
From that come the metrics that matter:
- Recall (sensitivity): of all fraud, what fraction did we catch?
TP / (TP + FN) - Precision: of everything we flagged, what fraction was fraud?
TP / (TP + FP) - ROC AUC: probability the model ranks a random fraud above a random legit transaction. Threshold-independent, but flattering on imbalanced data.
- PR AUC (average precision): area under the precision–recall curve. The headline metric for rare-event detection. A random model scores roughly the positive rate (≈ 0.03 here), so 0.33 is a ~10× lift.
import numpy as npfrom sklearn.dummy import DummyClassifierfrom sklearn.metrics import accuracy_score, average_precision_score, recall_score
rng = np.random.default_rng(0)y = (rng.random(10_000) < 0.031).astype(int) # 3.1% positivesX = rng.normal(size=(10_000, 3)) # features irrelevant for a dummy
dummy = DummyClassifier(strategy="most_frequent").fit(X, y) # always predicts 0pred = dummy.predict(X)
print(f"accuracy: {accuracy_score(y, pred):.3f}") # -> accuracy: 0.969print(f"recall: {recall_score(y, pred):.3f}") # -> recall: 0.000print(f"PR AUC: {average_precision_score(y, dummy.predict_proba(X)[:, 1]):.3f}") # -> PR AUC: 0.031Gotcha — always compare against a dummy baseline.
DummyClassifiertells you what “no skill” scores on your metric. If your model barely beats it, you don’t have a model yet.
Comparing Models, and Feature Engineering Still Matters
Logistic regression is linear in its inputs: it can only learn “risk goes up (or down) as hour increases”. But fraud risk is high between midnight and 5 a.m. and flat otherwise — not a straight line. Tree ensembles like random forests learn such thresholds on their own. Alternatively, you engineer the right feature and keep the simpler, more explainable model.
import numpy as npimport pandas as pdfrom sklearn.compose import ColumnTransformerfrom sklearn.ensemble import HistGradientBoostingClassifier, RandomForestClassifierfrom sklearn.linear_model import LogisticRegressionfrom sklearn.metrics import average_precision_score, roc_auc_scorefrom sklearn.model_selection import train_test_splitfrom sklearn.pipeline import Pipelinefrom sklearn.preprocessing import FunctionTransformer, OneHotEncoder, StandardScaler
def make_data(n=20_000, seed=7): rng = np.random.default_rng(seed) df = pd.DataFrame({ "amount": rng.lognormal(4.0, 1.0, n).round(2), "hour": rng.integers(0, 24, n), "merchant": rng.choice(["grocery", "fuel", "electronics", "travel", "gaming"], n, p=[0.35, 0.25, 0.15, 0.15, 0.10]), "amount_vs_history": rng.lognormal(0.0, 0.6, n).round(3), "is_foreign": (rng.random(n) < 0.08).astype(int), }) risk = {"grocery": 0.0, "fuel": 0.0, "electronics": 1.0, "travel": 0.6, "gaming": 1.2} logit = (-6.0 + 2.5 * np.log(df["amount_vs_history"]) + 2.0 * df["is_foreign"] + 2.2 * df["hour"].isin([0, 1, 2, 3, 4]).astype(int) + df["merchant"].map(risk) + 0.4 * (np.log(df["amount"]) - 4.0)) df["is_fraud"] = (rng.random(n) < 1 / (1 + np.exp(-logit))).astype(int) return df
def add_is_night(X: pd.DataFrame) -> pd.DataFrame: """Domain feature: 1 for transactions between 00:00 and 04:59.""" return X.assign(is_night=X["hour"].between(0, 4).astype(int))
df = make_data()X, y = df.drop(columns="is_fraud"), df["is_fraud"]X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, stratify=y, random_state=42)
def build(model, engineered=False): passthrough = ["is_foreign", "hour"] + (["is_night"] if engineered else []) steps = [("features", FunctionTransformer(add_is_night))] if engineered else [] steps += [ ("prep", ColumnTransformer([ ("num", Pipeline([("log", FunctionTransformer(np.log1p)), ("scale", StandardScaler())]), ["amount", "amount_vs_history"]), ("cat", OneHotEncoder(handle_unknown="ignore"), ["merchant"]), ("raw", "passthrough", passthrough), ])), ("model", model), ] return Pipeline(steps)
candidates = { "logreg (raw hour)": build(LogisticRegression(max_iter=1_000, class_weight="balanced")), "logreg (+ is_night)": build(LogisticRegression(max_iter=1_000, class_weight="balanced"), engineered=True), "random forest": build(RandomForestClassifier(n_estimators=300, min_samples_leaf=5, class_weight="balanced_subsample", n_jobs=-1, random_state=0)), "gradient boosting": build(HistGradientBoostingClassifier(learning_rate=0.05, max_iter=300, class_weight="balanced", random_state=0)),}
for name, pipe in candidates.items(): proba = pipe.fit(X_train, y_train).predict_proba(X_test)[:, 1] print(f"{name:22s} ROC AUC={roc_auc_score(y_test, proba):.3f} PR AUC={average_precision_score(y_test, proba):.3f}")Expected output (exact figures can vary slightly across library versions):
logreg (raw hour) ROC AUC=0.902 PR AUC=0.333logreg (+ is_night) ROC AUC=0.917 PR AUC=0.398random forest ROC AUC=0.904 PR AUC=0.353gradient boosting ROC AUC=0.907 PR AUC=0.299One honest domain feature lifted the linear model past both ensembles. That’s a recurring lesson in applied ML: features usually matter more than the algorithm — and the simpler model is easier to explain to a regulator.
Tip — Which model first? For tabular data: start with logistic regression as an explainable baseline, then try gradient-boosted trees (
HistGradientBoostingClassifier, or XGBoost/LightGBM). Deep learning (Part 10) rarely wins on tabular data; it shines on text, images, audio and sequences.
Deep Dive: The Threshold Is a Business Decision
predict() uses a 0.5 probability threshold. That number has nothing to do with your business. The right threshold depends on what errors cost. Suppose a missed fraud costs ₹5,000 on average, and a manual review of a flagged transaction costs ₹150. Then you choose the threshold that minimises total cost on validation data:
import numpy as npimport pandas as pdfrom sklearn.compose import ColumnTransformerfrom sklearn.linear_model import LogisticRegressionfrom sklearn.metrics import precision_recall_curvefrom sklearn.model_selection import train_test_splitfrom sklearn.pipeline import Pipelinefrom sklearn.preprocessing import FunctionTransformer, OneHotEncoder, StandardScaler
def make_data(n=20_000, seed=7): rng = np.random.default_rng(seed) df = pd.DataFrame({ "amount": rng.lognormal(4.0, 1.0, n).round(2), "hour": rng.integers(0, 24, n), "merchant": rng.choice(["grocery", "fuel", "electronics", "travel", "gaming"], n, p=[0.35, 0.25, 0.15, 0.15, 0.10]), "amount_vs_history": rng.lognormal(0.0, 0.6, n).round(3), "is_foreign": (rng.random(n) < 0.08).astype(int), }) risk = {"grocery": 0.0, "fuel": 0.0, "electronics": 1.0, "travel": 0.6, "gaming": 1.2} logit = (-6.0 + 2.5 * np.log(df["amount_vs_history"]) + 2.0 * df["is_foreign"] + 2.2 * df["hour"].isin([0, 1, 2, 3, 4]).astype(int) + df["merchant"].map(risk) + 0.4 * (np.log(df["amount"]) - 4.0)) df["is_fraud"] = (rng.random(n) < 1 / (1 + np.exp(-logit))).astype(int) return df
df = make_data().assign(is_night=lambda d: d["hour"].between(0, 4).astype(int))X, y = df.drop(columns="is_fraud"), df["is_fraud"]
# Three-way split: train (fit), validation (choose threshold), test (final report).X_tmp, X_test, y_tmp, y_test = train_test_split(X, y, test_size=0.2, stratify=y, random_state=1)X_train, X_val, y_train, y_val = train_test_split(X_tmp, y_tmp, test_size=0.25, stratify=y_tmp, random_state=1)
clf = Pipeline([ ("prep", ColumnTransformer([ ("num", Pipeline([("log", FunctionTransformer(np.log1p)), ("scale", StandardScaler())]), ["amount", "amount_vs_history"]), ("cat", OneHotEncoder(handle_unknown="ignore"), ["merchant"]), ("raw", "passthrough", ["is_foreign", "is_night"]), ])), ("model", LogisticRegression(max_iter=1_000, class_weight="balanced")),]).fit(X_train, y_train)
COST_MISSED_FRAUD = 5_000 # ₹ lost per false negativeCOST_REVIEW = 150 # ₹ per flagged transaction (true or false positive)
def total_cost(y_true, proba, threshold): flagged = proba >= threshold missed = ((~flagged) & (y_true == 1)).sum() return missed * COST_MISSED_FRAUD + flagged.sum() * COST_REVIEW
val_proba = clf.predict_proba(X_val)[:, 1]thresholds = np.linspace(0.05, 0.95, 91)costs = [total_cost(y_val.to_numpy(), val_proba, t) for t in thresholds]best = thresholds[int(np.argmin(costs))]print(f"best threshold on validation: {best:.2f}") # e.g. 0.54
# Report on the untouched test set — once.test_proba = clf.predict_proba(X_test)[:, 1]for t in (0.5, best): flagged = test_proba >= t caught = (flagged & (y_test == 1)).sum() print(f"t={t:.2f}: flagged={flagged.sum():4d} caught={caught}/{y_test.sum()} " f"cost=₹{total_cost(y_test.to_numpy(), test_proba, t):,}")# t=0.50: flagged= 772 caught=98/123 cost=₹240,800# t=0.54: flagged= 683 caught=96/123 cost=₹237,450
# The precision/recall trade-off at every threshold, for the curious:precision, recall, pr_thresholds = precision_recall_curve(y_val, val_proba)print(len(precision) == len(recall) == len(pr_thresholds) + 1) # -> TrueGotcha — never pick a threshold (or anything else) on the test set. Every decision made by looking at test results leaks information into the model-building process, and your reported numbers become optimistic. Use a validation split or cross-validation for decisions; touch the test set once.
Cross-Validation and Hyperparameter Search
A single validation split is noisy, especially with only a few hundred positives. K-fold cross-validation trains K models, each validated on a different fold, and averages the scores. GridSearchCV wraps that around a hyperparameter search. Because the whole pipeline is the estimator, preprocessing is re-fitted inside every fold — no leakage.
import numpy as npimport pandas as pdfrom sklearn.compose import ColumnTransformerfrom sklearn.linear_model import LogisticRegressionfrom sklearn.model_selection import GridSearchCV, StratifiedKFold, cross_val_scorefrom sklearn.pipeline import Pipelinefrom sklearn.preprocessing import FunctionTransformer, OneHotEncoder, StandardScaler
def make_data(n=20_000, seed=7): rng = np.random.default_rng(seed) df = pd.DataFrame({ "amount": rng.lognormal(4.0, 1.0, n).round(2), "hour": rng.integers(0, 24, n), "merchant": rng.choice(["grocery", "fuel", "electronics", "travel", "gaming"], n, p=[0.35, 0.25, 0.15, 0.15, 0.10]), "amount_vs_history": rng.lognormal(0.0, 0.6, n).round(3), "is_foreign": (rng.random(n) < 0.08).astype(int), }) risk = {"grocery": 0.0, "fuel": 0.0, "electronics": 1.0, "travel": 0.6, "gaming": 1.2} logit = (-6.0 + 2.5 * np.log(df["amount_vs_history"]) + 2.0 * df["is_foreign"] + 2.2 * df["hour"].isin([0, 1, 2, 3, 4]).astype(int) + df["merchant"].map(risk) + 0.4 * (np.log(df["amount"]) - 4.0)) df["is_fraud"] = (rng.random(n) < 1 / (1 + np.exp(-logit))).astype(int) return df
df = make_data().assign(is_night=lambda d: d["hour"].between(0, 4).astype(int))X, y = df.drop(columns="is_fraud"), df["is_fraud"]
pipe = Pipeline([ ("prep", ColumnTransformer([ ("num", Pipeline([("log", FunctionTransformer(np.log1p)), ("scale", StandardScaler())]), ["amount", "amount_vs_history"]), ("cat", OneHotEncoder(handle_unknown="ignore"), ["merchant"]), ("raw", "passthrough", ["is_foreign", "is_night"]), ])), ("model", LogisticRegression(max_iter=1_000, class_weight="balanced")),])
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=0)
scores = cross_val_score(pipe, X, y, cv=cv, scoring="average_precision", n_jobs=-1)print(f"PR AUC: {scores.mean():.3f} ± {scores.std():.3f}") # e.g. PR AUC: 0.371 ± 0.034
# Hyperparameters of nested steps are addressed as <step>__<param> (double underscore).search = GridSearchCV( pipe, param_grid={"model__C": [0.01, 0.1, 1.0, 10.0]}, # inverse regularisation strength scoring="average_precision", cv=cv, n_jobs=-1,)search.fit(X, y)print(search.best_params_) # e.g. {'model__C': 10.0}print(f"best CV PR AUC: {search.best_score_:.3f}")Tip — time matters for fraud. Random K-fold lets a model train on September and validate on August. For temporal data, prefer
TimeSeriesSplitor a simple “train on the past, test on the most recent weeks” split. It’s a more honest preview of production, where the model always predicts the future.
Deep Dive: Data Leakage — When a Model Is Too Good
Leakage is any information in training that won’t be available at prediction time. It produces spectacular offline metrics and useless production models. It is the most expensive mistake in applied ML, and it’s almost never a code bug — it’s a design bug.
import numpy as npimport pandas as pdfrom sklearn.ensemble import HistGradientBoostingClassifierfrom sklearn.metrics import average_precision_scorefrom sklearn.model_selection import train_test_split
rng = np.random.default_rng(3)n = 20_000df = pd.DataFrame({ "amount_vs_history": rng.lognormal(0, 0.6, n), "is_foreign": (rng.random(n) < 0.08).astype(int),})logit = -5.0 + 2.5 * np.log(df["amount_vs_history"]) + 2.0 * df["is_foreign"]df["is_fraud"] = (rng.random(n) < 1 / (1 + np.exp(-logit))).astype(int)
# A column from the case-management system: was a chargeback filed?# Chargebacks are filed AFTER fraud happens, weeks later. At scoring time it's always unknown.df["chargeback_filed"] = ((df["is_fraud"] == 1) & (rng.random(n) < 0.9)).astype(int)
def evaluate(columns): X_tr, X_te, y_tr, y_te = train_test_split(df[columns], df["is_fraud"], test_size=0.25, stratify=df["is_fraud"], random_state=0) model = HistGradientBoostingClassifier(random_state=0).fit(X_tr, y_tr) return average_precision_score(y_te, model.predict_proba(X_te)[:, 1])
honest = evaluate(["amount_vs_history", "is_foreign"])leaky = evaluate(["amount_vs_history", "is_foreign", "chargeback_filed"])print(f"honest PR AUC: {honest:.2f}")print(f"leaky PR AUC: {leaky:.2f} <- 'amazing', and worthless in production")print(leaky > honest + 0.4) # -> True| Leak | How it happens | Prevention |
|---|---|---|
| Target leakage | A feature encodes the outcome (chargebacks, “account closed”, investigator notes) | For each feature ask: would I know this at the moment of scoring? |
| Temporal leakage | Features computed using future rows (Part 8’s transform("mean") over all rows) | Point-in-time features: shift, expanding, rolling on sorted data |
| Preprocessing leakage | Scaling/encoding fitted on the full dataset before splitting | Put preprocessing in a Pipeline; split first |
| Duplicate leakage | Same customer or near-identical rows in train and test | Group-aware splits: GroupKFold, GroupShuffleSplit by customer |
Tip — be suspicious of great results. If a first model scores PR AUC 0.98 on a hard problem, the right reaction isn’t celebration — it’s to inspect feature importances and hunt for the leak.
Persisting and Serving the Model
import joblibimport numpy as npimport pandas as pdfrom sklearn.compose import ColumnTransformerfrom sklearn.linear_model import LogisticRegressionfrom sklearn.pipeline import Pipelinefrom sklearn.preprocessing import OneHotEncoder, StandardScaler
# A tiny trained pipeline (any pipeline from this part works the same way).train = pd.DataFrame({ "amount_vs_history": [0.5, 1.0, 1.2, 4.0, 6.0, 0.8, 5.5, 0.9], "merchant": ["grocery", "fuel", "grocery", "gaming", "electronics", "fuel", "gaming", "grocery"], "is_fraud": [0, 0, 0, 1, 1, 0, 1, 0],})pipe = Pipeline([ ("prep", ColumnTransformer([ ("num", StandardScaler(), ["amount_vs_history"]), ("cat", OneHotEncoder(handle_unknown="ignore"), ["merchant"]), ])), ("model", LogisticRegression()),]).fit(train.drop(columns="is_fraud"), train["is_fraud"])
# Save: ONE artefact containing preprocessing + model.joblib.dump(pipe, "sentinel_v1.joblib")
# Load (in the scoring service) and predict on a raw, single transaction.loaded = joblib.load("sentinel_v1.joblib")incoming = {"amount_vs_history": 5.0, "merchant": "crypto"} # unseen merchant category!score = loaded.predict_proba(pd.DataFrame([incoming]))[0, 1]print(0.5 < score <= 1.0) # -> True (handle_unknown="ignore" kept it safe)Gotcha — pickles are code.
joblib/picklefiles can execute arbitrary code when loaded: never load one from an untrusted source. They’re also tied to library versions, so pin scikit-learn exactly in the serving environment. For cross-language serving (say, scoring from a Java service), export to ONNX withskl2onnxand run it with ONNX Runtime’s Java API — Part 10 discusses the options.
Visualising the Trade-off
A precision–recall curve shows every threshold at once. Two models, one plot, saved to a file:
import matplotlibmatplotlib.use("Agg") # render to file; drop this line in notebooksimport matplotlib.pyplot as pltimport numpy as npfrom sklearn.datasets import make_classificationfrom sklearn.ensemble import RandomForestClassifierfrom sklearn.linear_model import LogisticRegressionfrom sklearn.metrics import PrecisionRecallDisplayfrom sklearn.model_selection import train_test_split
X, y = make_classification(n_samples=20_000, n_features=8, n_informative=4, weights=[0.97], flip_y=0.01, random_state=0)X_tr, X_te, y_tr, y_te = train_test_split(X, y, stratify=y, random_state=0)
fig, ax = plt.subplots(figsize=(6, 5.5))models = { "Logistic regression": (LogisticRegression(max_iter=1_000), "#2a78d6"), "Random forest": (RandomForestClassifier(n_estimators=200, random_state=0, n_jobs=-1), "#eb6834"),}for name, (model, color) in models.items(): model.fit(X_tr, y_tr) PrecisionRecallDisplay.from_estimator(model, X_te, y_te, name=name, ax=ax, color=color, linewidth=2)
ax.axhline(y_te.mean(), color="#8a8a85", linestyle="--", linewidth=1, label="No-skill baseline")ax.set_title("Precision–recall on held-out data")ax.grid(alpha=0.25)ax.spines[["top", "right"]].set_visible(False)ax.legend(frameon=False, loc="upper center", bbox_to_anchor=(0.5, -0.15), ncol=2) # below the plot, clear of the curvesfig.tight_layout()fig.savefig("pr_curve.png", dpi=150)print("saved pr_curve.png") # -> saved pr_curve.pngYour Own Estimator: Duck Typing at Work
Because scikit-learn relies on the estimator protocol, your own classes plug straight into pipelines. Inheriting BaseEstimator gives you get_params/set_params (needed for grid search and cloning), and TransformerMixin gives you fit_transform for free.
import numpy as npimport pandas as pdfrom sklearn.base import BaseEstimator, TransformerMixin
class AmountRatio(BaseEstimator, TransformerMixin): """Adds amount / median amount of the merchant, learned at fit time."""
def __init__(self, column="amount", group="merchant"): self.column = column # store params unchanged: sklearn convention self.group = group
def fit(self, X: pd.DataFrame, y=None): self.medians_ = X.groupby(self.group)[self.column].median() # learned state self.global_median_ = X[self.column].median() return self # always return self
def transform(self, X: pd.DataFrame) -> pd.DataFrame: medians = X[self.group].map(self.medians_).fillna(self.global_median_) return X.assign(amount_ratio=X[self.column] / medians)
train = pd.DataFrame({"merchant": ["a", "a", "b", "b"], "amount": [10.0, 30.0, 100.0, 300.0]})new = pd.DataFrame({"merchant": ["a", "b", "zzz"], "amount": [40.0, 100.0, 65.0]})
ratio = AmountRatio().fit(train)print(ratio.transform(new)["amount_ratio"].tolist()) # -> [2.0, 0.5, 1.0]print(ratio.get_params()) # -> {'column': 'amount', 'group': 'merchant'}Tips, Tricks & Gotchas
Tip —
set_config(transform_output="pandas")makes transformers return DataFrames with readable column names instead of bare NumPy arrays. Invaluable for debugging pipelines;pipe[:-1].get_feature_names_out()lists the final features.
Gotcha —
fit_transformon test data. Callingscaler.fit_transform(X_test)re-learns statistics from the test set. Alwaysfit(orfit_transform) on train and onlytransformon everything else — or let aPipelinehandle it.
Tip — explainability.
LogisticRegression.coef_gives signed feature weights;sklearn.inspection.permutation_importanceworks for any model. For per-prediction explanations (“why was this transaction flagged?”), look at SHAP. Fraud teams and regulators will ask.
Gotcha — probabilities aren’t always calibrated. A random forest’s
0.8doesn’t mean an 80% chance of fraud. If downstream systems interpret scores as probabilities, wrap the model inCalibratedClassifierCV.
Key Takeaways
| Concept | Remember |
|---|---|
| Estimator API | fit learns, transform/predict apply; learned attributes end in _ |
| Splits | Stratified train/validation/test; touch test once |
| Pipelines | Preprocessing + model as one estimator; prevents leakage; one artefact |
ColumnTransformer | Route columns to scalers, encoders, passthrough |
| Metrics | Accuracy lies on imbalance; use PR AUC, precision, recall, a dummy baseline |
| Thresholds | Choose from business costs on validation data, not 0.5 |
| Cross-validation | StratifiedKFold, GridSearchCV with step__param; time-aware splits for temporal data |
| Leakage | Ask “would I know this at scoring time?”; be suspicious of great scores |
| Persistence | joblib for Python serving (trusted sources only); ONNX for cross-language |
Story Closing
Sentinel v1 shipped as a logistic regression with eleven features, a cost-optimised threshold, and a model card documenting every feature’s point-in-time logic. In its first month in shadow mode, it caught most confirmed fraud at a review volume the analysts could actually staff — and every flag came with the feature values that drove it. The fraud team bought Arjun coffee for a week.
Then came the next request. The fraud analysts wanted the model to read the free-text descriptions merchants attach to transactions, and the product team wanted an assistant that could explain flagged cases in plain language. That meant neural networks, embeddings and LLM APIs — and a whole new set of questions for a Java architect: what’s a tensor? Why won’t Python threads use all my cores? How do I call an LLM API five hundred times concurrently?
In Part 10, the final part, Arjun dives into PyTorch, concurrency and the AI toolkit — and figures out how Python models meet JVM services in production.
This is Part 9 of a 10-part series: “Python for Java Developers: From Streams to Tensors.”