Story Opening

Two requests landed on Arjun’s desk in the same week.

The fraud analysts wanted a neural network that combined transaction features with the free-text descriptions merchants attach to payments. And the product team wanted “Sentinel Assist”: for every flagged transaction, call an LLM to draft a plain-language explanation for the reviewer — about five hundred calls per batch.

Arjun’s first attempt at the second request was a plain loop. Each LLM call took about two seconds. Five hundred calls: seventeen minutes. His Java instincts said thread pool. His Python research said something about a “GIL” that made threads useless. Both were half right.

“Different problems, different tools,” said Priya. “For the network, you need PyTorch. For the LLM calls, you need asyncio. And you need to know which is which.”

This final part covers the deep-learning core you need to read and write PyTorch, then the concurrency model every Java developer must relearn in Python, and finally the patterns that glue AI services into production systems.


Java → Python: The Quick Map

Java worldPython world
ND4J / DJL NDArraytorch.Tensor (NumPy-like, GPU-capable, differentiable)
Hand-derived gradientsAutograd: loss.backward()
A model class with forward()torch.nn.Module
Iterator<Batch>torch.utils.data.DataLoader
ExecutorService (threads)concurrent.futures.ThreadPoolExecutor — for I/O only (the GIL)
Separate JVM processesProcessPoolExecutor / multiprocessing — for CPU-bound work
CompletableFuture, virtual threadsasyncio coroutines: async def / await
Semaphoreasyncio.Semaphore
ONNX Runtime Java, DJLThe bridge from Python-trained models to JVM services

Tensors: NumPy That Can Learn

A PyTorch tensor is an n-dimensional array like NumPy’s, with two extra powers: it can live on a GPU, and it can record the operations applied to it so gradients can be computed automatically. Almost everything you learned in Part 7 — dtypes, shapes, broadcasting, indexing, axes (called dim here) — carries over.

import numpy as np
import torch
x = torch.tensor([[1.0, 2.0], [3.0, 4.0]])
print(x.shape, x.dtype) # -> torch.Size([2, 2]) torch.float32 (float32 by default, not float64)
print((x * 2 + 1).tolist()) # -> [[3.0, 5.0], [7.0, 9.0]] (vectorised, like NumPy)
print(x.sum(dim=0).tolist()) # -> [4.0, 6.0] ('dim' plays the role of NumPy's 'axis')
print((x @ x.T).tolist()) # -> [[5.0, 11.0], [11.0, 25.0]] (matrix multiply)
print(x[:, 1].tolist()) # -> [2.0, 4.0]
z = torch.zeros(3, 4)
print(z.reshape(2, -1).shape) # -> torch.Size([2, 6])
print(torch.rand(2, 3).unsqueeze(0).shape) # -> torch.Size([1, 2, 3]) (add a batch dimension)
# NumPy interop: from_numpy and .numpy() SHARE memory (on CPU) — Part 7's views, again.
arr = np.array([1.0, 2.0, 3.0])
t = torch.from_numpy(arr)
t[0] = 99.0
print(arr[0]) # -> 99.0 (same buffer!)
safe = torch.tensor(arr) # torch.tensor() always copies
print(safe.dtype) # -> torch.float64 (dtype preserved from NumPy)
# Devices: move data and models with .to(device). Code stays the same on CPU or GPU.
device = "cuda" if torch.cuda.is_available() else ("mps" if torch.backends.mps.is_available() else "cpu")
print(x.to(device).device.type in {"cpu", "cuda", "mps"}) # -> True

Gotcha — dtype mismatches. NumPy defaults to float64; PyTorch models default to float32. Feeding a float64 tensor into a float32 model raises RuntimeError: mat1 and mat2 must have the same dtype. Convert at the boundary: torch.tensor(arr, dtype=torch.float32).

Gotcha — device mismatches. All tensors in one operation must be on the same device. “Expected all tensors to be on the same device” means you moved the model to the GPU but not the batch (or vice versa).


Deep Dive: Autograd — Gradients for Free

Training a neural network means adjusting its parameters to reduce a loss. To know which direction to adjust, you need the gradient of the loss with respect to every parameter. Deriving those by hand for millions of parameters is impossible; autograd does it automatically.

When a tensor has requires_grad=True, PyTorch records every operation applied to it in a computation graph. Calling .backward() on a scalar result walks that graph in reverse (the chain rule — “backpropagation”) and stores each gradient in the leaf tensor’s .grad.

import torch
w = torch.tensor(3.0, requires_grad=True) # a "parameter"
x = torch.tensor(2.0) # an input (no gradient needed)
y = w * x ** 2 + 5 * w # y = w·x² + 5w (graph recorded)
y.backward() # dy/dw = x² + 5 = 9
print(y.item()) # -> 27.0
print(w.grad.item()) # -> 9.0

Now the whole learning loop in miniature. We’ll learn the slope and intercept of a line from noisy data using nothing but autograd and gradient descent:

import torch
torch.manual_seed(0)
x = torch.linspace(0, 10, 200)
y = 2.5 * x + 4.0 + torch.randn(200) # true slope 2.5, intercept 4.0, plus noise
w = torch.zeros(1, requires_grad=True) # start from nothing
b = torch.zeros(1, requires_grad=True)
lr = 0.01 # learning rate: step size
for step in range(2_000):
pred = w * x + b # 1. forward pass
loss = ((pred - y) ** 2).mean() # 2. loss: mean squared error
loss.backward() # 3. backward pass: fills w.grad and b.grad
with torch.no_grad(): # 4. update WITHOUT recording these ops
w -= lr * w.grad
b -= lr * b.grad
w.grad.zero_() # 5. reset: gradients ACCUMULATE by default
b.grad.zero_()
print(round(w.item(), 1), round(b.item(), 1)) # -> 2.5 4.1 (close to the true 2.5 and 4.0; noise shifts the fit slightly)

Every neural network you will ever train runs these five steps: forward → loss → backward → update → zero gradients. The libraries add layers, optimisers and data loading around them, but the skeleton never changes.

Gotcha — forgetting zero_grad(). Gradients accumulate across backward() calls (useful for some advanced tricks). Forget to reset them and each step uses the sum of all previous gradients — training diverges or behaves bizarrely.


A Real Training Loop: nn.Module, DataLoader, Optimiser

Here is the standard shape of PyTorch training code — the shape you’ll recognise in every repository and paper implementation. We train a small multi-layer perceptron (MLP) on synthetic fraud data.

import numpy as np
import torch
from torch import nn
from torch.utils.data import DataLoader, TensorDataset
from sklearn.metrics import average_precision_score
torch.manual_seed(0)
rng = np.random.default_rng(0)
# --- 1. Data: synthetic features with a non-linear fraud signal ----------------
n, n_features = 20_000, 6
X = rng.normal(size=(n, n_features)).astype(np.float32)
logit = -4.0 + 1.5 * X[:, 0] + 2.0 * (X[:, 1] * X[:, 2] > 1.0) + np.abs(X[:, 3])
y = (rng.random(n) < 1 / (1 + np.exp(-logit))).astype(np.float32)
split = int(0.8 * n)
train_ds = TensorDataset(torch.from_numpy(X[:split]), torch.from_numpy(y[:split]))
test_X, test_y = torch.from_numpy(X[split:]), y[split:]
# DataLoader: an iterator of shuffled mini-batches (Part 5's lazy iteration, for tensors).
train_loader = DataLoader(train_ds, batch_size=256, shuffle=True)
# --- 2. Model: subclass nn.Module, declare layers in __init__, wire them in forward ---
class FraudMLP(nn.Module):
def __init__(self, n_in: int, hidden: int = 32):
super().__init__()
self.net = nn.Sequential(
nn.Linear(n_in, hidden), # weights and biases are registered automatically
nn.ReLU(),
nn.Dropout(0.1), # behaves differently in train vs eval mode
nn.Linear(hidden, hidden),
nn.ReLU(),
nn.Linear(hidden, 1), # one output: a raw score ("logit")
)
def forward(self, x: torch.Tensor) -> torch.Tensor:
return self.net(x).squeeze(-1) # (batch, 1) -> (batch,)
model = FraudMLP(n_features)
print(sum(p.numel() for p in model.parameters())) # -> 1313 (trainable parameters)
# --- 3. Loss and optimiser ------------------------------------------------------
pos_weight = torch.tensor((1 - y.mean()) / y.mean()) # up-weight the rare class (Part 9)
loss_fn = nn.BCEWithLogitsLoss(pos_weight=pos_weight) # sigmoid + binary cross-entropy, numerically stable
optimizer = torch.optim.AdamW(model.parameters(), lr=3e-3)
# --- 4. Train ---------------------------------------------------------------------
for epoch in range(8):
model.train() # enable dropout etc.
running = 0.0
for xb, yb in train_loader:
optimizer.zero_grad() # step 5 from the mini loop, done first by convention
loss = loss_fn(model(xb), yb) # model(xb) calls __call__ -> forward (Part 3's callables)
loss.backward()
optimizer.step() # the optimiser applies the update rule
running += loss.item() * len(xb)
if epoch % 2 == 1:
print(f"epoch {epoch + 1}: loss {running / len(train_ds):.4f}")
# --- 5. Evaluate ------------------------------------------------------------------
model.eval() # disable dropout
with torch.no_grad(): # no graph recording: faster, less memory
scores = torch.sigmoid(model(test_X)).numpy() # logits -> probabilities
print(f"test PR AUC: {average_precision_score(test_y, scores):.3f} (base rate {test_y.mean():.3f})")
torch.save(model.state_dict(), "fraud_mlp.pt") # save the WEIGHTS, not the pickled object

Tip — model.train() and model.eval() don’t train or evaluate anything. They flip a mode flag that changes how layers like Dropout and BatchNorm behave. Forgetting eval() before inference gives noisy, non-deterministic predictions.

Tip — torch.no_grad() is a context manager (Part 5) that switches graph recording off and restores it afterwards. Use it (or torch.inference_mode()) for every evaluation and inference path.

StepCodeWhy
Reset gradientsoptimizer.zero_grad()Gradients accumulate
Forwardout = model(xb)Compute predictions
Lossloss = loss_fn(out, yb)One scalar to minimise
Backwardloss.backward()Autograd computes all gradients
Updateoptimizer.step()Apply the optimiser’s rule (AdamW, SGD…)

Deep Dive: The GIL, Threads and Processes

Here’s where Java developers get burned. CPython has a Global Interpreter Lock (GIL): only one thread executes Python bytecode at a time. It exists because CPython’s memory management (reference counting) isn’t thread-safe without it.

The consequences:

  • CPU-bound pure-Python code doesn’t get faster with threads. Four threads computing in Python take as long as one (or longer).
  • I/O-bound code does benefit from threads. A thread waiting on a socket, file or time.sleep releases the GIL, so other threads run.
  • C extensions release the GIL. NumPy, pandas internals and PyTorch kernels run their heavy loops outside the GIL — which is why vectorised code (Parts 7–8) can use multiple cores internally.
import time
from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor
def cpu_work(n: int) -> int:
"""Pure-Python CPU-bound work: holds the GIL the whole time."""
total = 0
for i in range(n):
total += i * i % 7
return total
def io_work(_: int) -> str:
"""I/O-bound work: sleeping (like waiting on a network call) releases the GIL."""
time.sleep(0.2)
return "ok"
def timed(label, fn):
start = time.perf_counter()
fn()
print(f"{label:28s} {time.perf_counter() - start:.2f}s")
if __name__ == "__main__": # REQUIRED for ProcessPoolExecutor on macOS/Windows
jobs = [2_000_000] * 4
timed("CPU, sequential", lambda: [cpu_work(n) for n in jobs])
with ThreadPoolExecutor(max_workers=4) as pool: # ≈ Executors.newFixedThreadPool(4)
timed("CPU, 4 threads (GIL!)", lambda: list(pool.map(cpu_work, jobs)))
with ProcessPoolExecutor(max_workers=4) as pool: # 4 interpreters, 4 GILs
timed("CPU, 4 processes", lambda: list(pool.map(cpu_work, jobs)))
timed("I/O, sequential (8 x 0.2s)", lambda: [io_work(i) for i in range(8)])
with ThreadPoolExecutor(max_workers=8) as pool:
timed("I/O, 8 threads", lambda: list(pool.map(io_work, range(8))))
# Typical output on a 4+ core laptop:
# CPU, sequential 1.10s
# CPU, 4 threads (GIL!) 1.12s <- no speed-up
# CPU, 4 processes 0.35s <- real parallelism
# I/O, sequential (8 x 0.2s) 1.60s
# I/O, 8 threads 0.20s <- threads are fine for I/O
WorkloadUseJava analogue
Numeric / array mathsVectorise (NumPy, pandas, PyTorch) — parallel inside C—
CPU-bound pure PythonProcessPoolExecutor, multiprocessingMultiple JVMs / fork-join (without the GIL problem)
Blocking I/O (legacy SDKs, files)ThreadPoolExecutorFixed thread pool
Many concurrent network callsasyncioVirtual threads / CompletableFuture / reactive

Gotcha — processes don’t share memory. Arguments and results are pickled and copied between processes. Sending a 2 GB DataFrame to each worker costs more than the work itself. Pass file paths or small chunks, not big objects.

Tip — the GIL is going away (slowly). Python 3.13 shipped an experimental free-threaded build (python3.13t) without the GIL, and Python 3.14 made it officially supported (PEP 779), though still optional and not the default. Library support is growing; for now, design with the GIL in mind and treat free-threading as an upgrade path.


Deep Dive: asyncio — Fanning Out 500 LLM Calls

Arjun’s LLM problem is pure waiting: each call spends ~2 seconds blocked on the network. For that, Python’s best tool is asyncio: a single-threaded event loop running thousands of coroutines, each of which yields control whenever it awaits I/O. If you’ve used Netty, Vert.x, WebFlux or CompletableFuture chains, the model is familiar — but the syntax reads like ordinary sequential code, much like Java’s virtual threads.

  • async def defines a coroutine function. Calling it creates a coroutine object; nothing runs yet.
  • await suspends the current coroutine until the awaited operation finishes, letting others run.
  • asyncio.run(main()) starts the event loop.
  • asyncio.gather(...) / TaskGroup run coroutines concurrently.
  • asyncio.Semaphore(n) caps concurrency — essential for API rate limits.

The example below simulates the LLM API with asyncio.sleep so it runs anywhere; swap in a real async client (httpx.AsyncClient, or a provider SDK’s async client) and the structure is identical.

import asyncio
import random
import time
random.seed(1)
class RateLimitError(Exception):
pass
async def call_llm(txn_id: str) -> str:
"""Stand-in for an async HTTP call to an LLM API (~0.2s, occasionally rate-limited)."""
await asyncio.sleep(0.2) # awaiting I/O: the event loop runs other tasks
if random.random() < 0.05:
raise RateLimitError("429 Too Many Requests")
return f"{txn_id}: unusual amount for this customer at an unusual hour."
async def explain(txn_id: str, limiter: asyncio.Semaphore, retries: int = 3) -> str:
for attempt in range(1, retries + 1):
async with limiter: # at most N calls in flight (rate limit)
try:
return await asyncio.wait_for(call_llm(txn_id), timeout=5) # per-call timeout
except RateLimitError:
pass
await asyncio.sleep(0.05 * 2 ** attempt) # exponential backoff OUTSIDE the semaphore
return f"{txn_id}: explanation unavailable"
async def main() -> None:
txn_ids = [f"T{i:03d}" for i in range(500)]
limiter = asyncio.Semaphore(50) # 50 concurrent requests
start = time.perf_counter()
results = await asyncio.gather(*(explain(t, limiter) for t in txn_ids))
elapsed = time.perf_counter() - start
print(len(results), results[0])
# -> 500 T000: unusual amount for this customer at an unusual hour.
print(f"{elapsed:.1f}s instead of ~{0.2 * len(txn_ids):.0f}s sequentially")
# e.g. 2.3s instead of ~100s sequentially
asyncio.run(main())

Gotcha — never block the event loop. Calling a synchronous function inside a coroutine — time.sleep, requests.get, a heavy pandas operation — freezes every other coroutine until it returns. Use async libraries (httpx, aiofiles, async SDK clients), or push blocking calls to a thread with await asyncio.to_thread(blocking_fn, arg).

Gotcha — “coroutine was never awaited”. Calling explain(t, limiter) without await (or without passing it to gather) just creates a coroutine object and throws it away. Python warns, but your code silently did nothing.

Tip — Jupyter already runs an event loop. In a notebook, write await main() directly in a cell instead of asyncio.run(main()), which would fail with “cannot be called from a running event loop”.


The analysts’ text request rests on embeddings: a model maps each text to a vector such that similar meanings land close together. Search then becomes Part 7’s cosine similarity. In production you’d use a neural embedding model; here, to keep the example self-contained and offline, we build a classic stand-in (TF-IDF + truncated SVD, a technique called latent semantic analysis) with scikit-learn. The search code is identical either way.

import numpy as np
from sklearn.decomposition import TruncatedSVD
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import Normalizer
descriptions = [
"monthly grocery shopping at neighbourhood supermarket",
"fuel purchase at highway petrol station",
"gift cards bought in bulk, resold online",
"electronics store laptop and phone purchase",
"bulk prepaid gift card purchase late at night",
"supermarket weekly groceries and household items",
"online gaming credits and gift card top-up",
"petrol and car wash at fuel station",
]
# 'embed' maps text -> unit-length dense vectors (here 4-dimensional).
embed = make_pipeline(
TfidfVectorizer(ngram_range=(1, 2)),
TruncatedSVD(n_components=4, random_state=0),
Normalizer(), # unit length: dot product == cosine similarity
)
vectors = embed.fit_transform(descriptions) # (8, 4)
def search(query: str, k: int = 3) -> list[str]:
q = embed.transform([query])[0] # (4,)
scores = vectors @ q # cosine similarity against every description
return [descriptions[i] for i in np.argsort(-scores)[:k]]
results = search("buying lots of gift cards")
print(vectors.shape) # -> (8, 4)
print(all("gift card" in r for r in results)) # -> True
for r in results:
print("-", r)

With a real neural embedding model the only change is the embed function. For example, with the sentence-transformers package (it downloads a model on first use):

try:
from sentence_transformers import SentenceTransformer # pip/uv add sentence-transformers
model = SentenceTransformer("all-MiniLM-L6-v2") # small, fast, widely used
vecs = model.encode(["gift card fraud", "grocery run"], normalize_embeddings=True)
print(vecs.shape) # (2, 384)
except ImportError:
print("sentence-transformers not installed - see the TF-IDF example above")

From there, a vector database (pgvector, Redis, FAISS, Qdrant…) stores millions of vectors behind an approximate nearest-neighbour index, and retrieval-augmented generation (RAG) feeds the top results into an LLM prompt so its answers are grounded in your data.


Structured LLM Output with Pydantic

LLMs return text. Production systems need typed data. The reliable pattern ties together Part 6 and this part: describe the output with a Pydantic model, send its JSON Schema to the LLM (most provider APIs accept one as a “response format” or a tool definition), then validate the reply — and retry or fall back when validation fails.

import json
from typing import Literal
from pydantic import BaseModel, Field, ValidationError
class FraudExplanation(BaseModel):
txn_id: str
verdict: Literal["likely_fraud", "needs_review", "likely_legit"]
confidence: float = Field(ge=0, le=1)
reasons: list[str] = Field(min_length=1, max_length=5)
schema = FraudExplanation.model_json_schema() # hand this to the LLM API
print(sorted(schema["properties"])) # -> ['confidence', 'reasons', 'txn_id', 'verdict']
def fake_llm(prompt: str) -> str:
"""Stand-in for a real API call; returns what a model might send back."""
return json.dumps({
"txn_id": "T042",
"verdict": "needs_review",
"confidence": 0.72,
"reasons": ["Amount is 6x the customer's median", "First transaction at this merchant"],
})
reply = fake_llm(f"Explain transaction T042. Respond as JSON matching: {json.dumps(schema)}")
explanation = FraudExplanation.model_validate_json(reply) # typed, validated object
print(explanation.verdict, explanation.confidence) # -> needs_review 0.72
# A malformed reply fails loudly instead of corrupting downstream systems:
try:
FraudExplanation.model_validate_json('{"txn_id": "T1", "verdict": "maybe", "confidence": 1.4, "reasons": []}')
except ValidationError as e:
print(e.error_count(), "validation errors") # -> 3 validation errors

Tip — Frameworks like LangChain, LlamaIndex, Pydantic AI and the provider SDKs wrap this loop (schema → call → validate → retry). Understanding the raw pattern first makes them far less magical — and makes debugging them possible.


Python Models Meet JVM Services

Ledgerline’s payment authorisation path is Java, with a strict latency budget. Sentinel’s models are trained in Python. There are four mainstream ways to connect them:

ApproachHowLatencyBest for
Export to ONNX, run in Javatorch.onnx.export / skl2onnx → ONNX Runtime Java APILowest (in-process)Hot paths with strict latency budgets
Python model serviceA small HTTP/gRPC service (FastAPI, BentoML, Triton, TorchServe)+1 network hopComplex models, GPUs, frequent retraining
Streaming scoringPython consumers on Kafka score events and publish resultsAsynchronousNear-real-time enrichment, alerts
Batch scoringSpark / scheduled jobs write scores to a tableHoursRisk reports, offline features
import torch
from torch import nn
model = nn.Sequential(nn.Linear(6, 16), nn.ReLU(), nn.Linear(16, 1)).eval()
example = torch.randn(1, 6) # an example input defines the graph's shapes
try:
torch.onnx.export(
model, (example,), "sentinel_mlp.onnx",
input_names=["features"], output_names=["logit"],
dynamic_axes={"features": {0: "batch"}}, # allow any batch size at inference time
)
print("exported sentinel_mlp.onnx")
except Exception as e: # the exporter needs the 'onnx' package installed
print(f"ONNX export unavailable here ({type(e).__name__}); run: uv add onnx onnxscript")

On the Java side, scoring is a few lines with the com.microsoft.onnxruntime:onnxruntime dependency:

try (var env = OrtEnvironment.getEnvironment();
var session = env.createSession("sentinel_mlp.onnx", new OrtSession.SessionOptions());
var input = OnnxTensor.createTensor(env, new float[][]{{0.3f, 1.2f, -0.5f, 0.0f, 2.1f, 0.7f}});
var result = session.run(Map.of("features", input))) {
float logit = ((float[][]) result.get(0).getValue())[0][0];
double fraudProbability = 1.0 / (1.0 + Math.exp(-logit));
}

Tip — keep feature logic in one place. The classic production bug is “training–serving skew”: Python computes a feature one way during training, Java recomputes it slightly differently at scoring time. Either export the preprocessing into the ONNX graph, compute features in one shared service or feature store, or test both implementations against the same golden dataset.


The Ecosystem Map

NeedGo-to tools
Interactive explorationJupyter / JupyterLab, VS Code notebooks
DataFrames at scalepandas, Polars, DuckDB
Classical MLscikit-learn, XGBoost, LightGBM
Deep learningPyTorch (JAX in research)
Pre-trained modelsHugging Face transformers, sentence-transformers
Local LLM inferenceOllama, llama.cpp, vLLM
LLM application frameworksLangChain / LangGraph, LlamaIndex, Pydantic AI, provider SDKs
Experiment tracking & model registryMLflow, Weights & Biases
Model servingFastAPI, BentoML, Triton, ONNX Runtime

Tips, Tricks & Gotchas

Tip — reproducibility. Set seeds for Python, NumPy and PyTorch (random.seed, np.random.default_rng(seed), torch.manual_seed) and record library versions with uv.lock. GPU kernels can still be non-deterministic; torch.use_deterministic_algorithms(True) trades speed for repeatability.

Gotcha — .item() and .numpy() in hot loops force a device-to-host sync on GPUs. Accumulate on the device and convert once at the end of an epoch.

Tip — start with a pre-trained model. For text and images, fine-tuning a pre-trained model (or simply using its embeddings as features for Part 9’s scikit-learn models) beats training from scratch almost every time.

Gotcha — notebook state. Notebooks let you run cells out of order, so a variable can hold a value no current code produces. Restart the kernel and “Run All” before trusting a result — the notebook equivalent of a clean build.


Key Takeaways

ConceptRemember
TensorsNumPy-like, float32 by default, live on a device; from_numpy shares memory
Autogradrequires_grad, backward(), .grad; gradients accumulate — zero them
Training loopzero_grad → forward → loss → backward → step; train()/eval(); no_grad()
GILThreads don’t speed up CPU-bound Python; processes or vectorisation do
ThreadsFine for blocking I/O
asyncioBest for many concurrent network calls; Semaphore for rate limits; never block the loop
EmbeddingsText → vectors; search = cosine similarity; RAG feeds results to an LLM
Structured outputPydantic schema in, validated object out
JVM integrationONNX in-process, a model service, streaming or batch — avoid training–serving skew

Story Closing

Sentinel Assist went live with an asyncio fan-out capped at fifty concurrent requests: five hundred explanations in under thirty seconds, every one validated against a Pydantic schema before a reviewer saw it. The text-aware fraud model shipped as an ONNX file scored inside the Java authorisation service, adding four milliseconds to the payment path. Its preprocessing was exported in the same graph, so training–serving skew was impossible by construction.

Six weeks earlier, Arjun had assumed working = baseline made a copy. Now he was reviewing Priya’s team’s pull requests — pointing out a leaky transform, a missing model.eval(), a blocking call inside a coroutine. He still wrote Java most days. But when the next problem involved data, he reached for Python first, and it no longer felt like a foreign language.

It felt like another dialect of the same craft — one with fewer semicolons, much better collections, and a remarkable ecosystem of people who had already solved the hard parts.


Where to Go Next

  • Practise the stack on real data: pick a public dataset (Kaggle, UCI, your own logs) and take it from raw CSV to a validated model using only what’s in Parts 7–9.
  • Go deeper on ML: Aurélien Géron’s Hands-On Machine Learning and the official scikit-learn user guide.
  • Go deeper on deep learning: the PyTorch tutorials and fast.ai’s Practical Deep Learning for Coders.
  • Build an AI application: combine Part 10’s embeddings, asyncio and Pydantic into a small RAG service over your own documents — the series AI: Through an Architect’s Lens covers the architecture side.

This is Part 10 of a 10-part series: “Python for Java Developers: From Streams to Tensors.”