"Into the Deep End" — PyTorch, Concurrency and the AI Toolkit
Tensors, autograd and a hand-written PyTorch training loop. Then the questions a Java architect asks: the GIL, threads vs processes, asyncio for fanning out LLM calls, embeddings and semantic search, validated structured output, and how Python models reach JVM services.
Story Opening
Two requests landed on Arjun’s desk in the same week.
The fraud analysts wanted a neural network that combined transaction features with the free-text descriptions merchants attach to payments. And the product team wanted “Sentinel Assist”: for every flagged transaction, call an LLM to draft a plain-language explanation for the reviewer — about five hundred calls per batch.
Arjun’s first attempt at the second request was a plain loop. Each LLM call took about two seconds. Five hundred calls: seventeen minutes. His Java instincts said thread pool. His Python research said something about a “GIL” that made threads useless. Both were half right.
“Different problems, different tools,” said Priya. “For the network, you need PyTorch. For the LLM calls, you need asyncio. And you need to know which is which.”
This final part covers the deep-learning core you need to read and write PyTorch, then the concurrency model every Java developer must relearn in Python, and finally the patterns that glue AI services into production systems.
Java → Python: The Quick Map
| Java world | Python world |
|---|---|
ND4J / DJL NDArray | torch.Tensor (NumPy-like, GPU-capable, differentiable) |
| Hand-derived gradients | Autograd: loss.backward() |
A model class with forward() | torch.nn.Module |
Iterator<Batch> | torch.utils.data.DataLoader |
ExecutorService (threads) | concurrent.futures.ThreadPoolExecutor — for I/O only (the GIL) |
| Separate JVM processes | ProcessPoolExecutor / multiprocessing — for CPU-bound work |
CompletableFuture, virtual threads | asyncio coroutines: async def / await |
Semaphore | asyncio.Semaphore |
| ONNX Runtime Java, DJL | The bridge from Python-trained models to JVM services |
Tensors: NumPy That Can Learn
A PyTorch tensor is an n-dimensional array like NumPy’s, with two extra powers: it can live on a GPU, and it can record the operations applied to it so gradients can be computed automatically. Almost everything you learned in Part 7 — dtypes, shapes, broadcasting, indexing, axes (called dim here) — carries over.
import numpy as npimport torch
x = torch.tensor([[1.0, 2.0], [3.0, 4.0]])print(x.shape, x.dtype) # -> torch.Size([2, 2]) torch.float32 (float32 by default, not float64)
print((x * 2 + 1).tolist()) # -> [[3.0, 5.0], [7.0, 9.0]] (vectorised, like NumPy)print(x.sum(dim=0).tolist()) # -> [4.0, 6.0] ('dim' plays the role of NumPy's 'axis')print((x @ x.T).tolist()) # -> [[5.0, 11.0], [11.0, 25.0]] (matrix multiply)print(x[:, 1].tolist()) # -> [2.0, 4.0]
z = torch.zeros(3, 4)print(z.reshape(2, -1).shape) # -> torch.Size([2, 6])print(torch.rand(2, 3).unsqueeze(0).shape) # -> torch.Size([1, 2, 3]) (add a batch dimension)
# NumPy interop: from_numpy and .numpy() SHARE memory (on CPU) — Part 7's views, again.arr = np.array([1.0, 2.0, 3.0])t = torch.from_numpy(arr)t[0] = 99.0print(arr[0]) # -> 99.0 (same buffer!)safe = torch.tensor(arr) # torch.tensor() always copiesprint(safe.dtype) # -> torch.float64 (dtype preserved from NumPy)
# Devices: move data and models with .to(device). Code stays the same on CPU or GPU.device = "cuda" if torch.cuda.is_available() else ("mps" if torch.backends.mps.is_available() else "cpu")print(x.to(device).device.type in {"cpu", "cuda", "mps"}) # -> TrueGotcha — dtype mismatches. NumPy defaults to
float64; PyTorch models default tofloat32. Feeding a float64 tensor into a float32 model raisesRuntimeError: mat1 and mat2 must have the same dtype. Convert at the boundary:torch.tensor(arr, dtype=torch.float32).
Gotcha — device mismatches. All tensors in one operation must be on the same device. “Expected all tensors to be on the same device” means you moved the model to the GPU but not the batch (or vice versa).
Deep Dive: Autograd — Gradients for Free
Training a neural network means adjusting its parameters to reduce a loss. To know which direction to adjust, you need the gradient of the loss with respect to every parameter. Deriving those by hand for millions of parameters is impossible; autograd does it automatically.
When a tensor has requires_grad=True, PyTorch records every operation applied to it in a computation graph. Calling .backward() on a scalar result walks that graph in reverse (the chain rule — “backpropagation”) and stores each gradient in the leaf tensor’s .grad.
import torch
w = torch.tensor(3.0, requires_grad=True) # a "parameter"x = torch.tensor(2.0) # an input (no gradient needed)
y = w * x ** 2 + 5 * w # y = w·x² + 5w (graph recorded)y.backward() # dy/dw = x² + 5 = 9
print(y.item()) # -> 27.0print(w.grad.item()) # -> 9.0Now the whole learning loop in miniature. We’ll learn the slope and intercept of a line from noisy data using nothing but autograd and gradient descent:
import torch
torch.manual_seed(0)x = torch.linspace(0, 10, 200)y = 2.5 * x + 4.0 + torch.randn(200) # true slope 2.5, intercept 4.0, plus noise
w = torch.zeros(1, requires_grad=True) # start from nothingb = torch.zeros(1, requires_grad=True)lr = 0.01 # learning rate: step size
for step in range(2_000): pred = w * x + b # 1. forward pass loss = ((pred - y) ** 2).mean() # 2. loss: mean squared error loss.backward() # 3. backward pass: fills w.grad and b.grad
with torch.no_grad(): # 4. update WITHOUT recording these ops w -= lr * w.grad b -= lr * b.grad w.grad.zero_() # 5. reset: gradients ACCUMULATE by default b.grad.zero_()
print(round(w.item(), 1), round(b.item(), 1)) # -> 2.5 4.1 (close to the true 2.5 and 4.0; noise shifts the fit slightly)Every neural network you will ever train runs these five steps: forward → loss → backward → update → zero gradients. The libraries add layers, optimisers and data loading around them, but the skeleton never changes.
Gotcha — forgetting
zero_grad(). Gradients accumulate acrossbackward()calls (useful for some advanced tricks). Forget to reset them and each step uses the sum of all previous gradients — training diverges or behaves bizarrely.
A Real Training Loop: nn.Module, DataLoader, Optimiser
Here is the standard shape of PyTorch training code — the shape you’ll recognise in every repository and paper implementation. We train a small multi-layer perceptron (MLP) on synthetic fraud data.
import numpy as npimport torchfrom torch import nnfrom torch.utils.data import DataLoader, TensorDatasetfrom sklearn.metrics import average_precision_score
torch.manual_seed(0)rng = np.random.default_rng(0)
# --- 1. Data: synthetic features with a non-linear fraud signal ----------------n, n_features = 20_000, 6X = rng.normal(size=(n, n_features)).astype(np.float32)logit = -4.0 + 1.5 * X[:, 0] + 2.0 * (X[:, 1] * X[:, 2] > 1.0) + np.abs(X[:, 3])y = (rng.random(n) < 1 / (1 + np.exp(-logit))).astype(np.float32)
split = int(0.8 * n)train_ds = TensorDataset(torch.from_numpy(X[:split]), torch.from_numpy(y[:split]))test_X, test_y = torch.from_numpy(X[split:]), y[split:]
# DataLoader: an iterator of shuffled mini-batches (Part 5's lazy iteration, for tensors).train_loader = DataLoader(train_ds, batch_size=256, shuffle=True)
# --- 2. Model: subclass nn.Module, declare layers in __init__, wire them in forward ---class FraudMLP(nn.Module): def __init__(self, n_in: int, hidden: int = 32): super().__init__() self.net = nn.Sequential( nn.Linear(n_in, hidden), # weights and biases are registered automatically nn.ReLU(), nn.Dropout(0.1), # behaves differently in train vs eval mode nn.Linear(hidden, hidden), nn.ReLU(), nn.Linear(hidden, 1), # one output: a raw score ("logit") )
def forward(self, x: torch.Tensor) -> torch.Tensor: return self.net(x).squeeze(-1) # (batch, 1) -> (batch,)
model = FraudMLP(n_features)print(sum(p.numel() for p in model.parameters())) # -> 1313 (trainable parameters)
# --- 3. Loss and optimiser ------------------------------------------------------pos_weight = torch.tensor((1 - y.mean()) / y.mean()) # up-weight the rare class (Part 9)loss_fn = nn.BCEWithLogitsLoss(pos_weight=pos_weight) # sigmoid + binary cross-entropy, numerically stableoptimizer = torch.optim.AdamW(model.parameters(), lr=3e-3)
# --- 4. Train ---------------------------------------------------------------------for epoch in range(8): model.train() # enable dropout etc. running = 0.0 for xb, yb in train_loader: optimizer.zero_grad() # step 5 from the mini loop, done first by convention loss = loss_fn(model(xb), yb) # model(xb) calls __call__ -> forward (Part 3's callables) loss.backward() optimizer.step() # the optimiser applies the update rule running += loss.item() * len(xb) if epoch % 2 == 1: print(f"epoch {epoch + 1}: loss {running / len(train_ds):.4f}")
# --- 5. Evaluate ------------------------------------------------------------------model.eval() # disable dropoutwith torch.no_grad(): # no graph recording: faster, less memory scores = torch.sigmoid(model(test_X)).numpy() # logits -> probabilities
print(f"test PR AUC: {average_precision_score(test_y, scores):.3f} (base rate {test_y.mean():.3f})")
torch.save(model.state_dict(), "fraud_mlp.pt") # save the WEIGHTS, not the pickled objectTip —
model.train()andmodel.eval()don’t train or evaluate anything. They flip a mode flag that changes how layers likeDropoutandBatchNormbehave. Forgettingeval()before inference gives noisy, non-deterministic predictions.
Tip —
torch.no_grad()is a context manager (Part 5) that switches graph recording off and restores it afterwards. Use it (ortorch.inference_mode()) for every evaluation and inference path.
| Step | Code | Why |
|---|---|---|
| Reset gradients | optimizer.zero_grad() | Gradients accumulate |
| Forward | out = model(xb) | Compute predictions |
| Loss | loss = loss_fn(out, yb) | One scalar to minimise |
| Backward | loss.backward() | Autograd computes all gradients |
| Update | optimizer.step() | Apply the optimiser’s rule (AdamW, SGD…) |
Deep Dive: The GIL, Threads and Processes
Here’s where Java developers get burned. CPython has a Global Interpreter Lock (GIL): only one thread executes Python bytecode at a time. It exists because CPython’s memory management (reference counting) isn’t thread-safe without it.
The consequences:
- CPU-bound pure-Python code doesn’t get faster with threads. Four threads computing in Python take as long as one (or longer).
- I/O-bound code does benefit from threads. A thread waiting on a socket, file or
time.sleepreleases the GIL, so other threads run. - C extensions release the GIL. NumPy, pandas internals and PyTorch kernels run their heavy loops outside the GIL — which is why vectorised code (Parts 7–8) can use multiple cores internally.
import timefrom concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor
def cpu_work(n: int) -> int: """Pure-Python CPU-bound work: holds the GIL the whole time.""" total = 0 for i in range(n): total += i * i % 7 return total
def io_work(_: int) -> str: """I/O-bound work: sleeping (like waiting on a network call) releases the GIL.""" time.sleep(0.2) return "ok"
def timed(label, fn): start = time.perf_counter() fn() print(f"{label:28s} {time.perf_counter() - start:.2f}s")
if __name__ == "__main__": # REQUIRED for ProcessPoolExecutor on macOS/Windows jobs = [2_000_000] * 4
timed("CPU, sequential", lambda: [cpu_work(n) for n in jobs]) with ThreadPoolExecutor(max_workers=4) as pool: # ≈ Executors.newFixedThreadPool(4) timed("CPU, 4 threads (GIL!)", lambda: list(pool.map(cpu_work, jobs))) with ProcessPoolExecutor(max_workers=4) as pool: # 4 interpreters, 4 GILs timed("CPU, 4 processes", lambda: list(pool.map(cpu_work, jobs)))
timed("I/O, sequential (8 x 0.2s)", lambda: [io_work(i) for i in range(8)]) with ThreadPoolExecutor(max_workers=8) as pool: timed("I/O, 8 threads", lambda: list(pool.map(io_work, range(8))))
# Typical output on a 4+ core laptop:# CPU, sequential 1.10s# CPU, 4 threads (GIL!) 1.12s <- no speed-up# CPU, 4 processes 0.35s <- real parallelism# I/O, sequential (8 x 0.2s) 1.60s# I/O, 8 threads 0.20s <- threads are fine for I/O| Workload | Use | Java analogue |
|---|---|---|
| Numeric / array maths | Vectorise (NumPy, pandas, PyTorch) — parallel inside C | — |
| CPU-bound pure Python | ProcessPoolExecutor, multiprocessing | Multiple JVMs / fork-join (without the GIL problem) |
| Blocking I/O (legacy SDKs, files) | ThreadPoolExecutor | Fixed thread pool |
| Many concurrent network calls | asyncio | Virtual threads / CompletableFuture / reactive |
Gotcha — processes don’t share memory. Arguments and results are pickled and copied between processes. Sending a 2 GB DataFrame to each worker costs more than the work itself. Pass file paths or small chunks, not big objects.
Tip — the GIL is going away (slowly). Python 3.13 shipped an experimental free-threaded build (
python3.13t) without the GIL, and Python 3.14 made it officially supported (PEP 779), though still optional and not the default. Library support is growing; for now, design with the GIL in mind and treat free-threading as an upgrade path.
Deep Dive: asyncio — Fanning Out 500 LLM Calls
Arjun’s LLM problem is pure waiting: each call spends ~2 seconds blocked on the network. For that, Python’s best tool is asyncio: a single-threaded event loop running thousands of coroutines, each of which yields control whenever it awaits I/O. If you’ve used Netty, Vert.x, WebFlux or CompletableFuture chains, the model is familiar — but the syntax reads like ordinary sequential code, much like Java’s virtual threads.
async defdefines a coroutine function. Calling it creates a coroutine object; nothing runs yet.awaitsuspends the current coroutine until the awaited operation finishes, letting others run.asyncio.run(main())starts the event loop.asyncio.gather(...)/TaskGrouprun coroutines concurrently.asyncio.Semaphore(n)caps concurrency — essential for API rate limits.
The example below simulates the LLM API with asyncio.sleep so it runs anywhere; swap in a real async client (httpx.AsyncClient, or a provider SDK’s async client) and the structure is identical.
import asyncioimport randomimport time
random.seed(1)
class RateLimitError(Exception): pass
async def call_llm(txn_id: str) -> str: """Stand-in for an async HTTP call to an LLM API (~0.2s, occasionally rate-limited).""" await asyncio.sleep(0.2) # awaiting I/O: the event loop runs other tasks if random.random() < 0.05: raise RateLimitError("429 Too Many Requests") return f"{txn_id}: unusual amount for this customer at an unusual hour."
async def explain(txn_id: str, limiter: asyncio.Semaphore, retries: int = 3) -> str: for attempt in range(1, retries + 1): async with limiter: # at most N calls in flight (rate limit) try: return await asyncio.wait_for(call_llm(txn_id), timeout=5) # per-call timeout except RateLimitError: pass await asyncio.sleep(0.05 * 2 ** attempt) # exponential backoff OUTSIDE the semaphore return f"{txn_id}: explanation unavailable"
async def main() -> None: txn_ids = [f"T{i:03d}" for i in range(500)] limiter = asyncio.Semaphore(50) # 50 concurrent requests
start = time.perf_counter() results = await asyncio.gather(*(explain(t, limiter) for t in txn_ids)) elapsed = time.perf_counter() - start
print(len(results), results[0]) # -> 500 T000: unusual amount for this customer at an unusual hour. print(f"{elapsed:.1f}s instead of ~{0.2 * len(txn_ids):.0f}s sequentially") # e.g. 2.3s instead of ~100s sequentially
asyncio.run(main())Gotcha — never block the event loop. Calling a synchronous function inside a coroutine —
time.sleep,requests.get, a heavy pandas operation — freezes every other coroutine until it returns. Use async libraries (httpx,aiofiles, async SDK clients), or push blocking calls to a thread withawait asyncio.to_thread(blocking_fn, arg).
Gotcha — “coroutine was never awaited”. Calling
explain(t, limiter)withoutawait(or without passing it togather) just creates a coroutine object and throws it away. Python warns, but your code silently did nothing.
Tip — Jupyter already runs an event loop. In a notebook, write
await main()directly in a cell instead ofasyncio.run(main()), which would fail with “cannot be called from a running event loop”.
Embeddings and Semantic Search
The analysts’ text request rests on embeddings: a model maps each text to a vector such that similar meanings land close together. Search then becomes Part 7’s cosine similarity. In production you’d use a neural embedding model; here, to keep the example self-contained and offline, we build a classic stand-in (TF-IDF + truncated SVD, a technique called latent semantic analysis) with scikit-learn. The search code is identical either way.
import numpy as npfrom sklearn.decomposition import TruncatedSVDfrom sklearn.feature_extraction.text import TfidfVectorizerfrom sklearn.pipeline import make_pipelinefrom sklearn.preprocessing import Normalizer
descriptions = [ "monthly grocery shopping at neighbourhood supermarket", "fuel purchase at highway petrol station", "gift cards bought in bulk, resold online", "electronics store laptop and phone purchase", "bulk prepaid gift card purchase late at night", "supermarket weekly groceries and household items", "online gaming credits and gift card top-up", "petrol and car wash at fuel station",]
# 'embed' maps text -> unit-length dense vectors (here 4-dimensional).embed = make_pipeline( TfidfVectorizer(ngram_range=(1, 2)), TruncatedSVD(n_components=4, random_state=0), Normalizer(), # unit length: dot product == cosine similarity)vectors = embed.fit_transform(descriptions) # (8, 4)
def search(query: str, k: int = 3) -> list[str]: q = embed.transform([query])[0] # (4,) scores = vectors @ q # cosine similarity against every description return [descriptions[i] for i in np.argsort(-scores)[:k]]
results = search("buying lots of gift cards")print(vectors.shape) # -> (8, 4)print(all("gift card" in r for r in results)) # -> Truefor r in results: print("-", r)With a real neural embedding model the only change is the embed function. For example, with the sentence-transformers package (it downloads a model on first use):
try: from sentence_transformers import SentenceTransformer # pip/uv add sentence-transformers
model = SentenceTransformer("all-MiniLM-L6-v2") # small, fast, widely used vecs = model.encode(["gift card fraud", "grocery run"], normalize_embeddings=True) print(vecs.shape) # (2, 384)except ImportError: print("sentence-transformers not installed - see the TF-IDF example above")From there, a vector database (pgvector, Redis, FAISS, Qdrant…) stores millions of vectors behind an approximate nearest-neighbour index, and retrieval-augmented generation (RAG) feeds the top results into an LLM prompt so its answers are grounded in your data.
Structured LLM Output with Pydantic
LLMs return text. Production systems need typed data. The reliable pattern ties together Part 6 and this part: describe the output with a Pydantic model, send its JSON Schema to the LLM (most provider APIs accept one as a “response format” or a tool definition), then validate the reply — and retry or fall back when validation fails.
import jsonfrom typing import Literal
from pydantic import BaseModel, Field, ValidationError
class FraudExplanation(BaseModel): txn_id: str verdict: Literal["likely_fraud", "needs_review", "likely_legit"] confidence: float = Field(ge=0, le=1) reasons: list[str] = Field(min_length=1, max_length=5)
schema = FraudExplanation.model_json_schema() # hand this to the LLM APIprint(sorted(schema["properties"])) # -> ['confidence', 'reasons', 'txn_id', 'verdict']
def fake_llm(prompt: str) -> str: """Stand-in for a real API call; returns what a model might send back.""" return json.dumps({ "txn_id": "T042", "verdict": "needs_review", "confidence": 0.72, "reasons": ["Amount is 6x the customer's median", "First transaction at this merchant"], })
reply = fake_llm(f"Explain transaction T042. Respond as JSON matching: {json.dumps(schema)}")explanation = FraudExplanation.model_validate_json(reply) # typed, validated objectprint(explanation.verdict, explanation.confidence) # -> needs_review 0.72
# A malformed reply fails loudly instead of corrupting downstream systems:try: FraudExplanation.model_validate_json('{"txn_id": "T1", "verdict": "maybe", "confidence": 1.4, "reasons": []}')except ValidationError as e: print(e.error_count(), "validation errors") # -> 3 validation errorsTip — Frameworks like LangChain, LlamaIndex, Pydantic AI and the provider SDKs wrap this loop (schema → call → validate → retry). Understanding the raw pattern first makes them far less magical — and makes debugging them possible.
Python Models Meet JVM Services
Ledgerline’s payment authorisation path is Java, with a strict latency budget. Sentinel’s models are trained in Python. There are four mainstream ways to connect them:
| Approach | How | Latency | Best for |
|---|---|---|---|
| Export to ONNX, run in Java | torch.onnx.export / skl2onnx → ONNX Runtime Java API | Lowest (in-process) | Hot paths with strict latency budgets |
| Python model service | A small HTTP/gRPC service (FastAPI, BentoML, Triton, TorchServe) | +1 network hop | Complex models, GPUs, frequent retraining |
| Streaming scoring | Python consumers on Kafka score events and publish results | Asynchronous | Near-real-time enrichment, alerts |
| Batch scoring | Spark / scheduled jobs write scores to a table | Hours | Risk reports, offline features |
import torchfrom torch import nn
model = nn.Sequential(nn.Linear(6, 16), nn.ReLU(), nn.Linear(16, 1)).eval()example = torch.randn(1, 6) # an example input defines the graph's shapes
try: torch.onnx.export( model, (example,), "sentinel_mlp.onnx", input_names=["features"], output_names=["logit"], dynamic_axes={"features": {0: "batch"}}, # allow any batch size at inference time ) print("exported sentinel_mlp.onnx")except Exception as e: # the exporter needs the 'onnx' package installed print(f"ONNX export unavailable here ({type(e).__name__}); run: uv add onnx onnxscript")On the Java side, scoring is a few lines with the com.microsoft.onnxruntime:onnxruntime dependency:
try (var env = OrtEnvironment.getEnvironment(); var session = env.createSession("sentinel_mlp.onnx", new OrtSession.SessionOptions()); var input = OnnxTensor.createTensor(env, new float[][]{{0.3f, 1.2f, -0.5f, 0.0f, 2.1f, 0.7f}}); var result = session.run(Map.of("features", input))) { float logit = ((float[][]) result.get(0).getValue())[0][0]; double fraudProbability = 1.0 / (1.0 + Math.exp(-logit));}Tip — keep feature logic in one place. The classic production bug is “training–serving skew”: Python computes a feature one way during training, Java recomputes it slightly differently at scoring time. Either export the preprocessing into the ONNX graph, compute features in one shared service or feature store, or test both implementations against the same golden dataset.
The Ecosystem Map
| Need | Go-to tools |
|---|---|
| Interactive exploration | Jupyter / JupyterLab, VS Code notebooks |
| DataFrames at scale | pandas, Polars, DuckDB |
| Classical ML | scikit-learn, XGBoost, LightGBM |
| Deep learning | PyTorch (JAX in research) |
| Pre-trained models | Hugging Face transformers, sentence-transformers |
| Local LLM inference | Ollama, llama.cpp, vLLM |
| LLM application frameworks | LangChain / LangGraph, LlamaIndex, Pydantic AI, provider SDKs |
| Experiment tracking & model registry | MLflow, Weights & Biases |
| Model serving | FastAPI, BentoML, Triton, ONNX Runtime |
Tips, Tricks & Gotchas
Tip — reproducibility. Set seeds for Python, NumPy and PyTorch (
random.seed,np.random.default_rng(seed),torch.manual_seed) and record library versions withuv.lock. GPU kernels can still be non-deterministic;torch.use_deterministic_algorithms(True)trades speed for repeatability.
Gotcha —
.item()and.numpy()in hot loops force a device-to-host sync on GPUs. Accumulate on the device and convert once at the end of an epoch.
Tip — start with a pre-trained model. For text and images, fine-tuning a pre-trained model (or simply using its embeddings as features for Part 9’s scikit-learn models) beats training from scratch almost every time.
Gotcha — notebook state. Notebooks let you run cells out of order, so a variable can hold a value no current code produces. Restart the kernel and “Run All” before trusting a result — the notebook equivalent of a clean build.
Key Takeaways
| Concept | Remember |
|---|---|
| Tensors | NumPy-like, float32 by default, live on a device; from_numpy shares memory |
| Autograd | requires_grad, backward(), .grad; gradients accumulate — zero them |
| Training loop | zero_grad → forward → loss → backward → step; train()/eval(); no_grad() |
| GIL | Threads don’t speed up CPU-bound Python; processes or vectorisation do |
| Threads | Fine for blocking I/O |
asyncio | Best for many concurrent network calls; Semaphore for rate limits; never block the loop |
| Embeddings | Text → vectors; search = cosine similarity; RAG feeds results to an LLM |
| Structured output | Pydantic schema in, validated object out |
| JVM integration | ONNX in-process, a model service, streaming or batch — avoid training–serving skew |
Story Closing
Sentinel Assist went live with an asyncio fan-out capped at fifty concurrent requests: five hundred explanations in under thirty seconds, every one validated against a Pydantic schema before a reviewer saw it. The text-aware fraud model shipped as an ONNX file scored inside the Java authorisation service, adding four milliseconds to the payment path. Its preprocessing was exported in the same graph, so training–serving skew was impossible by construction.
Six weeks earlier, Arjun had assumed working = baseline made a copy. Now he was reviewing Priya’s team’s pull requests — pointing out a leaky transform, a missing model.eval(), a blocking call inside a coroutine. He still wrote Java most days. But when the next problem involved data, he reached for Python first, and it no longer felt like a foreign language.
It felt like another dialect of the same craft — one with fewer semicolons, much better collections, and a remarkable ecosystem of people who had already solved the hard parts.
Where to Go Next
- Practise the stack on real data: pick a public dataset (Kaggle, UCI, your own logs) and take it from raw CSV to a validated model using only what’s in Parts 7–9.
- Go deeper on ML: Aurélien Géron’s Hands-On Machine Learning and the official scikit-learn user guide.
- Go deeper on deep learning: the PyTorch tutorials and fast.ai’s Practical Deep Learning for Coders.
- Build an AI application: combine Part 10’s embeddings,
asyncioand Pydantic into a small RAG service over your own documents — the series AI: Through an Architect’s Lens covers the architecture side.
This is Part 10 of a 10-part series: “Python for Java Developers: From Streams to Tensors.”