Story Opening

The prototype was a notebook and one long prompt: “You are a helpful assistant for Shelfwise”, followed by every product in the catalog, pasted in as text. Nothing in it mentioned returns, warranties or delivery, so when someone asked about a broken mixer grinder, the model did what language models do with a gap. It filled it, plausibly.

Kabir’s first instinct was the one every developer has: fix the prompt. Add “never make things up”, maybe in capitals. Lena shook her head before he’d finished the sentence. “It can’t decline to answer from a document it’s never seen. Give it the documents, and give it a way to say it has none.”


Java → Kotlin: The Quick Map

This part maps familiar Spring and search concepts onto Spring AI 2.0:

What you knowSpring AI 2.0Note
A DataSource bean configured from propertiesChatModel and EmbeddingModel beans from a starterOllama here; a property switch for hosted providers
RestTemplate calls to a model’s HTTP APIChatClient fluent APIprompt().user(…).call()
Lucene or Elasticsearch full-text searchVectorStore.similaritySearchMeaning, not matching words
A SQL WHERE clauseFilterExpressionBuilderMetadata filters, translated to Qdrant’s
A servlet filter chain around a callAdvisors around a ChatClient callRetrievalAugmentationAdvisor adds retrieved context
Hand-written JSON for “function calling”@Tool methodsSpring generates the schema and runs the call
@KafkaListener with a Jackson payload@KafkaListener + ErrorHandlingDeserializer + an exhaustive whenThe shared sealed events from Part 14

Conceptual Deep-Dive

A model is a function of its prompt

A chat model doesn’t look anything up. It continues the text it’s given, and its answer is only as good as that text. So the work in a product assistant is not in the model call but around it: deciding what goes into the prompt, and what the model may call. That pattern is retrieval-augmented generation (RAG), and it has two halves:

graph LR subgraph Ingest["Ingest (continuous)"] K[("Kafka
product-events")] --> L["@KafkaListener"] P["policies/*.md"] --> I L --> I["ProductIndexer"] I -->|"text → vector"| E1["EmbeddingModel
nomic-embed-text"] E1 --> Q[("Qdrant
products_springai")] end subgraph Ask["Ask (per question)"] U["Question"] --> E2["EmbeddingModel"] E2 -->|"nearest vectors,
score ≥ 0.6"| Q Q --> A["Augment the prompt"] A --> C["ChatModel
qwen3:4b-instruct"] C <-->|"@Tool"| T["Catalog API
live price"] C --> R["Answer + sources"] end

An embedding model turns a text into a vector, here 768 numbers, such that texts with similar meaning point in similar directions. Indexing stores each product’s and policy’s vector in Qdrant with its text and metadata. Asking embeds the question the same way, finds the nearest stored vectors, and puts their text into the prompt as context. “Gluten-free pasta under ₹200” lands near the millet pasta’s description although the words differ, and a metadata filter can still demand a price under 20,000 paise exactly.

The shift for a Java developer: you don’t make an LLM correct by instructing it; you make it correct by controlling its inputs. Retrieval decides what it knows, a similarity threshold decides when it knows nothing, and tools decide what it can look up. The prompt’s instructions only tell it how to use those.

One port, two adapters

The service talks to AI libraries through two Kotlin interfaces, and Part 16 implements them a second time with LangChain4j:

// Fragment of ports/Ports.kt
// What the service needs from an AI library, in its own words. Part 15 implements both ports with
// Spring AI; Part 16 implements them again with LangChain4j, behind a profile.
interface ProductIndexer {
fun upsert(product: ProductSnapshot)
fun delete(sku: String)
fun upsertPolicy(name: String, text: String)
}
data class ProductHit(val sku: String, val name: String, val pricePaise: Long, val score: Double)
data class Answer(val text: String, val sources: List<String>)
interface ProductAssistant {
fun search(query: String, maxPricePaise: Long? = null, dietaryTag: String? = null, limit: Int = 5): List<ProductHit>
fun ask(question: String): Answer
}

The adapters never share a request path, and they don’t share storage either. Each gets its own Qdrant collection (products_springai here), because the two libraries store a document’s text under different payload keys and each would read the other’s points as empty. Each gets its own Kafka consumer group (product-ai-springai): a group with no committed offsets starts at the earliest retained record (with auto.offset.reset=earliest; Kafka’s default is latest), so switching profiles replays the topic and rebuilds that adapter’s index, as far back as the topic’s retention reaches (seven days by default; the catalog’s republish fills in the rest). Both use the same embedding model; vectors from different models aren’t comparable, and mixing them degrades retrieval without any error.


Technical Explanation

The build: Spring AI 2.0 on Boot 4.1

// Fragment of build.gradle.kts
dependencyManagement {
imports {
mavenBom("org.springframework.ai:spring-ai-bom:${libs.findVersion("spring-ai").get().requiredVersion}")
}
dependencies {
// One Qdrant client for both AI libraries, matching the server's minor version.
dependency("io.qdrant:client:${libs.findVersion("qdrant-client").get().requiredVersion}")
}
}
// Fragment of build.gradle.kts
// Spring AI: a chat and an embedding model from Ollama, Qdrant as the vector store, and the RAG building blocks.
implementation("org.springframework.ai:spring-ai-starter-model-ollama")
implementation("org.springframework.ai:spring-ai-starter-vector-store-qdrant")
implementation("org.springframework.ai:spring-ai-rag")

Spring AI 2.0.1 is the current release for Boot 4.0 and 4.1 (2.1 milestones target Boot 4.2); its BOM manages the starters. The pinned Qdrant client solves a dependency puzzle: Spring AI 2.0.1’s Qdrant store depends on io.qdrant:client 1.18.0, LangChain4j 1.21’s on 1.17.0, and the server is 1.19.1. Qdrant releases its Java client in step with the server and recommends matching minor versions, so the build pins 1.19.0: the Spring AI tests run against it here, and Part 16’s adapter shares the pin.

Documents with stable identities

// Fragment of ports/IndexedDocuments.kt
// How products and policies become documents, shared by both adapters so they index the same text.
object IndexedDocuments {
// Deterministic point ids: the same SKU always maps to the same point, so a redelivered or
// replayed event overwrites instead of duplicating. That's the idempotence Kafka needs.
fun productId(sku: String): String = UUID.nameUUIDFromBytes("product:$sku".encodeToByteArray()).toString()
fun policyId(name: String): String = UUID.nameUUIDFromBytes("policy:$name".encodeToByteArray()).toString()
fun productText(p: ProductSnapshot): String = buildString {
appendLine("${p.name} (SKU ${p.sku})")
if (p.description.isNotBlank()) appendLine(p.description)
append("Category: ${p.category.lowercase().replace('_', ' ')}. ")
append("Price: ₹%d.%02d.".format(p.pricePaise / 100, p.pricePaise % 100))
if (p.dietaryTags.isNotEmpty()) append(" Dietary: ${p.dietaryTags.joinToString()}.")
}
fun productMetadata(p: ProductSnapshot): Map<String, Any> = mapOf(
"type" to "product",
"sku" to p.sku,
"name" to p.name,
"category" to p.category,
// An Int, not a Long: Spring AI's Qdrant store writes Long values as strings, and a numeric
// range filter on a string field matches nothing. Paise fit in an Int up to ₹2 crore.
"pricePaise" to Math.toIntExact(p.pricePaise),
"dietaryTags" to p.dietaryTags,
)
}

The point id is a name-based UUID derived from the SKU: the same product always maps to the same point. Kafka delivers at least once, Part 14’s events are keyed by SKU so one product’s create, updates and delete arrive in order, the republish sends every product again, and a new consumer group replays what the topic retains; each of those overwrites a point instead of adding a duplicate. That’s the consumer side of Part 14’s idempotence requirement, achieved without a deduplication table. The text is what gets embedded; the metadata is what filters match. pricePaise is stored as an Int, because Spring AI’s Qdrant store writes a Long as a string. (Math.toIntExact would throw for a price above ₹2 crore, and the error handler would retry, then skip the event; a Double would avoid that edge entirely.)

Consuming product events

// Fragment of ingest/ProductEventListener.kt
// Keeps the vector index in step with the catalog. Exhaustive over the sealed hierarchy: a new event
// type in shared/product-events doesn't compile here until it's handled.
@Component
class ProductEventListener(private val indexer: ProductIndexer) {
private val log = LoggerFactory.getLogger(javaClass)
// A plain (blocking) listener: indexing calls the embedding model and Qdrant synchronously. With the
// default AckMode.BATCH, offsets are committed once every record from a poll has been processed.
@KafkaListener(topics = [ProductEvents.TOPIC], groupId = "\${assistant.consumer-group}")
fun on(event: ProductEvent) {
when (event) {
is ProductCreated -> indexer.upsert(event.product)
is ProductUpdated -> indexer.upsert(event.product)
is ProductDeleted -> indexer.delete(event.sku)
}
log.info("Indexed {} for {}", event.type, event.sku)
}
}

The when is exhaustive over Part 14’s sealed hierarchy, so a new event type fails this module’s compilation once it upgrades shared/product-events. What it can’t do is decode an event type that’s newer than the module it was built with, and Kafka will keep redelivering a record the deserializer can’t read: the classic poison pill. Spring Kafka’s ErrorHandlingDeserializer wraps the real deserializer and turns that failure into a record the default error handler logs and skips:

kafka:
bootstrap-servers: localhost:9092
consumer:
auto-offset-reset: earliest # a new consumer group replays the topic and rebuilds the index
value-deserializer: org.springframework.kafka.support.serializer.ErrorHandlingDeserializer
properties:
spring.deserializer.value.delegate.class: com.shelfwise.events.ProductEventDeserializer

Spring Kafka also accepts suspend listeners, but then it switches the container to manual acknowledgement with out-of-order commits, and in-flight work can outlive a container stop (spring-kafka issue #4519); with synchronous indexing there’s nothing to gain. For throughput, a batch listener (List<ProductEvent>) and one vectorStore.add per batch would be the next step; Ollama embeds a batch in one request.

Indexing with Spring AI

// Fragment of springai/SpringAiProductIndexer.kt
@Component
@Profile("!lc4j")
class SpringAiProductIndexer(private val vectorStore: VectorStore) : ProductIndexer {
// add() embeds the text with the EmbeddingModel and upserts the point: same id, new vector and payload.
override fun upsert(product: ProductSnapshot) = vectorStore.add(
listOf(
Document.builder()
.id(IndexedDocuments.productId(product.sku))
.text(IndexedDocuments.productText(product))
.metadata(IndexedDocuments.productMetadata(product))
.build(),
),
)
override fun delete(sku: String) = vectorStore.delete(listOf(IndexedDocuments.productId(sku)))
override fun upsertPolicy(name: String, text: String) = vectorStore.add(
listOf(
Document.builder()
.id(IndexedDocuments.policyId(name))
.text(text)
.metadata(mapOf("type" to "policy", "policy" to name))
.build(),
),
)
}

VectorStore.add embeds each document with the configured EmbeddingModel and upserts it into Qdrant; delete takes the same ids. With initialize-schema: true, the store creates the collection on startup, sized to the embedding model’s dimensions (768 for nomic-embed-text). Policies go through the same port: PolicyLoader indexes classpath:policies/*.md at startup with deterministic ids, so restarts overwrite them.

Searching and answering

// Fragment of springai/SpringAiProductAssistant.kt
@Component
@Profile("!lc4j")
class SpringAiProductAssistant(
private val vectorStore: VectorStore,
chatClient: ChatClient.Builder,
catalogTools: CatalogTools,
@Value("\${assistant.rag.similarity-threshold}") similarityThreshold: Double,
) : ProductAssistant {
private val withContext = ContextualQueryAugmenter.builder()
.promptTemplate(PromptTemplate(AssistantPrompts.WITH_CONTEXT.trimIndent()))
.build()
private val chat = chatClient
.defaultSystem(AssistantPrompts.SYSTEM.trimIndent())
.defaultAdvisors(
RetrievalAugmentationAdvisor.builder()
.documentRetriever(
VectorStoreDocumentRetriever.builder()
.vectorStore(vectorStore)
.similarityThreshold(similarityThreshold) // weak matches are worse than none
.topK(4)
.build(),
)
// With nothing relevant retrieved, the model must not improvise a return policy, but a
// tool may still answer. ContextualQueryAugmenter's own empty-context prompt drops the
// user's question, so the empty case is handled here: a Kotlin lambda for a Java interface.
.queryAugmenter { query, documents ->
if (documents.isEmpty()) Query(AssistantPrompts.withoutContext(query.text())) else withContext.augment(query, documents)
}
.build(),
)
.defaultTools(catalogTools)
.build()
override fun search(query: String, maxPricePaise: Long?, dietaryTag: String?, limit: Int): List<ProductHit> {
val filter = FilterExpressionBuilder().run {
listOfNotNull(
eq("type", "product"),
maxPricePaise?.let { lte("pricePaise", it) },
dietaryTag?.let { eq("dietaryTags", it) }, // matches if any element of the list equals it
).reduce { all, condition -> and(all, condition) }.build()
}
val request = SearchRequest.builder().query(query).topK(limit).filterExpression(filter).build()
return vectorStore.similaritySearch(request).map { doc ->
ProductHit(
sku = doc.metadata["sku"] as String,
name = doc.metadata["name"] as String,
pricePaise = (doc.metadata["pricePaise"] as Number).toLong(),
score = doc.score ?: 0.0,
)
}
}
override fun ask(question: String): Answer {
val response = chat.prompt().user(question).call().chatClientResponse()
@Suppress("UNCHECKED_CAST")
val sources = response.context()[RetrievalAugmentationAdvisor.DOCUMENT_CONTEXT] as List<Document>? ?: emptyList()
return Answer(
text = response.chatResponse()?.result?.output?.text.orEmpty(),
sources = sources.map { it.metadata["sku"] as String? ?: it.metadata["policy"] as String }.distinct(),
)
}
}

search is semantic search with exact filters. FilterExpressionBuilder builds a portable expression, which the Qdrant store translates into Qdrant’s own filter; run and reduce combine however many conditions the caller supplied. lte("pricePaise", it) passes a Long and still matches the Int stored in the payload, because the converter parses every range value as a Double: filters accept a Long, but writes turn one into a string.

ask assembles the RAG pipeline from Spring AI’s building blocks. VectorStoreDocumentRetriever fetches up to four documents scoring at least the configured threshold. A query augmenter then rewrites the question to include them: with documents, ContextualQueryAugmenter and the WITH_CONTEXT template; without any, an instruction not to improvise, plus the question itself. The augmenter is a Kotlin lambda because QueryAugmenter is a Java interface with one abstract method, which Kotlin converts automatically (a Kotlin interface would need to be declared fun interface, Part 5). defaultTools registers the catalog tools, and the advisor returns the retrieved documents in the response context under DOCUMENT_CONTEXT, which become the answer’s sources. Those make a wrong answer diagnosable: was the right document not retrieved, or retrieved and ignored? The ?. chain in ask is there because Spring AI 2.0 annotates its API with JSpecify, so chatResponse() and friends are honestly nullable in Kotlin instead of platform types (Parts 2 and 12).

// Fragment of ports/AssistantPrompts.kt
// The instructions both adapters give the model, so Part 16 compares libraries, not prompts.
object AssistantPrompts {
const val SYSTEM = """
You are Shelfwise's product assistant for store staff and customers.
Answer from the context you are given and from your tools; never invent policies, prices or products.
When asked what a product costs now, call the currentPrice tool with its SKU.
Prices are in Indian rupees. Mention SKUs when you recommend products. Keep answers to one or two sentences.
"""
// {context} and {query} are filled in by the RAG step.
const val WITH_CONTEXT = """
Context information is below.
---------------------
{context}
---------------------
Answer the query using the context, and your tools where they help.
If neither the context nor a tool gives the answer, say that you don't know.
Don't start with phrases like "Based on the context".
Query: {query}
Answer:
"""
// Used when retrieval finds nothing relevant: no improvising, but a tool may still answer.
fun withoutContext(query: String) = """
No store documents match this query.
If one of your tools can answer it, use the tool. Otherwise, politely say that you can't help with it.
Query: $query
""".trimIndent()
}

The {context} and {query} placeholders are Spring AI’s template syntax. The texts are const vals so that annotations can use them too (Part 16’s LangChain4j service takes its system message from one), which is why .trimIndent() appears at each use site instead of in the declaration.

A tool for live prices

// Fragment of tools/CatalogTools.kt
data class CatalogProduct(val sku: String, val name: String, val price: String)
// The catalog's REST API, as an HTTP interface (Part 14). Blocking is fine here: tools run on the
// thread that serves the request, a virtual thread.
interface CatalogClient {
@GetExchange("/api/products/{sku}")
fun product(@PathVariable sku: String): CatalogProduct
}
// Indexed prices can be minutes old; a tool lets the model ask the catalog for the price right now.
@Component
class CatalogTools(private val catalog: CatalogClient) {
@Tool(description = "Get the current price of a Shelfwise product from the catalog, by SKU. Use it when asked what something costs now.")
fun currentPrice(@ToolParam(description = "The product SKU, for example SHW-1001") sku: String): String =
try {
catalog.product(sku).let { "${it.name} (${it.sku}) costs ${it.price}" }
} catch (e: HttpClientErrorException.NotFound) {
"There is no product with SKU $sku"
}
}

@Tool turns the method into a function the model can call: Spring AI derives a JSON schema from the Kotlin signature and the @ToolParam description, sends it with the request, runs the method when the model asks for it, and sends the result back for the final answer. Kotlin’s nullability can shape that schema, with one catch the Gotchas show. If the catalog is down, the RestClient call throws, and by default Spring AI sends the exception’s message back to the model as the tool’s result rather than failing the request. The model then usually says it couldn’t check the price, but the message (for a connection failure, the catalog’s internal URL) is now in the model’s context and can surface in an answer; catch it in the tool and return something neutral when that matters.

Observability

Spring AI records model calls, vector searches and advisors as Micrometer observations. With the metrics endpoint exposed, one question produced these meters: gen_ai.client.operation and gen_ai.client.token.usage (tagged with gen_ai.operation.name chat or embedding, the model names, and gen_ai.token.type input, output or total), db.vector.client.operation, spring.ai.chat.client and spring.ai.advisor. Token usage per model is the number to watch once a hosted model costs money.


Step-by-Step Hands-On: Ask It Something

Code: kotlin-for-java-survivors/services/product-ai-service, tag kotlin-for-java-survivors/part-15.

Step 1 — Infrastructure. infra/compose.yaml gains Qdrant and Ollama:

qdrant:
image: qdrant/qdrant:v1.19.1
ports:
- "6333:6333" # REST API and web dashboard (http://localhost:6333/dashboard)
- "6334:6334" # gRPC, which the Java clients use
volumes:
- qdrant-data:/qdrant/storage
ollama:
# Runs models on the CPU inside Docker. On a Mac, a natively installed Ollama uses the GPU and is
# much faster; it listens on the same port.
image: ollama/ollama:0.35.1
ports:
- "11434:11434"
volumes:
- ollama-models:/root/.ollama

With spring.ai.ollama.init.pull-model-strategy: when_missing, the service downloads nomic-embed-text and the chat model on first start (about 2.8 GB together).

Step 2 — Fill the index. The catalog’s seed data came from a Flyway migration, so no events ever described it. The catalog gains a backfill endpoint that republishes every product’s current state (in one transaction, which is fine for sixteen products; a real catalog would page through them). Event-driven systems need a backfill path like this from day one: anything created outside the event stream doesn’t exist for consumers.

// Fragment of product-catalog-service/.../ProductService.kt
// Rows that never went through create(), such as Flyway's seed data, have no events. A republish
// emits every product's current state, which idempotent consumers absorb safely.
@Transactional
fun republishAll(): Int = products.findAll().onEach { product ->
events.publishEvent(ProductUpdated(Uuid.random(), product.sku, now(), product.snapshot()))
}.size
Terminal window
docker compose -f infra/compose.yaml up -d postgres kafka qdrant ollama
./gradlew :services:product-catalog-service:bootRun # port 8080
./gradlew :services:product-ai-service:bootRun # port 8081
curl -s -X POST localhost:8080/api/products/republish
# {"republished":16}

Within a few seconds, the AI service’s log shows sixteen Indexed product-updated lines, and Qdrant’s dashboard at localhost:6333/dashboard shows 19 points in products_springai, sixteen products and three policies, each a 768-dimensional vector compared by cosine distance. A point’s payload, through Qdrant’s REST API (text shortened):

Terminal window
curl -s -X POST localhost:6333/collections/products_springai/points/scroll -H 'Content-Type: application/json' \
-d '{"limit":1,"with_payload":true,"filter":{"must":[{"key":"sku","match":{"value":"SHW-1003"}}]}}'
# … "payload":{"type":"product","sku":"SHW-1003","name":"Gluten-Free Millet Pasta 250 g","category":"PASTA_AND_GRAINS",
# "dietaryTags":["gluten-free","vegan"],"pricePaise":19500,"doc_content":"Gluten-Free Millet Pasta 250 g (SKU SHW-1003)\nMade from mill…"}

The metadata sits flat in the payload, and the text is under doc_content, Spring AI’s key. Part 16 comes back to that key.

Step 3 — Search.

Terminal window
curl -s -G localhost:8081/api/assistant/search --data-urlencode "q=gluten-free pasta" \
--data-urlencode "maxPricePaise=20000" --data-urlencode "tag=gluten-free"
# [{"sku":"SHW-1003","name":"Gluten-Free Millet Pasta 250 g","pricePaise":19500,"score":0.7688671350479126}]

The quinoa pasta is gluten-free too and close to the query, but at ₹250 the filter removes it.

Step 4 — Ask. These are real answers from qwen3:4b-instruct on four CPU cores in Docker; timings include retrieval and, for the price, a tool call.

Terminal window
curl -s -X POST localhost:8081/api/assistant/ask -H 'Content-Type: application/json' \
-d '{"question":"Who won the cricket match yesterday?"}'
# {"text":"I can't help with that query.","sources":[]}
QuestionAnswerSourcesTime
”Is there any gluten-free pasta under ₹200?""Yes, the Gluten-Free Millet Pasta (SKU SHW-1003) is gluten-free and costs ₹195.00, which is under ₹200.”SHW-1003, SHW-1004, SHW-1005, SHW-100114.1 s
”My mixer grinder broke after three weeks. You have a 30-day no-questions-asked return policy, right?""Shelfwise does not have a “no-questions-asked” return policy for small appliances. While unused mixer grinders can be returned within 10 days of purchase for a full refund, the appliance must be in original packaging and the bill must be provided. After that period, returns are not allowed under any circumstances.”returns-small-appliances, returns-food-and-groceries5.8 s
”What is the delivery charge on a ₹300 order?""The delivery charge on a ₹300 order is ₹40, as delivery is free only for orders of ₹499 or more. …“delivery-and-orders, SHW-10107.7 s
”What does SHW-1001 cost right now?""The Durum Wheat Penne 500 g (SHW-1001) costs ₹165.00.” Repriced in the catalog to ₹159, the same question answered ₹159.00.none: the tool answered5.0 s
”Who won the cricket match yesterday?""I can’t help with that query.”none0.7 s

The second row is the demo rehearsal, rerun: the assistant contradicts the invented policy with the real one, and names it. It isn’t word-perfect. The policy says used appliances are not returnable under any circumstances; the answer turned that into “after that period, returns are not allowed under any circumstances”, close but not the same rule. Grounding narrows distortion, it doesn’t remove it. The answer also misses something a person would add, that a mixer grinder failing after three weeks is a warranty repair; small models answer the question asked, and the policy document’s warranty paragraph was in its context. A larger model, or a prompt that asks for “what the customer can do instead”, improves that, and it’s the kind of quality judgement Part 16’s comparison needs a fixed test set for.

Step 5 — Tests without a model.

The tests run the real Spring AI wiring, real Qdrant and real Kafka, with deterministic stand-ins for the models: BagOfWordsEmbeddingModel hashes words into a 512-dimensional vector, so texts that share words are close, and a chat model records the prompt it was given:

// Fragment of FakeModels.kt
// Records the prompt it was given and answers with a fixed sentence: enough to check what the
// RAG advisor put in front of the model.
class RecordingChatModel : ChatModel {
var lastPrompt: Prompt? = null
override fun call(prompt: Prompt): ChatResponse {
lastPrompt = prompt
return ChatResponse(listOf(Generation(AssistantMessage("(fake model answer)"))))
}
}

A composed annotation, @AssistantSpringTest, gives every test class the same configuration, and so one cached context and one set of containers. It sets spring.ai.model.chat=none and spring.ai.model.embedding=none, which switch off Ollama’s auto-configuration, and imports the containers; Qdrant’s @ServiceConnection comes from spring-ai-spring-boot-testcontainers:

// Fragment of TestcontainersConfiguration.kt
// Kafka and Qdrant, connected by @ServiceConnection. Qdrant's connection details come from
// spring-ai-spring-boot-testcontainers, which knows the vector store's properties.
@TestConfiguration(proxyBeanMethods = false)
class TestcontainersConfiguration {
@Bean
@ServiceConnection
fun kafka() = KafkaContainer("apache/kafka-native:4.2.2")
@Bean
@ServiceConnection
fun qdrant() = QdrantContainer("qdrant/qdrant:v1.19.1")
}
// Fragment of AssistantTest.kt
@Test
fun `answers are grounded in retrieved policies`() {
val answer = assistant.ask("Can I return a mixer grinder that stopped working?")
assertTrue("returns-small-appliances" in answer.sources)
val prompt = chatModel.lastPrompt!!.contents
assertTrue("There is no \"no questions asked\" return policy." in prompt) // the policy text reached the model
}
@Test
fun `with nothing relevant retrieved, the model is told to decline`() {
val answer = assistant.ask("Who won the cricket match yesterday?")
assertEquals(emptyList(), answer.sources)
val prompt = chatModel.lastPrompt!!.contents
assertTrue("No store documents match this query." in prompt)
assertTrue("Query: Who won the cricket match yesterday?" in prompt) // the question survives, for tools
}

What they can’t check is the wording of a real model’s answer, so they check what the model was given instead: the policy text in the prompt, or the no-documents instruction with the question intact. The suite also proves filtered search, idempotent replays and deletes, and that a record no deserializer can read is skipped rather than retried forever.


Tips, Tricks & Gotchas

Gotcha — Kotlin nullability decides “required” in a tool schema, until @ToolParam overrides it. Spring AI builds the tool’s JSON schema with Spring’s nullness support, which reads Kotlin’s types: fun stock(sku: String, store: String?) produces "required" : [ "sku" ], where a plain Java String would be required unless annotated @Nullable. But an explicit annotation is checked first, and @ToolParam’s required defaults to true: add @ToolParam(description = …) to store: String? and it’s required again. For an optional parameter, write @ToolParam(required = false) store: String? and let the method cope with null (ToolSchemaTest checks all three).

Gotcha — const val and trimIndent() don’t mix. Java’s text blocks strip incidental indentation at compile time; Kotlin’s raw strings keep it, and const val PROMPT = """…""".trimIndent() fails with “const ‘val’ initializer must be a constant value”. Prompts that annotations need (LangChain4j’s @SystemMessage, in Part 16) stay indented in the declaration and get trimmed where code uses them, as AssistantPrompts does.

Gotcha — a Long in metadata becomes a string in Qdrant. Spring AI 2.0.1’s Qdrant store writes Long values as strings (its source says so: “use String representation”), so the value comes back as "42", and a range filter, which the converter sends as a number, matches nothing. This bites Java’s boxed Long as much as Kotlin’s, and money in paise is exactly where JVM developers reach for Long. QdrantLongMetadataTest pins it down; the index stores pricePaise as an Int.

Gotcha — ContextualQueryAugmenter drops the question when nothing is retrieved. By default (allowEmptyContext is false), its empty-context prompt replaces the user’s query entirely. The model then can’t call a tool for a question it never saw: “What does SHW-1001 cost right now?” retrieved nothing above the threshold and was refused, until the custom augmenter kept the question.

Gotcha — qwen3:4b is now a thinking model. The plain tag resolves to the Qwen3 Thinking fine-tune (ollama show reports a thinking capability), which writes its reasoning into the answer, ending in </think>, and took 25–35 seconds per question here; spring.ai.ollama.chat.think: false doesn’t change that. qwen3:4b-instruct is the non-thinking variant: same size, tool support, plain answers in 5–15 seconds.

Gotcha — a similarity threshold belongs to one embedding model. Scores aren’t comparable across models: related products score about 0.65–0.8 with nomic-embed-text, while the tests’ bag-of-words embedding needed 0.25. Calibrate the threshold on real questions whenever the embedding model changes, and re-embed the whole collection at the same time.


Debugging and Troubleshooting

SymptomLikely causeFix
The assistant answers confidently about things that aren’t in the store’s documentsNo retrieval, or allowEmptyContext(true) with a permissive promptRAG with a similarity threshold; an explicit no-documents instruction
Every answer is “I can’t help”Threshold too high for this embedding model, or nothing indexedCheck vectorStore.similaritySearch scores; check the collection’s point count
A numeric metadata filter matches nothingThe field was stored as a string (a Long value)Store Int or Double
The model never calls a toolThe prompt says “only from the context”, or the empty-context prompt dropped the questionMention the tool in the system prompt; keep the question in the empty case
The consumer logs the same deserialization error foreverA poison pill without ErrorHandlingDeserializerWrap the deserializer; the default error handler logs and skips
Retrieval quality dropped after a configuration changeDifferent embedding model than the one that built the collectionSame model everywhere; new model, new collection
Startup fails connecting to localhost:11434Ollama not running, or a model still downloadingdocker compose up -d ollama; watch the pull in the log
io.qdrant NoSuchMethodError or gRPC errorsTwo Qdrant client versions on the classpathPin one (io.qdrant:client:1.19.0) in dependency management

Key Takeaways

ConceptRemember
RAGCorrectness comes from the inputs: retrieval, a threshold for “nothing relevant”, tools for live data
Ports and adaptersProductIndexer and ProductAssistant in the service’s terms; one collection and one consumer group per adapter
Idempotent indexingName-based UUIDs from the SKU: redeliveries, republishes and replays overwrite
Kafka consumerBlocking listener, exhaustive when, ErrorHandlingDeserializer for poison pills
Spring AI pipelineVectorStoreDocumentRetriever + a query augmenter + RetrievalAugmentationAdvisor; sources from DOCUMENT_CONTEXT
Kotlin and Java APIsA lambda implements Spring AI’s QueryAugmenter; run and reduce assemble filters
Tools@Tool + @ToolParam on a Kotlin method; mention them in the prompt
TrapsLong metadata → string in Qdrant; empty-context prompt drops the question; @ToolParam overrides nullability; qwen3:4b thinks
TestingReal Qdrant and Kafka, fake models: assert what the model was given
OperationsSame embedding model everywhere; Micrometer meters for tokens and latency

Story Closing

The rerun of the demo went the way demos are supposed to go. The assistant found the millet pasta, answered from the real return policy instead of inventing one, checked a live price through the catalog, and declined to talk about cricket. The AI team’s lead asked for the test suite before she asked for anything else.

The architecture review board met the following Tuesday. The slides were fine; the first question was not about them. “Spring AI reached 1.0 in May last year and 2.0 this summer. LangChain4j moves faster. What if we’ve picked the wrong one?”

Kabir had expected the question. The service talked to Spring AI through two interfaces, almost everywhere. The exceptions were small and he knew them: the @Tool annotations on CatalogTools, and the {context} and {query} placeholders in the shared prompts, which are Spring AI’s template syntax.

In Part 16, the final part, Kabir implements the same ports with LangChain4j, runs one test suite against both adapters, and compares them honestly: abstractions, RAG flexibility, Spring integration, Kotlin ergonomics and upgrade cadence.


This is Part 15 of a 16-part series: “Kotlin for Java Survivors: Life After Semicolons.”