"The Assistant" — Spring AI 2.0, Qdrant and RAG
The prototype invented a return policy, and Kabir builds the service that can't. We build product-ai-service on Spring AI 2.0: a Kafka consumer that keeps a Qdrant index in step with the catalog, filtered semantic search, answers grounded in store policies, a tool for live prices, and a local model through Ollama, with the traps we hit on the way.
Story Opening
The prototype was a notebook and one long prompt: “You are a helpful assistant for Shelfwise”, followed by every product in the catalog, pasted in as text. Nothing in it mentioned returns, warranties or delivery, so when someone asked about a broken mixer grinder, the model did what language models do with a gap. It filled it, plausibly.
Kabir’s first instinct was the one every developer has: fix the prompt. Add “never make things up”, maybe in capitals. Lena shook her head before he’d finished the sentence. “It can’t decline to answer from a document it’s never seen. Give it the documents, and give it a way to say it has none.”
Java → Kotlin: The Quick Map
This part maps familiar Spring and search concepts onto Spring AI 2.0:
| What you know | Spring AI 2.0 | Note |
|---|---|---|
A DataSource bean configured from properties | ChatModel and EmbeddingModel beans from a starter | Ollama here; a property switch for hosted providers |
RestTemplate calls to a model’s HTTP API | ChatClient fluent API | prompt().user(…).call() |
| Lucene or Elasticsearch full-text search | VectorStore.similaritySearch | Meaning, not matching words |
A SQL WHERE clause | FilterExpressionBuilder | Metadata filters, translated to Qdrant’s |
| A servlet filter chain around a call | Advisors around a ChatClient call | RetrievalAugmentationAdvisor adds retrieved context |
| Hand-written JSON for “function calling” | @Tool methods | Spring generates the schema and runs the call |
@KafkaListener with a Jackson payload | @KafkaListener + ErrorHandlingDeserializer + an exhaustive when | The shared sealed events from Part 14 |
Conceptual Deep-Dive
A model is a function of its prompt
A chat model doesn’t look anything up. It continues the text it’s given, and its answer is only as good as that text. So the work in a product assistant is not in the model call but around it: deciding what goes into the prompt, and what the model may call. That pattern is retrieval-augmented generation (RAG), and it has two halves:
product-events")] --> L["@KafkaListener"] P["policies/*.md"] --> I L --> I["ProductIndexer"] I -->|"text → vector"| E1["EmbeddingModel
nomic-embed-text"] E1 --> Q[("Qdrant
products_springai")] end subgraph Ask["Ask (per question)"] U["Question"] --> E2["EmbeddingModel"] E2 -->|"nearest vectors,
score ≥ 0.6"| Q Q --> A["Augment the prompt"] A --> C["ChatModel
qwen3:4b-instruct"] C <-->|"@Tool"| T["Catalog API
live price"] C --> R["Answer + sources"] end
An embedding model turns a text into a vector, here 768 numbers, such that texts with similar meaning point in similar directions. Indexing stores each product’s and policy’s vector in Qdrant with its text and metadata. Asking embeds the question the same way, finds the nearest stored vectors, and puts their text into the prompt as context. “Gluten-free pasta under ₹200” lands near the millet pasta’s description although the words differ, and a metadata filter can still demand a price under 20,000 paise exactly.
The shift for a Java developer: you don’t make an LLM correct by instructing it; you make it correct by controlling its inputs. Retrieval decides what it knows, a similarity threshold decides when it knows nothing, and tools decide what it can look up. The prompt’s instructions only tell it how to use those.
One port, two adapters
The service talks to AI libraries through two Kotlin interfaces, and Part 16 implements them a second time with LangChain4j:
// Fragment of ports/Ports.kt// What the service needs from an AI library, in its own words. Part 15 implements both ports with// Spring AI; Part 16 implements them again with LangChain4j, behind a profile.interface ProductIndexer { fun upsert(product: ProductSnapshot)
fun delete(sku: String)
fun upsertPolicy(name: String, text: String)}
data class ProductHit(val sku: String, val name: String, val pricePaise: Long, val score: Double)
data class Answer(val text: String, val sources: List<String>)
interface ProductAssistant { fun search(query: String, maxPricePaise: Long? = null, dietaryTag: String? = null, limit: Int = 5): List<ProductHit>
fun ask(question: String): Answer}The adapters never share a request path, and they don’t share storage either. Each gets its own Qdrant collection (products_springai here), because the two libraries store a document’s text under different payload keys and each would read the other’s points as empty. Each gets its own Kafka consumer group (product-ai-springai): a group with no committed offsets starts at the earliest retained record (with auto.offset.reset=earliest; Kafka’s default is latest), so switching profiles replays the topic and rebuilds that adapter’s index, as far back as the topic’s retention reaches (seven days by default; the catalog’s republish fills in the rest). Both use the same embedding model; vectors from different models aren’t comparable, and mixing them degrades retrieval without any error.
Technical Explanation
The build: Spring AI 2.0 on Boot 4.1
// Fragment of build.gradle.ktsdependencyManagement { imports { mavenBom("org.springframework.ai:spring-ai-bom:${libs.findVersion("spring-ai").get().requiredVersion}") } dependencies { // One Qdrant client for both AI libraries, matching the server's minor version. dependency("io.qdrant:client:${libs.findVersion("qdrant-client").get().requiredVersion}") }}// Fragment of build.gradle.kts // Spring AI: a chat and an embedding model from Ollama, Qdrant as the vector store, and the RAG building blocks. implementation("org.springframework.ai:spring-ai-starter-model-ollama") implementation("org.springframework.ai:spring-ai-starter-vector-store-qdrant") implementation("org.springframework.ai:spring-ai-rag")Spring AI 2.0.1 is the current release for Boot 4.0 and 4.1 (2.1 milestones target Boot 4.2); its BOM manages the starters. The pinned Qdrant client solves a dependency puzzle: Spring AI 2.0.1’s Qdrant store depends on io.qdrant:client 1.18.0, LangChain4j 1.21’s on 1.17.0, and the server is 1.19.1. Qdrant releases its Java client in step with the server and recommends matching minor versions, so the build pins 1.19.0: the Spring AI tests run against it here, and Part 16’s adapter shares the pin.
Documents with stable identities
// Fragment of ports/IndexedDocuments.kt// How products and policies become documents, shared by both adapters so they index the same text.object IndexedDocuments { // Deterministic point ids: the same SKU always maps to the same point, so a redelivered or // replayed event overwrites instead of duplicating. That's the idempotence Kafka needs. fun productId(sku: String): String = UUID.nameUUIDFromBytes("product:$sku".encodeToByteArray()).toString()
fun policyId(name: String): String = UUID.nameUUIDFromBytes("policy:$name".encodeToByteArray()).toString()
fun productText(p: ProductSnapshot): String = buildString { appendLine("${p.name} (SKU ${p.sku})") if (p.description.isNotBlank()) appendLine(p.description) append("Category: ${p.category.lowercase().replace('_', ' ')}. ") append("Price: ₹%d.%02d.".format(p.pricePaise / 100, p.pricePaise % 100)) if (p.dietaryTags.isNotEmpty()) append(" Dietary: ${p.dietaryTags.joinToString()}.") }
fun productMetadata(p: ProductSnapshot): Map<String, Any> = mapOf( "type" to "product", "sku" to p.sku, "name" to p.name, "category" to p.category, // An Int, not a Long: Spring AI's Qdrant store writes Long values as strings, and a numeric // range filter on a string field matches nothing. Paise fit in an Int up to ₹2 crore. "pricePaise" to Math.toIntExact(p.pricePaise), "dietaryTags" to p.dietaryTags, )}The point id is a name-based UUID derived from the SKU: the same product always maps to the same point. Kafka delivers at least once, Part 14’s events are keyed by SKU so one product’s create, updates and delete arrive in order, the republish sends every product again, and a new consumer group replays what the topic retains; each of those overwrites a point instead of adding a duplicate. That’s the consumer side of Part 14’s idempotence requirement, achieved without a deduplication table. The text is what gets embedded; the metadata is what filters match. pricePaise is stored as an Int, because Spring AI’s Qdrant store writes a Long as a string. (Math.toIntExact would throw for a price above ₹2 crore, and the error handler would retry, then skip the event; a Double would avoid that edge entirely.)
Consuming product events
// Fragment of ingest/ProductEventListener.kt// Keeps the vector index in step with the catalog. Exhaustive over the sealed hierarchy: a new event// type in shared/product-events doesn't compile here until it's handled.@Componentclass ProductEventListener(private val indexer: ProductIndexer) { private val log = LoggerFactory.getLogger(javaClass)
// A plain (blocking) listener: indexing calls the embedding model and Qdrant synchronously. With the // default AckMode.BATCH, offsets are committed once every record from a poll has been processed. @KafkaListener(topics = [ProductEvents.TOPIC], groupId = "\${assistant.consumer-group}") fun on(event: ProductEvent) { when (event) { is ProductCreated -> indexer.upsert(event.product) is ProductUpdated -> indexer.upsert(event.product) is ProductDeleted -> indexer.delete(event.sku) } log.info("Indexed {} for {}", event.type, event.sku) }}The when is exhaustive over Part 14’s sealed hierarchy, so a new event type fails this module’s compilation once it upgrades shared/product-events. What it can’t do is decode an event type that’s newer than the module it was built with, and Kafka will keep redelivering a record the deserializer can’t read: the classic poison pill. Spring Kafka’s ErrorHandlingDeserializer wraps the real deserializer and turns that failure into a record the default error handler logs and skips:
kafka: bootstrap-servers: localhost:9092 consumer: auto-offset-reset: earliest # a new consumer group replays the topic and rebuilds the index value-deserializer: org.springframework.kafka.support.serializer.ErrorHandlingDeserializer properties: spring.deserializer.value.delegate.class: com.shelfwise.events.ProductEventDeserializerSpring Kafka also accepts suspend listeners, but then it switches the container to manual acknowledgement with out-of-order commits, and in-flight work can outlive a container stop (spring-kafka issue #4519); with synchronous indexing there’s nothing to gain. For throughput, a batch listener (List<ProductEvent>) and one vectorStore.add per batch would be the next step; Ollama embeds a batch in one request.
Indexing with Spring AI
// Fragment of springai/SpringAiProductIndexer.kt@Component@Profile("!lc4j")class SpringAiProductIndexer(private val vectorStore: VectorStore) : ProductIndexer { // add() embeds the text with the EmbeddingModel and upserts the point: same id, new vector and payload. override fun upsert(product: ProductSnapshot) = vectorStore.add( listOf( Document.builder() .id(IndexedDocuments.productId(product.sku)) .text(IndexedDocuments.productText(product)) .metadata(IndexedDocuments.productMetadata(product)) .build(), ), )
override fun delete(sku: String) = vectorStore.delete(listOf(IndexedDocuments.productId(sku)))
override fun upsertPolicy(name: String, text: String) = vectorStore.add( listOf( Document.builder() .id(IndexedDocuments.policyId(name)) .text(text) .metadata(mapOf("type" to "policy", "policy" to name)) .build(), ), )}VectorStore.add embeds each document with the configured EmbeddingModel and upserts it into Qdrant; delete takes the same ids. With initialize-schema: true, the store creates the collection on startup, sized to the embedding model’s dimensions (768 for nomic-embed-text). Policies go through the same port: PolicyLoader indexes classpath:policies/*.md at startup with deterministic ids, so restarts overwrite them.
Searching and answering
// Fragment of springai/SpringAiProductAssistant.kt@Component@Profile("!lc4j")class SpringAiProductAssistant( private val vectorStore: VectorStore, chatClient: ChatClient.Builder, catalogTools: CatalogTools, @Value("\${assistant.rag.similarity-threshold}") similarityThreshold: Double,) : ProductAssistant { private val withContext = ContextualQueryAugmenter.builder() .promptTemplate(PromptTemplate(AssistantPrompts.WITH_CONTEXT.trimIndent())) .build()
private val chat = chatClient .defaultSystem(AssistantPrompts.SYSTEM.trimIndent()) .defaultAdvisors( RetrievalAugmentationAdvisor.builder() .documentRetriever( VectorStoreDocumentRetriever.builder() .vectorStore(vectorStore) .similarityThreshold(similarityThreshold) // weak matches are worse than none .topK(4) .build(), ) // With nothing relevant retrieved, the model must not improvise a return policy, but a // tool may still answer. ContextualQueryAugmenter's own empty-context prompt drops the // user's question, so the empty case is handled here: a Kotlin lambda for a Java interface. .queryAugmenter { query, documents -> if (documents.isEmpty()) Query(AssistantPrompts.withoutContext(query.text())) else withContext.augment(query, documents) } .build(), ) .defaultTools(catalogTools) .build()
override fun search(query: String, maxPricePaise: Long?, dietaryTag: String?, limit: Int): List<ProductHit> { val filter = FilterExpressionBuilder().run { listOfNotNull( eq("type", "product"), maxPricePaise?.let { lte("pricePaise", it) }, dietaryTag?.let { eq("dietaryTags", it) }, // matches if any element of the list equals it ).reduce { all, condition -> and(all, condition) }.build() } val request = SearchRequest.builder().query(query).topK(limit).filterExpression(filter).build() return vectorStore.similaritySearch(request).map { doc -> ProductHit( sku = doc.metadata["sku"] as String, name = doc.metadata["name"] as String, pricePaise = (doc.metadata["pricePaise"] as Number).toLong(), score = doc.score ?: 0.0, ) } }
override fun ask(question: String): Answer { val response = chat.prompt().user(question).call().chatClientResponse() @Suppress("UNCHECKED_CAST") val sources = response.context()[RetrievalAugmentationAdvisor.DOCUMENT_CONTEXT] as List<Document>? ?: emptyList() return Answer( text = response.chatResponse()?.result?.output?.text.orEmpty(), sources = sources.map { it.metadata["sku"] as String? ?: it.metadata["policy"] as String }.distinct(), ) }}search is semantic search with exact filters. FilterExpressionBuilder builds a portable expression, which the Qdrant store translates into Qdrant’s own filter; run and reduce combine however many conditions the caller supplied. lte("pricePaise", it) passes a Long and still matches the Int stored in the payload, because the converter parses every range value as a Double: filters accept a Long, but writes turn one into a string.
ask assembles the RAG pipeline from Spring AI’s building blocks. VectorStoreDocumentRetriever fetches up to four documents scoring at least the configured threshold. A query augmenter then rewrites the question to include them: with documents, ContextualQueryAugmenter and the WITH_CONTEXT template; without any, an instruction not to improvise, plus the question itself. The augmenter is a Kotlin lambda because QueryAugmenter is a Java interface with one abstract method, which Kotlin converts automatically (a Kotlin interface would need to be declared fun interface, Part 5). defaultTools registers the catalog tools, and the advisor returns the retrieved documents in the response context under DOCUMENT_CONTEXT, which become the answer’s sources. Those make a wrong answer diagnosable: was the right document not retrieved, or retrieved and ignored? The ?. chain in ask is there because Spring AI 2.0 annotates its API with JSpecify, so chatResponse() and friends are honestly nullable in Kotlin instead of platform types (Parts 2 and 12).
// Fragment of ports/AssistantPrompts.kt// The instructions both adapters give the model, so Part 16 compares libraries, not prompts.object AssistantPrompts { const val SYSTEM = """ You are Shelfwise's product assistant for store staff and customers. Answer from the context you are given and from your tools; never invent policies, prices or products. When asked what a product costs now, call the currentPrice tool with its SKU. Prices are in Indian rupees. Mention SKUs when you recommend products. Keep answers to one or two sentences. """
// {context} and {query} are filled in by the RAG step. const val WITH_CONTEXT = """ Context information is below. --------------------- {context} --------------------- Answer the query using the context, and your tools where they help. If neither the context nor a tool gives the answer, say that you don't know. Don't start with phrases like "Based on the context".
Query: {query}
Answer: """
// Used when retrieval finds nothing relevant: no improvising, but a tool may still answer. fun withoutContext(query: String) = """ No store documents match this query. If one of your tools can answer it, use the tool. Otherwise, politely say that you can't help with it.
Query: $query """.trimIndent()}The {context} and {query} placeholders are Spring AI’s template syntax. The texts are const vals so that annotations can use them too (Part 16’s LangChain4j service takes its system message from one), which is why .trimIndent() appears at each use site instead of in the declaration.
A tool for live prices
// Fragment of tools/CatalogTools.ktdata class CatalogProduct(val sku: String, val name: String, val price: String)
// The catalog's REST API, as an HTTP interface (Part 14). Blocking is fine here: tools run on the// thread that serves the request, a virtual thread.interface CatalogClient { @GetExchange("/api/products/{sku}") fun product(@PathVariable sku: String): CatalogProduct}
// Indexed prices can be minutes old; a tool lets the model ask the catalog for the price right now.@Componentclass CatalogTools(private val catalog: CatalogClient) { @Tool(description = "Get the current price of a Shelfwise product from the catalog, by SKU. Use it when asked what something costs now.") fun currentPrice(@ToolParam(description = "The product SKU, for example SHW-1001") sku: String): String = try { catalog.product(sku).let { "${it.name} (${it.sku}) costs ${it.price}" } } catch (e: HttpClientErrorException.NotFound) { "There is no product with SKU $sku" }}@Tool turns the method into a function the model can call: Spring AI derives a JSON schema from the Kotlin signature and the @ToolParam description, sends it with the request, runs the method when the model asks for it, and sends the result back for the final answer. Kotlin’s nullability can shape that schema, with one catch the Gotchas show. If the catalog is down, the RestClient call throws, and by default Spring AI sends the exception’s message back to the model as the tool’s result rather than failing the request. The model then usually says it couldn’t check the price, but the message (for a connection failure, the catalog’s internal URL) is now in the model’s context and can surface in an answer; catch it in the tool and return something neutral when that matters.
Observability
Spring AI records model calls, vector searches and advisors as Micrometer observations. With the metrics endpoint exposed, one question produced these meters: gen_ai.client.operation and gen_ai.client.token.usage (tagged with gen_ai.operation.name chat or embedding, the model names, and gen_ai.token.type input, output or total), db.vector.client.operation, spring.ai.chat.client and spring.ai.advisor. Token usage per model is the number to watch once a hosted model costs money.
Step-by-Step Hands-On: Ask It Something
Code: kotlin-for-java-survivors/services/product-ai-service, tag kotlin-for-java-survivors/part-15.
Step 1 — Infrastructure. infra/compose.yaml gains Qdrant and Ollama:
qdrant: image: qdrant/qdrant:v1.19.1 ports: - "6333:6333" # REST API and web dashboard (http://localhost:6333/dashboard) - "6334:6334" # gRPC, which the Java clients use volumes: - qdrant-data:/qdrant/storage
ollama: # Runs models on the CPU inside Docker. On a Mac, a natively installed Ollama uses the GPU and is # much faster; it listens on the same port. image: ollama/ollama:0.35.1 ports: - "11434:11434" volumes: - ollama-models:/root/.ollamaWith spring.ai.ollama.init.pull-model-strategy: when_missing, the service downloads nomic-embed-text and the chat model on first start (about 2.8 GB together).
Step 2 — Fill the index. The catalog’s seed data came from a Flyway migration, so no events ever described it. The catalog gains a backfill endpoint that republishes every product’s current state (in one transaction, which is fine for sixteen products; a real catalog would page through them). Event-driven systems need a backfill path like this from day one: anything created outside the event stream doesn’t exist for consumers.
// Fragment of product-catalog-service/.../ProductService.kt // Rows that never went through create(), such as Flyway's seed data, have no events. A republish // emits every product's current state, which idempotent consumers absorb safely. @Transactional fun republishAll(): Int = products.findAll().onEach { product -> events.publishEvent(ProductUpdated(Uuid.random(), product.sku, now(), product.snapshot())) }.sizedocker compose -f infra/compose.yaml up -d postgres kafka qdrant ollama./gradlew :services:product-catalog-service:bootRun # port 8080./gradlew :services:product-ai-service:bootRun # port 8081
curl -s -X POST localhost:8080/api/products/republish# {"republished":16}Within a few seconds, the AI service’s log shows sixteen Indexed product-updated lines, and Qdrant’s dashboard at localhost:6333/dashboard shows 19 points in products_springai, sixteen products and three policies, each a 768-dimensional vector compared by cosine distance. A point’s payload, through Qdrant’s REST API (text shortened):
curl -s -X POST localhost:6333/collections/products_springai/points/scroll -H 'Content-Type: application/json' \ -d '{"limit":1,"with_payload":true,"filter":{"must":[{"key":"sku","match":{"value":"SHW-1003"}}]}}'# … "payload":{"type":"product","sku":"SHW-1003","name":"Gluten-Free Millet Pasta 250 g","category":"PASTA_AND_GRAINS",# "dietaryTags":["gluten-free","vegan"],"pricePaise":19500,"doc_content":"Gluten-Free Millet Pasta 250 g (SKU SHW-1003)\nMade from mill…"}The metadata sits flat in the payload, and the text is under doc_content, Spring AI’s key. Part 16 comes back to that key.
Step 3 — Search.
curl -s -G localhost:8081/api/assistant/search --data-urlencode "q=gluten-free pasta" \ --data-urlencode "maxPricePaise=20000" --data-urlencode "tag=gluten-free"# [{"sku":"SHW-1003","name":"Gluten-Free Millet Pasta 250 g","pricePaise":19500,"score":0.7688671350479126}]The quinoa pasta is gluten-free too and close to the query, but at ₹250 the filter removes it.
Step 4 — Ask. These are real answers from qwen3:4b-instruct on four CPU cores in Docker; timings include retrieval and, for the price, a tool call.
curl -s -X POST localhost:8081/api/assistant/ask -H 'Content-Type: application/json' \ -d '{"question":"Who won the cricket match yesterday?"}'# {"text":"I can't help with that query.","sources":[]}| Question | Answer | Sources | Time |
|---|---|---|---|
| ”Is there any gluten-free pasta under ₹200?" | "Yes, the Gluten-Free Millet Pasta (SKU SHW-1003) is gluten-free and costs ₹195.00, which is under ₹200.” | SHW-1003, SHW-1004, SHW-1005, SHW-1001 | 14.1 s |
| ”My mixer grinder broke after three weeks. You have a 30-day no-questions-asked return policy, right?" | "Shelfwise does not have a “no-questions-asked” return policy for small appliances. While unused mixer grinders can be returned within 10 days of purchase for a full refund, the appliance must be in original packaging and the bill must be provided. After that period, returns are not allowed under any circumstances.” | returns-small-appliances, returns-food-and-groceries | 5.8 s |
| ”What is the delivery charge on a ₹300 order?" | "The delivery charge on a ₹300 order is ₹40, as delivery is free only for orders of ₹499 or more. …“ | delivery-and-orders, SHW-1010 | 7.7 s |
| ”What does SHW-1001 cost right now?" | "The Durum Wheat Penne 500 g (SHW-1001) costs ₹165.00.” Repriced in the catalog to ₹159, the same question answered ₹159.00. | none: the tool answered | 5.0 s |
| ”Who won the cricket match yesterday?" | "I can’t help with that query.” | none | 0.7 s |
The second row is the demo rehearsal, rerun: the assistant contradicts the invented policy with the real one, and names it. It isn’t word-perfect. The policy says used appliances are not returnable under any circumstances; the answer turned that into “after that period, returns are not allowed under any circumstances”, close but not the same rule. Grounding narrows distortion, it doesn’t remove it. The answer also misses something a person would add, that a mixer grinder failing after three weeks is a warranty repair; small models answer the question asked, and the policy document’s warranty paragraph was in its context. A larger model, or a prompt that asks for “what the customer can do instead”, improves that, and it’s the kind of quality judgement Part 16’s comparison needs a fixed test set for.
Step 5 — Tests without a model.
The tests run the real Spring AI wiring, real Qdrant and real Kafka, with deterministic stand-ins for the models: BagOfWordsEmbeddingModel hashes words into a 512-dimensional vector, so texts that share words are close, and a chat model records the prompt it was given:
// Fragment of FakeModels.kt// Records the prompt it was given and answers with a fixed sentence: enough to check what the// RAG advisor put in front of the model.class RecordingChatModel : ChatModel { var lastPrompt: Prompt? = null
override fun call(prompt: Prompt): ChatResponse { lastPrompt = prompt return ChatResponse(listOf(Generation(AssistantMessage("(fake model answer)")))) }}A composed annotation, @AssistantSpringTest, gives every test class the same configuration, and so one cached context and one set of containers. It sets spring.ai.model.chat=none and spring.ai.model.embedding=none, which switch off Ollama’s auto-configuration, and imports the containers; Qdrant’s @ServiceConnection comes from spring-ai-spring-boot-testcontainers:
// Fragment of TestcontainersConfiguration.kt// Kafka and Qdrant, connected by @ServiceConnection. Qdrant's connection details come from// spring-ai-spring-boot-testcontainers, which knows the vector store's properties.@TestConfiguration(proxyBeanMethods = false)class TestcontainersConfiguration { @Bean @ServiceConnection fun kafka() = KafkaContainer("apache/kafka-native:4.2.2")
@Bean @ServiceConnection fun qdrant() = QdrantContainer("qdrant/qdrant:v1.19.1")}// Fragment of AssistantTest.kt @Test fun `answers are grounded in retrieved policies`() { val answer = assistant.ask("Can I return a mixer grinder that stopped working?")
assertTrue("returns-small-appliances" in answer.sources) val prompt = chatModel.lastPrompt!!.contents assertTrue("There is no \"no questions asked\" return policy." in prompt) // the policy text reached the model }
@Test fun `with nothing relevant retrieved, the model is told to decline`() { val answer = assistant.ask("Who won the cricket match yesterday?")
assertEquals(emptyList(), answer.sources) val prompt = chatModel.lastPrompt!!.contents assertTrue("No store documents match this query." in prompt) assertTrue("Query: Who won the cricket match yesterday?" in prompt) // the question survives, for tools }What they can’t check is the wording of a real model’s answer, so they check what the model was given instead: the policy text in the prompt, or the no-documents instruction with the question intact. The suite also proves filtered search, idempotent replays and deletes, and that a record no deserializer can read is skipped rather than retried forever.
Tips, Tricks & Gotchas
Gotcha — Kotlin nullability decides “required” in a tool schema, until
@ToolParamoverrides it. Spring AI builds the tool’s JSON schema with Spring’s nullness support, which reads Kotlin’s types:fun stock(sku: String, store: String?)produces"required" : [ "sku" ], where a plain JavaStringwould be required unless annotated@Nullable. But an explicit annotation is checked first, and@ToolParam’srequireddefaults totrue: add@ToolParam(description = …)tostore: String?and it’s required again. For an optional parameter, write@ToolParam(required = false) store: String?and let the method cope withnull(ToolSchemaTestchecks all three).
Gotcha —
const valandtrimIndent()don’t mix. Java’s text blocks strip incidental indentation at compile time; Kotlin’s raw strings keep it, andconst val PROMPT = """…""".trimIndent()fails with “const ‘val’ initializer must be a constant value”. Prompts that annotations need (LangChain4j’s@SystemMessage, in Part 16) stay indented in the declaration and get trimmed where code uses them, asAssistantPromptsdoes.
Gotcha — a
Longin metadata becomes a string in Qdrant. Spring AI 2.0.1’s Qdrant store writesLongvalues as strings (its source says so: “use String representation”), so the value comes back as"42", and a range filter, which the converter sends as a number, matches nothing. This bites Java’s boxedLongas much as Kotlin’s, and money in paise is exactly where JVM developers reach forLong.QdrantLongMetadataTestpins it down; the index storespricePaiseas anInt.
Gotcha —
ContextualQueryAugmenterdrops the question when nothing is retrieved. By default (allowEmptyContextisfalse), its empty-context prompt replaces the user’s query entirely. The model then can’t call a tool for a question it never saw: “What does SHW-1001 cost right now?” retrieved nothing above the threshold and was refused, until the custom augmenter kept the question.
Gotcha —
qwen3:4bis now a thinking model. The plain tag resolves to the Qwen3 Thinking fine-tune (ollama showreports athinkingcapability), which writes its reasoning into the answer, ending in</think>, and took 25–35 seconds per question here;spring.ai.ollama.chat.think: falsedoesn’t change that.qwen3:4b-instructis the non-thinking variant: same size, tool support, plain answers in 5–15 seconds.
Gotcha — a similarity threshold belongs to one embedding model. Scores aren’t comparable across models: related products score about 0.65–0.8 with
nomic-embed-text, while the tests’ bag-of-words embedding needed 0.25. Calibrate the threshold on real questions whenever the embedding model changes, and re-embed the whole collection at the same time.
Debugging and Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| The assistant answers confidently about things that aren’t in the store’s documents | No retrieval, or allowEmptyContext(true) with a permissive prompt | RAG with a similarity threshold; an explicit no-documents instruction |
| Every answer is “I can’t help” | Threshold too high for this embedding model, or nothing indexed | Check vectorStore.similaritySearch scores; check the collection’s point count |
| A numeric metadata filter matches nothing | The field was stored as a string (a Long value) | Store Int or Double |
| The model never calls a tool | The prompt says “only from the context”, or the empty-context prompt dropped the question | Mention the tool in the system prompt; keep the question in the empty case |
| The consumer logs the same deserialization error forever | A poison pill without ErrorHandlingDeserializer | Wrap the deserializer; the default error handler logs and skips |
| Retrieval quality dropped after a configuration change | Different embedding model than the one that built the collection | Same model everywhere; new model, new collection |
Startup fails connecting to localhost:11434 | Ollama not running, or a model still downloading | docker compose up -d ollama; watch the pull in the log |
io.qdrant NoSuchMethodError or gRPC errors | Two Qdrant client versions on the classpath | Pin one (io.qdrant:client:1.19.0) in dependency management |
Key Takeaways
| Concept | Remember |
|---|---|
| RAG | Correctness comes from the inputs: retrieval, a threshold for “nothing relevant”, tools for live data |
| Ports and adapters | ProductIndexer and ProductAssistant in the service’s terms; one collection and one consumer group per adapter |
| Idempotent indexing | Name-based UUIDs from the SKU: redeliveries, republishes and replays overwrite |
| Kafka consumer | Blocking listener, exhaustive when, ErrorHandlingDeserializer for poison pills |
| Spring AI pipeline | VectorStoreDocumentRetriever + a query augmenter + RetrievalAugmentationAdvisor; sources from DOCUMENT_CONTEXT |
| Kotlin and Java APIs | A lambda implements Spring AI’s QueryAugmenter; run and reduce assemble filters |
| Tools | @Tool + @ToolParam on a Kotlin method; mention them in the prompt |
| Traps | Long metadata → string in Qdrant; empty-context prompt drops the question; @ToolParam overrides nullability; qwen3:4b thinks |
| Testing | Real Qdrant and Kafka, fake models: assert what the model was given |
| Operations | Same embedding model everywhere; Micrometer meters for tokens and latency |
Story Closing
The rerun of the demo went the way demos are supposed to go. The assistant found the millet pasta, answered from the real return policy instead of inventing one, checked a live price through the catalog, and declined to talk about cricket. The AI team’s lead asked for the test suite before she asked for anything else.
The architecture review board met the following Tuesday. The slides were fine; the first question was not about them. “Spring AI reached 1.0 in May last year and 2.0 this summer. LangChain4j moves faster. What if we’ve picked the wrong one?”
Kabir had expected the question. The service talked to Spring AI through two interfaces, almost everywhere. The exceptions were small and he knew them: the @Tool annotations on CatalogTools, and the {context} and {query} placeholders in the shared prompts, which are Spring AI’s template syntax.
In Part 16, the final part, Kabir implements the same ports with LangChain4j, runs one test suite against both adapters, and compares them honestly: abstractions, RAG flexibility, Spring integration, Kotlin ergonomics and upgrade cadence.
This is Part 15 of a 16-part series: “Kotlin for Java Survivors: Life After Semicolons.”