AI message not showing

#1 · closed · 6 comments

View on GitHub ↗

Khim3

I have cloned this repo, change the DENSE_MODEL from using Hugging face to OllamaEmbeddings (nomix-text-embed). Then i run the app but after uploading document, creating vector db store and typing my message, the AI message did not show up. DENSE_MODEL = "nomic-embed-text" SPARSE_MODEL = "Qdrant/bm25" LLM_MODEL = "llama3.2" LLM_TEMPERATURE = 0 import config from langchain_ollama import OllamaEmbeddings from langchain_qdrant import QdrantVectorStore, FastEmbedSparse, RetrievalMode from qdrant_client import QdrantClient from qdrant_client.http import models as qmodels class VectorDbManager: __client: QdrantClient __dense_embeddings: OllamaEmbeddings __sparse_embeddings: FastEmbedSparse def __init__(self): self.__client = QdrantClient(path=config.QDRANT_DB_PATH) self.__dense_embeddings = OllamaEmbeddings(model=config.DENSE_MODEL) self.__sparse_embeddings = FastEmbedSparse(model_name=config.SPARSE_MODEL) <img width="1890" height="1454" alt="Image" src="https://github.com/user-attachments/assets/5d030067-2ec5-4239-a629-333280d8a226" />

Comments

GiovanniPasq

Hi, To help you as best as possible, I’d like to ask you the following: 1. Have you checked that the markdown files are correctly extracted from the PDFs? PymuPdf4llm is a lightweight and very fast library, but it doesn’t work particularly well with PDFs that have complex layouts or many images. 2. Could you log the output of the chunk-retrieval and parent-retrieval tools? 3. I see that you wrote “llama3.2” in the config. After checking the models available on Ollama, are you sure it’s written correctly? On the Ollama page there is no model with that exact name. You can check the available tags here: https://ollama.com/library/llama3.2/tags You should specify something like “llama3.2:1b” or “llama3.2:latest”, etc. Also, remember that before using a model you need to run the command ollama pull "model-name". Let me know, Giovanni

Khim3

1. Yes, PymuPdf4llm is working properly 2. when running the app, there is no error appears, so I don't know where to log 3. with Ollama, when you say llama3.2, it would automatically run the latest version (3b version). I could see Ollama loading model to my GPU, i could see it see running but the output is empty. i doubt that the cause may come from nodes.py where answering logic is handled.

GiovanniPasq

Hi, I’m trying it and indeed I’m also getting an empty message. Let me analyze the problem and I’ll try to give you the solution as soon as possible

GiovanniPasq

Hi, I’ve identified where the issue comes from. It originates in the node responsible for aggregating the responses: ``` def aggregate_responses(state: State, llm): if not state.get("agent_answers"): return {"messages": [AIMessage(content="No answers were generated.")]} aggregation_prompt = get_aggregation_prompt(state["originalQuery"], state["agent_answers"]) synthesis_response = llm.invoke([SystemMessage(content=aggregation_prompt)]) return {"messages": [AIMessage(content=synthesis_response.content)]} ``` The problem is that all the messages to be aggregated are placed into a single system prompt. While Qwen3 is able to understand that the questions and answers are embedded inside that system prompt, Llama 3.2 is not. As a result, it processes the input as a system instruction and expects to receive the questions and answers separately, even though they are already embedded inside the system prompt. Try making the following changes and let me know how it goes. Inside the node.py script, replace the aggregate_responses function with this one: ``` def aggregate_responses(state: State, llm): if not state.get("agent_answers"): return {"messages": [AIMessage(content="No answers were generated.")]} aggregation_prompt = get_aggregation_prompt() sorted_answers = sorted(state["agent_answers"], key=lambda x: x["index"]) formatted_answers = "" for i, ans in enumerate(sorted_answers, start=1): formatted_answers += ( f"\nAnswer {i}:\n" f"{ans['answer']}\n" ) user_message = HumanMessage(content=f""" Original user question: {state["originalQuery"]} Retrieved answers: {formatted_answers} """) synthesis_response = llm.invoke([SystemMessage(content=aggregation_prompt)] + [user_message]) return {"messages": [AIMessage(content=synthesis_response.content)]} ``` Inside the prompts.py script, update the get_aggregation_prompt function with this version: ``` def get_aggregation_prompt() -> str: return f""" You are merging multiple retrieved answers into a final response. Rules: - Use ONLY the content provided in the retrieved answers. - Do NOT add new information, explanations, or assumptions. - Do NOT rephrase or paraphrase unless combining overlapping answers is required. Aggregation instructions: 1. If the answers cover different parts of the question: - Combine them into a single coherent response. - Preserve ALL details exactly as written. 2. If multiple answers contain overlapping or duplicate information: - Merge them carefully without removing details. 3. If an answer is irrelevant or empty: - Ignore it completely. Sources and citations: 4. Include source references ONLY if they already exist in the answers. 5. Do NOT invent, modify, or add new sources. 6. Place all source references ONLY at the end of the final answer. 7. Deduplicate sources if repeated. Failure handling: 8. If no usable answers are present: - Respond exactly with: "Sorry, I could not find any information to answer your question." Output: - Return ONLY the final answer. - Do NOT mention sub-questions. - Do NOT describe your reasoning. """ ``` Unfortunately, when using very small models (up to 8B), the system prompt has a significant weight, and each model requires specific optimizations. This issue does not occur when using larger models. This also happens with the nodes that generates the summary or rewrites the query. I will fix the other nodes as soon as possible, but the logic is the same as the one I already showed you.

GiovanniPasq

Hi, I’ve fixed the issue with Llama 3.2. I’ll push the updated code by the end of the day. Also, based on some tests, qwen3:4b-instruct-2507-q4_K_M performs better than Llama 3.2 (3B), so you may want to consider using it :)

GiovanniPasq

The issue has been fixed in the following commit: https://github.com/GiovanniPasq/agentic-rag-for-dummies/commit/02ad39af9be132909700440ceadd2c5b9a521b41 If you run into any other problems, feel free to open a new issue or reopen this one. 🙂 As mentioned earlier, when working with very small models, the system prompt becomes one of the most important — and most customization-dependent — components for achieving good performance.