Production RAG chatbot: retrieval, context limits, and evaluation
A production chatbot for author Ramazan Demir, answering in Turkish from his published writing. I built the retrieval context, input checks, reviewed corpus-update workflow, and regression benchmark.
Project at a glance
- Purpose
- Turkish answers from Ramazan Demir’s writing, with sources and an intended refusal when evidence is insufficient.
- My contribution
- Retrieval context, input validation, reviewed corpus updates, and FastAPI / Cloud Run.
- Status and checks
- Production, with a regression benchmark. Measured results are not published.
- Evidence
- Try the live chatbot (Turkish); implementation excerpts below.
You can try it here: Yapay Zeka ile Soru Cevap. Ask it something he has written about. The answer should sound like him, in Turkish, and it should stop when the sources do not cover the question.
I did not get that by picking a clever model. I got it by deciding what the model is allowed to see, and what it is allowed to do when that is not enough.
01The packet is the product.
The articles start as DOCX and PDF in a Drive folder. A scheduled job converts them, chunks them, and writes a diff of what is already in the system and what is not. I read that diff. Upload is a separate step.
Chunking is paragraph-aware. I do not want a quotation in one chunk and the sentence that explains it in the next, if they still fit together.
def _should_start_new_chunk(current, next_paragraph, config):
current_words = sum(p.word_count for p in current)
proposed_words = current_words + next_paragraph.word_count
if proposed_words > config.max_words:
return True
if current_words >= config.target_words and current_words >= config.min_words:
return True
return False
Each chunk keeps the title, year, category, and a content hash. The year matters. Word will happily stamp every rebuilt file with the day I resaved it, so a sidecar date map overrides the embedded date. Otherwise the bot cites 2013 for a piece written in 2021.
At query time I do not hand the model the search dump. A hit brings one neighbour on each side, then I cap the packet.
top_k = 8
neighbor_window = 1
max_context_chunks = 14
max_chunks_per_source = 4
max_sources = 4
context_chars_per_chunk = 1400
This implementation caps each source at four chunks to limit how much of any one article enters the context. These budgets are project choices, rather than universal retrieval limits.
02Some questions never reach the model.
A tap on the phone, a string of the same letter, a “teşekkürler” after the answer. None of those deserve a vector search. The guard is deterministic. No API call.
MIN_QUERY_CHARS = 3
REPETITIVE_CHAR_RATIO = 0.80
def validate_query_quality(query: str) -> str | None:
stripped = query.strip()
if len(stripped) < MIN_QUERY_CHARS:
return "query_too_short"
has_letter = any(ch.isalpha() for token in stripped.split() for ch in token)
if not has_letter:
return "query_no_meaningful_content"
lowered = stripped.lower()
max_count = max(lowered.count(ch) for ch in set(lowered))
if max_count / len(lowered) >= REPETITIVE_CHAR_RATIO:
return "query_repetitive"
return None
Thank-you messages short-circuit too. One sentence back, no headings, no fake sources. The expensive path is for an actual question.
03The instruction is a constraint, not a prompt trick.
Synthesis runs on a hosted model. I kept the weights there on purpose. The corpus is the asset. Retrieval lets the application supply explicit source context without training a model to imitate the author. The model sees the packet, and a short list of rules.
INSUFFICIENT_EVIDENCE_TEXT = "Bu konuda elimde yeterli delil yok."
# Answer only in Turkish, first person, and only from the packet.
# If the retrieved text has Arabic, quote it, then explain. Never invent Arabic.
# If a source states someone else's view, say it is not mine.
# If retrieval is thin, say so and stop. Do not guess.
# Do not fabricate source links, timestamps, or source cards.
The fallback line is the one I care about. A fluent miss is worse than “I don’t have enough here.” You can feel that on the live page if you ask outside the writing. It should refuse, not improvise.
04Reviewed corpus updates and regression checks.
Chunking and conversion run on a schedule. The vector store upload does not. First run is a dry proposal: what would be added, what would be deleted, no API call. I look at it. Execute is the second run.
The benchmark is the same habit. A file of questions, the same budget as production, and execute=False unless I ask for the model. The report lands on disk. If I change the chunker or the instruction, I want a diff against questions I already know, including the ones that used to fail.
@dataclass(frozen=True)
class HybridQueryBenchmarkRequest:
questions: str | Path
full_chunk_manifest: str | Path
chunk_files_dir: str | Path
output_dir: str | Path
top_k: int = 8
neighbor_window: int = 1
max_context_chunks: int = 14
execute: bool = False
The service in front of this is ordinary. FastAPI on Cloud Run, Google ID tokens, a usage limit. An answer I cannot trace to a caller is not one I want up.
If you build agents, this is the part I would start with. Not the model. The packet, the guard, and a refusal you can point at.