FROM THE DESK OFMERT’S MIND
FIELD NOTE / 002APPLIED AI
A QUESTION BOT, IN PRODUCTION

Production RAG chatbot: retrieval, context limits, and evaluation

A production chatbot for author Ramazan Demir, answering in Turkish from his published writing. I built the retrieval context, input checks, reviewed corpus-update workflow, and regression benchmark.

By RAG / EVALUATION / CLOUD RUN7 MIN READ

Project at a glance

Purpose
Turkish answers from Ramazan Demir’s writing, with sources and an intended refusal when evidence is insufficient.
My contribution
Retrieval context, input validation, reviewed corpus updates, and FastAPI / Cloud Run.
Status and checks
Production, with a regression benchmark. Measured results are not published.
Evidence
Try the live chatbot (Turkish); implementation excerpts below.

You can try it here: Yapay Zeka ile Soru Cevap. Ask it something he has written about. The answer should sound like him, in Turkish, and it should stop when the sources do not cover the question.

I did not get that by picking a clever model. I got it by deciding what the model is allowed to see, and what it is allowed to do when that is not enough.

01The packet is the product.

The articles start as DOCX and PDF in a Drive folder. A scheduled job converts them, chunks them, and writes a diff of what is already in the system and what is not. I read that diff. Upload is a separate step.

Chunking is paragraph-aware. I do not want a quotation in one chunk and the sentence that explains it in the next, if they still fit together.

ARTICLE CHUNKERWHEN TO BREAK
def _should_start_new_chunk(current, next_paragraph, config):
    current_words = sum(p.word_count for p in current)
    proposed_words = current_words + next_paragraph.word_count

    if proposed_words > config.max_words:
        return True
    if current_words >= config.target_words and current_words >= config.min_words:
        return True
    return False

Each chunk keeps the title, year, category, and a content hash. The year matters. Word will happily stamp every rebuilt file with the day I resaved it, so a sidecar date map overrides the embedded date. Otherwise the bot cites 2013 for a piece written in 2021.

At query time I do not hand the model the search dump. A hit brings one neighbour on each side, then I cap the packet.

CONTEXT BUDGETPER QUESTION
top_k                   = 8
neighbor_window         = 1
max_context_chunks      = 14
max_chunks_per_source   = 4
max_sources             = 4
context_chars_per_chunk = 1400

This implementation caps each source at four chunks to limit how much of any one article enters the context. These budgets are project choices, rather than universal retrieval limits.

02Some questions never reach the model.

A tap on the phone, a string of the same letter, a “teşekkürler” after the answer. None of those deserve a vector search. The guard is deterministic. No API call.

QUERY GUARDBEFORE RETRIEVAL
MIN_QUERY_CHARS = 3
REPETITIVE_CHAR_RATIO = 0.80

def validate_query_quality(query: str) -> str | None:
    stripped = query.strip()
    if len(stripped) < MIN_QUERY_CHARS:
        return "query_too_short"

    has_letter = any(ch.isalpha() for token in stripped.split() for ch in token)
    if not has_letter:
        return "query_no_meaningful_content"

    lowered = stripped.lower()
    max_count = max(lowered.count(ch) for ch in set(lowered))
    if max_count / len(lowered) >= REPETITIVE_CHAR_RATIO:
        return "query_repetitive"
    return None

Thank-you messages short-circuit too. One sentence back, no headings, no fake sources. The expensive path is for an actual question.

03The instruction is a constraint, not a prompt trick.

Synthesis runs on a hosted model. I kept the weights there on purpose. The corpus is the asset. Retrieval lets the application supply explicit source context without training a model to imitate the author. The model sees the packet, and a short list of rules.

SYSTEM INSTRUCTIONTHE PART THAT MATTERS
INSUFFICIENT_EVIDENCE_TEXT = "Bu konuda elimde yeterli delil yok."

# Answer only in Turkish, first person, and only from the packet.
# If the retrieved text has Arabic, quote it, then explain. Never invent Arabic.
# If a source states someone else's view, say it is not mine.
# If retrieval is thin, say so and stop. Do not guess.
# Do not fabricate source links, timestamps, or source cards.

The fallback line is the one I care about. A fluent miss is worse than “I don’t have enough here.” You can feel that on the live page if you ask outside the writing. It should refuse, not improvise.

04Reviewed corpus updates and regression checks.

Chunking and conversion run on a schedule. The vector store upload does not. First run is a dry proposal: what would be added, what would be deleted, no API call. I look at it. Execute is the second run.

The benchmark is the same habit. A file of questions, the same budget as production, and execute=False unless I ask for the model. The report lands on disk. If I change the chunker or the instruction, I want a diff against questions I already know, including the ones that used to fail.

BENCHMARK REQUESTDEFAULT IS NOT TO CALL
@dataclass(frozen=True)
class HybridQueryBenchmarkRequest:
    questions: str | Path
    full_chunk_manifest: str | Path
    chunk_files_dir: str | Path
    output_dir: str | Path
    top_k: int = 8
    neighbor_window: int = 1
    max_context_chunks: int = 14
    execute: bool = False

The service in front of this is ordinary. FastAPI on Cloud Run, Google ID tokens, a usage limit. An answer I cannot trace to a caller is not one I want up.

If you build agents, this is the part I would start with. Not the model. The packet, the guard, and a refusal you can point at.

END OF NOTE / 002RETURN TO THE INDEX ↑