Fine-tuning Llama 3.2 1B for grounded HDN drafts on Amazon Bedrock
Fine-tuned Llama 3.2 1B Instruct locally on customer-approved HDN packets, imported the weights to Bedrock, and rejected any draft whose names or amounts were not in the packet. Support questions outside the product docs were handed off.
Project at a glance
- Purpose
- Readable mortgage-application drafts grounded in customer-approved HDN fields.
- My contribution
- Packet parsing, local Llama 3.2 1B Instruct fine-tuning, Bedrock import, and validation of draft names and amounts.
- Status and checks
- The workflow described checks drafts against their source packet. Production deployment and measured quality are not independently confirmed here.
- Evidence
- Reduced packet, training-pair, and inference excerpts below; no customer data or evaluation scores published.
Datakeeper is the wallet. The customer pulls source data onto their phone and chooses what to share. In 2024 it became a certified HDN brondata supplier: IBL income, BRP, Mijn Pensioenoverzicht, and identification, later the other bronberichten. The mortgage vault keeps that share for ninety days. Only the fields the application needs go out. That rule is the whole design of the draft. The model does not get to know more than the customer sent, and it does not underwrite anything.
01Support, on the site that is already live.
The site loads Crisp. The bot in that widget answers from the product: what a data request is, which source a check uses, what the customer still has to do in the app. It is not the draft model. A support question that is not in the product docs gets a handoff, not a guess. Same rule I use everywhere else. If the packet cannot support the sentence, the sentence does not ship.
02HDN arrives as messages. The model sees fields.
HDN bronberichten are XML. The ones that matter here are LoondienstInkomstenBericht, BasisregistratiePersonenBericht, PensioenOverzichtBericht, and DigitaleIdentificatieBericht, requested through a BronAanvraagBericht. I do not hand that XML to the model. A parser lifts the fields the draft is allowed to use and drops the rest. BSN stays out unless a contract field actually requires it. Most of the time it does not.
def to_packet(messages: dict) -> dict:
income = messages["LoondienstInkomstenBericht"]
person = messages["BasisregistratiePersonenBericht"]
return {
"applicant": {
"name": person["naam"],
"city": person["woonplaats"],
},
"income": {
"employer": income["werkgever"],
"annual_gross": income["toetsinkomen"],
},
"purpose": "draft-mortgage-application",
}
The keys above are the packet the draft may quote. The live parser maps HDN taxonomy paths into that shape. It is not the schema, and it is not the whole message. The packet is what the app can show the customer as “this is what the draft was built from.” If a field is missing, the draft says it is missing. It does not interpolate a salary.
03The weights were trained where the data already was.
IBL, BRP, and pension-register data does not go to a public fine-tune API so I can save a GPU. The base is Llama 3.2 1B Instruct. It is on the Bedrock Custom Model Import list, it is small enough to fine-tune on one GPU, and a filled-template draft does not need a larger model. I fine-tuned it locally on pairs of packet and draft. Import will not take an arbitrary network trained from scratch, and it will not take a small model whose architecture is not on the list. The export is Hugging Face safetensors plus the tokenizer files. Those weights go to S3 in eu-central-1, then into an imported model. Inference uses that imported model ARN, not a foundation-model id.
The public walkthrough of this size of job is Muhammad Imran Zaman’s Fine-Tuning 1B LLaMA 3.2: LoRA on a 4-bit Llama 3.2 1B, with Unsloth, on one GPU. That is the training shape. It is not this draft. His set is counseling conversations, the checkpoint he loads is the quantized base rather than Instruct, and the post stops at a local save. The packet, the validator, and the Bedrock import are the parts that are specific here.
The training target is a fixed ontwerp. The packet’s figures go in the slots. The sentences around them are fixed, including the line that this is a draft and not an offer. The loss is on that text. The label is produced by filling the template. The model is trained to emit that text from the packet, not to calculate a mortgage. A template would copy the figures exactly. The model is there for the wording, including the missing-field lines. The check in the next section is what stops a figure that was not in the packet.
row = {
"input": packet,
"output": (
"Ontwerp-aanvraag. "
"Aanvrager: {name}, {city}. "
"Toetsinkomen loondienst: {annual_gross}. "
"Dit is een ontwerp, geen offerte."
).format(**flat(packet)),
}
Inference is not that machine. The app is on AWS, the same stack as the rest of Datakeeper, on Fargate. The request path is an API call with a timeout, a log line, and a way to turn it off. No GPU sitting next to the services for a document that is a draft. Bedrock’s completion body for an imported model uses max_gen_len. Imports after November 2025 also accept the OpenAI completion key, max_tokens. The call below is the Bedrock shape.
response = bedrock.invoke_model(
modelId=DRAFT_MODEL_ARN,
body=json.dumps({
"prompt": render(packet),
"max_gen_len": 600,
"temperature": 0,
}),
)
Temperature stays at zero. This is not a place for variety. Two runs on the same packet should read the same.
04Then I check the draft against the packet.
The model is not the last step. After it returns, I pull every amount and every name back out and compare them to the packet. A figure that is not in the packet fails the draft. That check is the control. The model can emit a number the template would not have. The validator cannot. The app shows the customer an ontwerp, with the sources named, not a contract and not an offer. A person still accepts the application. The model’s job is to put the shared brondata into a readable draft so the customer can see what will be sent.