/ ENGINEERING SYSTEMS
Back to Field Notes

FIELD NOTE / PRODUCTION ACCOUNT

What Actually Happens When You Fine-Tune an LLM on 12,000 Domain Samples: A Production Account

15 min readUpdated 2026-09-30

A granular account of QLoRA on Llama 3 8B: the failed runs, the measured gains, the serving surprises, and why fine-tuning was not the system bottleneck we thought it was.

Public-safe engineering note · no client or confidential implementation details

We did not start with a desire to own a model. We started with a production number that would not move. Our agentic audit intelligence platform was using GPT-4o through Azure OpenAI for accounting-document classification and evidence extraction. After few-shot examples, chain-of-thought prompting, explicit output contracts, and several rounds of prompt review, the task F1 had settled at 0.74.

That score was not a disaster. It was also not good enough for a workflow where the wrong document label can route evidence to the wrong workpaper and where an omitted field becomes somebody else’s manual investigation. We had two options: accept the plateau, or adapt a smaller open model that we could serve ourselves. The first framing was “can fine-tuning make the model smarter?” The framing that survived was narrower: “can fine-tuning make this model more consistent on the format and domain vocabulary our evaluation set measures?”

01
HYPOTHESISWe were not transferring intelligence. We were specializing behavior.

That distinction changed the proposal. We were not expecting an 8B model to become a smaller GPT-4o. We were expecting it to stop improvising around accounting labels, preserve fields in a known schema, and recognize recurring language in financial footnotes. Fine-tuning was a behavior intervention, not a general capability upgrade.

  1. QLoRA on 12,000 domain samples will produce better task F1 than few-shot prompting of a larger model on accounting-document classification.
  2. The cost-per-query reduction from a self-hosted 8B model versus an API-hosted larger model will offset fine-tuning infrastructure within 90 days.
  3. Catastrophic forgetting will not significantly degrade general reasoning at 4-bit NF4 precision.
  4. vLLM with continuous batching will achieve acceptable P99 latency at production query volume.

What turned out to be wrong: hypothesis three was only partly right, and hypothesis four was right only at average queue depth. Hypothesis one was true on the measured task, but its wording hid how uneven the gain was. Hypothesis two was directionally right, though the 90-day payback depended on traffic actually arriving in a stable, batchable shape.

DIAGRAM 1 / FINE-TUNING DECISION TREE
What failure mode?Format / vocabularygap?Reasoning capabilitygap?Retrieval bottleneck,not generation?specializecapabilityevidence5,000+ high-qualitylabeled samples?Cost or quality isthe constraint?Can a new base modelclose the gap?Fine-tuneSpecialize behavior.ContractsFew-shot + schema.Fix retrievalGeneration is not bottleneck.Accept gapDo not train around it.New baseChange capability.yesnoqualitycostyes
010.74 task F1 after prompt saturation
0212,000 domain samples available
0338% modeled unit-cost reduction
Given this information, would you make the same decision?

The decision was made with evidence, but not with complete evidence. The evaluation set suggested that format and terminology were holding us back. It did not tell us that production requests would over-index on document types where retrieval and OCR quality mattered more than the generator. A good outcome does not retroactively make the information state good.

Share a thought

Comments appear immediately. Email is optional and never shown.

Markdown & code fences supported
AUTH VIA: PUBLIC FORM

No comments yet. Be the first to share a thought.

01 / SYSTEM VIEWInspect the architecture Follow a request through retrieval, validation, and audit.02 / EVIDENCEReview the case studies See the engineering decisions and measured outcomes.03 / OPEN CHANNELStart a conversation Talk through a production GenAI problem.