FIELD NOTE / PRODUCTION ACCOUNT
What Actually Happens When You Fine-Tune an LLM on 12,000 Domain Samples: A Production Account
A granular account of QLoRA on Llama 3 8B: the failed runs, the measured gains, the serving surprises, and why fine-tuning was not the system bottleneck we thought it was.
We did not start with a desire to own a model. We started with a production number that would not move. Our agentic audit intelligence platform was using GPT-4o through Azure OpenAI for accounting-document classification and evidence extraction. After few-shot examples, chain-of-thought prompting, explicit output contracts, and several rounds of prompt review, the task F1 had settled at 0.74.
That score was not a disaster. It was also not good enough for a workflow where the wrong document label can route evidence to the wrong workpaper and where an omitted field becomes somebody else’s manual investigation. We had two options: accept the plateau, or adapt a smaller open model that we could serve ourselves. The first framing was “can fine-tuning make the model smarter?” The framing that survived was narrower: “can fine-tuning make this model more consistent on the format and domain vocabulary our evaluation set measures?”
That distinction changed the proposal. We were not expecting an 8B model to become a smaller GPT-4o. We were expecting it to stop improvising around accounting labels, preserve fields in a known schema, and recognize recurring language in financial footnotes. Fine-tuning was a behavior intervention, not a general capability upgrade.
- QLoRA on 12,000 domain samples will produce better task F1 than few-shot prompting of a larger model on accounting-document classification.
- The cost-per-query reduction from a self-hosted 8B model versus an API-hosted larger model will offset fine-tuning infrastructure within 90 days.
- Catastrophic forgetting will not significantly degrade general reasoning at 4-bit NF4 precision.
- vLLM with continuous batching will achieve acceptable P99 latency at production query volume.
What turned out to be wrong: hypothesis three was only partly right, and hypothesis four was right only at average queue depth. Hypothesis one was true on the measured task, but its wording hid how uneven the gain was. Hypothesis two was directionally right, though the 90-day payback depended on traffic actually arriving in a stable, batchable shape.
The decision was made with evidence, but not with complete evidence. The evaluation set suggested that format and terminology were holding us back. It did not tell us that production requests would over-index on document types where retrieval and OCR quality mattered more than the generator. A good outcome does not retroactively make the information state good.
Our 12,000 samples came from three sources: de-identified historical review examples, synthetic perturbations of approved accounting-document passages, and a smaller set of reviewer-authored edge cases. We removed client identifiers and engagement-specific references before any training artifact left the controlled environment. Each sample contained a document excerpt, the task instruction, a typed target label or extracted field set, and an evidence span. We kept the evidence span because a correct answer without a known supporting region was not enough for the workflow.
We split the data 80/10/10 by source family rather than by random row. That produced 9,600 training samples, 1,200 validation samples, and 1,200 internal test samples. A random split would have put near-duplicate footnote templates across all three sets and inflated confidence. The final score came from a separately curated set of 214 accounting-domain cases that none of the training-generation prompts could see.
We used 4-bit NF4 quantization, LoRA rank 16, alpha 32, dropout 0.05, and target modules on the attention Q, K, and V projections. Rank 16 was a compromise: enough capacity to encode label and phrasing behavior, small enough to keep the trainable fraction below one percent. Alpha twice the rank kept the adapter contribution visible without making the initial update so large that the run became hypersensitive to the learning rate. We began with a 1.6e-4 learning rate, effective batch size 16, cosine decay, and gradient clipping at 1.0.
Trainable params = layers × rank × Σ(input width + output width) for Q/K/V. Llama 3 8B: 32 × r × (8,192 + 5,120 + 5,120); memory is a heuristic for adapter weights, optimizer state, and batch activations.
Run 1 failed in a way that looked almost successful for the first few checkpoints. We increased the effective batch size to 32 to improve throughput and left the learning rate at 3e-4 because a generic configuration guide described that range as reasonable. At step 800 the validation loss spiked, gradient norms jumped above 18, and the model began emitting a repeated label token. The training process did not crash. It completed with a model that looked usable on easy examples and failed on the exact long-tail cases we cared about.
Run 2 changed a different variable. We reduced the learning rate to 8e-5 and kept the larger batch, expecting stability. It was stable, but the validation curve plateaued early. The adapter was underfitting: rank 8 and alpha 8 in that run gave us a conservative update with too little capacity for the nested extraction format. The documented learning-rate range had not been the useful boundary; the interaction between learning rate, batch, rank, and alpha was.
Run 3 returned to rank 16 and alpha 32, reduced the effective batch to 16, set learning rate to 1.6e-4, and enabled gradient accumulation instead of increasing the per-device batch. We also added a per-category validation gate rather than selecting the checkpoint with the lowest aggregate loss. The change fixed the divergence without hiding the categories that remained weak.
| RUN ID | CONFIG CHANGE | VAL F1 / CKPT 1 | VAL F1 / CKPT 2 | NOTES |
|---|---|---|---|---|
| R-01 | rank 16 · alpha 32 · lr 3e-4 · batch 32 | 0.68 | 0.52 | gradient norm spike at step 800; repeated label token |
| R-02 | rank 8 · alpha 8 · lr 8e-5 · batch 32 | 0.71 | 0.73 | stable but underfit; plateaued early |
| R-03 | rank 16 · alpha 32 · lr 1.6e-4 · batch 16 | 0.78 | 0.86 | first stable run; footnote slice still weak |
| R-04 | rank 16 · alpha 32 · lr 1.6e-4 · batch 16 + clip | 0.80 | 0.88 | best aggregate checkpoint; cross-border regression |
| R-05 | rank 32 · alpha 64 · lr 1.6e-4 · batch 16 | 0.79 | 0.87 | no meaningful gain; memory overhead doubled |
| R-06 | rank 16 · alpha 32 · lr 1.2e-4 · batch 16 | 0.79 | 0.87 | slower convergence; no production lift |
| R-07 | rank 16 · alpha 32 · lr 1.6e-4 · selected modules | 0.81 | 0.91 | final candidate; category gate passed |
| R-08 | final adapter + long-context prompt | 0.80 | 0.89 | prompt interaction reduced extraction on long cases |
Our 214-case evaluation measured classification, structured entity extraction, evidence-span retention, field completeness, anomaly flagging, and abstention on ambiguous documents. Two reviewers independently labeled the harder categories. Agreement was 0.91 on classification, 0.86 on extraction, and 0.71 on open-ended anomaly flagging. We treated the disagreement as a property of the task, not as noise to average away. The class imbalance was known: classification and extraction represented most production volume, while anomaly cases were fewer and more subjective. We reported macro F1, per-category F1, and a volume-weighted operational score so a strong extraction result could not erase a weak anomaly slice.
The final task F1 improved from 0.74 to 0.91 on the held-out evaluation set. That was a meaningful result, but it was not uniform. Structured entity extraction from financial footnotes improved from 0.61 to 0.94 because the adapter learned recurring field boundaries, abbreviations, and output order. Document classification improved from 0.76 to 0.93. Policy-reference extraction moved from 0.78 to 0.90.
Open-ended anomaly flagging improved from 0.69 to 0.76. That was useful, but modest. The task required deciding whether a pattern was unusual, not merely mapping language to a known label. The fine-tuned model became more consistent in how it expressed an anomaly; it did not gain a reliable causal theory of why an accounting pattern was unusual. We should have expected that.
Cross-border transaction disclosures regressed from 0.81 to 0.74. The training distribution contained too few jurisdictional variants and too many domestic disclosure templates. The adapter learned the dominant shape. This was not evidence that fine-tuning “broke” the model in a mysterious way. It was evidence that specialization has a boundary, and the boundary was visible only when the evaluation set stopped looking like the training set.
Hover is represented by the measurement labels; the frontier is the set of quality/cost choices not dominated by another point.
Hypothesis three was partially wrong. General reasoning on the small accounting-adjacent probe held within two points, but out-of-domain tasks degraded: general science explanation fell from 0.82 to 0.78 on our probe, and a small instruction-following slice fell from 0.88 to 0.83. The adapter did not erase general reasoning, but it changed the model’s prior toward the domain. That mattered for deployment. We could not route every request to the adapted model just because its accounting score was better.
Hypothesis four was correct at average queue depth. With vLLM continuous batching, average P99 latency was 238 milliseconds at the measured production rate. Under burst traffic, queue depth grew faster than our initial HPA policy expected, and P99 crossed 900 milliseconds before new replicas became ready. The model was not slow in isolation; the service was slow at the point where arrival rate, batching, GPU warm-up, and autoscaling met.
We served the model through vLLM with continuous batching, a capped maximum number of sequences, explicit swap-space limits, and a separate health probe that exercised a real structured output rather than checking only that the process was alive. The settings that mattered were maximum sequence length, batch-token limits, GPU memory utilization, and the readiness threshold used by the autoscaler. We spent less time tuning tensor parallelism because the 8B model fit our single-GPU shape.
One setting caused a silent correctness bug for three days: an output-token cap copied from an unrelated summarization service. It was low enough that long evidence fields were truncated, but high enough for the HTTP response to remain valid. The schema validator accepted the object because the required keys existed. The fix was not “increase the cap” alone. We added a completeness check for evidence spans, recorded finish reasons, and made truncation a hard failure for the extraction route.
The 38% per-query cost reduction came from a deliberately boring calculation. The GPT-4o baseline averaged $0.013 per query at the measured prompt and completion lengths. The self-hosted model cost $0.0081 per query when GPU time, tokenization, batching overhead, and API gateway overhead were allocated across observed volume. That is a $0.0049 reduction, or 37.7%, rounded to 38%. The estimate excludes the one-time dataset and training work; it includes the steady-state serving footprint.
At 1.05 million queries per month, the arithmetic was $13,650 for the API baseline versus $8,400 for the fine-tuned serving path. The $5,250 monthly difference was the reason the adaptation program could pay back its recurring operating cost. It did not make the system cheap in an absolute sense. It made the quality-preserving option less expensive than repeatedly paying for a larger general model.
We integrated the adapter into CI/CD as a versioned model artifact. Every candidate carried a dataset snapshot, base-model digest, adapter commit, tokenizer version, serving image, prompt contract, and evaluation receipt. The gate checked aggregate F1, per-category floors, evidence-span completeness, schema validity, P99 at a fixed concurrency, and a small out-of-domain probe. A candidate could improve the headline score and still be blocked.
In the first 30 days, the gate blocked a regression in cross-border disclosure routing. The new adapter improved the aggregate score by 0.02, but the category floor fell by 0.07 and the evidence-span retention rate fell below the release threshold. Without the per-category gate, the change would have reached client-facing tooling and the failure would probably have appeared as a reviewer complaint rather than a deployment failure. The constructed replay above is not a claim that a particular client incident occurred; it is the failure class the gate prevented.
Model versioning also changed rollback language. We did not say “roll back the model” as though one file controlled the behavior. We rolled back a bundle: model and adapter, tokenizer, prompt contract, vLLM image, routing rule, and evaluation receipt. When a domain update required a capability the adapted model did not yet cover, we temporarily routed that slice back to GPT-4o. We knew it was safe because the fallback had a separate regression receipt, the request type was explicitly allowlisted, and the trace recorded which model served each result.
Promote the adapter only for stable accounting slices; keep GPT-4o as a controlled fallback for uncovered document types. Unknown at decision time: the exact production mix of cross-border disclosures and the burst shape of monthly close.
The fallback was not a failure of self-hosting. It was the correct operational response to a capability boundary. The unsafe version would have been routing everything to the fine-tuned model because the average score looked better.
We would change the sample-construction decision first. We grouped samples by source family to avoid leakage, which was correct, but we allowed domestic footnote templates to dominate the training mix because they were cleaner and easier to label. That introduced a distribution skew. The model learned the stable majority and became less tolerant of jurisdictional variation. In retrospect, a smaller but deliberately stratified cross-border slice would have been more valuable than another thousand easy extraction examples.
We would also add operational disclosure packets to the evaluation map before training. Those packets were present in production, but absent from the test set. They combined scanned pages, abbreviated entity names, and reviewer annotations in a layout that was neither a clean footnote nor a normal classification document. The first post-deployment surprise was not a dramatic hallucination. It was a confident, well-formed answer over an input shape we had never measured.
The hypothesis that was right for the wrong reason was the cost hypothesis. The unit-cost reduction was real, but the savings did not come from the adapter alone. They came from choosing a model that fit a serving shape, batching requests, routing stable slices away from the API, and refusing to treat every query as a long-context general reasoning problem. If traffic had remained bursty and low-volume, the same model would have looked financially worse.
The uncomfortable conclusion is that fine-tuning was not the biggest bottleneck in the system. Three months later, we improved retrieval quality and recovered more value than the adapter work had produced. Better OCR normalization, document-aware chunking, jurisdiction metadata, and evidence-span ranking improved the end-to-end workflow because many “generation” errors were actually missing or poorly ranked evidence. The 0.91 model could only specialize over the context it received.
Fine-tuning solved behavior specialization on a narrow distribution. It did not solve retrieval misses, missing evidence, or open-ended reasoning. Unknown at decision time: retrieval quality was the dominant system bottleneck for a large share of production requests.
That is why the decision tree matters. “Should we fine-tune?” is a weak question because it assumes a single intervention. The useful question is which failure mode is limiting the workflow, what evidence would distinguish it from neighboring failures, and whether the team can operate the resulting model with a rollback target.
Usually. It specializes format, labels, and refusal patterns.
Schema contracts, constrained decoding, validators, few-shot examples.
Slice-level lift, held-out cases, cost and rollback evidence.
Whether the inconsistency is actually caused by retrieval or parser variance.
Sometimes, if the vocabulary appears repeatedly and the task is narrow.
Glossaries, query expansion, domain-aware retrieval, better chunking.
Slice-level lift, held-out cases, cost and rollback evidence.
Whether the model needs knowledge or only a better evidence path.
Rarely. Adapter training does not create a missing reasoning capability.
A different base model, decomposition, tools, verification, human review.
Slice-level lift, held-out cases, cost and rollback evidence.
Whether the task can be made smaller and more deterministic.
Yes, when the smaller model reaches the required quality and volume is real.
Caching, batching, prompt reduction, routing, provider negotiation.
Slice-level lift, held-out cases, cost and rollback evidence.
Serving and operational costs that replace the API bill.
Only for a repeated, narrow behavior; not as a general truthfulness patch.
Retrieval quality, citations, abstention contracts, claim-level gates.
Slice-level lift, held-out cases, cost and rollback evidence.
Whether the failure is unsupported context rather than model behavior.
No. Fine-tuning does not expand context or repair ingestion.
Parsing, hierarchical retrieval, summaries, long-context model, map/reduce.
Slice-level lift, held-out cases, cost and rollback evidence.
Where the information was lost before generation.
Often. Style is a behavior distribution with observable examples.
System prompts, style examples, post-processing, templates.
Slice-level lift, held-out cases, cost and rollback evidence.
Whether style matters enough to justify lifecycle ownership.
No. It can reduce repetitive work but cannot define accountability.
Confidence routing, review queues, audit logs, approval boundaries.
Slice-level lift, held-out cases, cost and rollback evidence.
Who owns the residual error when the score is not enough.
Fine-tuning is a tool with a specific problem it solves: behavior specialization on a narrow distribution. Teams that apply it to reasoning gaps, knowledge gaps, or retrieval failures will often get worse results at higher cost than teams that apply it to the right problem. The adapter can make a model more consistent. It cannot make absent evidence appear.
Share a thought
Comments appear immediately. Email is optional and never shown.
No comments yet. Be the first to share a thought.