THE PROBLEM
Why this system existed
High-volume domain tasks need strong quality and lower unit cost without sacrificing reproducibility or rollback.
Fine-tuning · MLOps · Cost Control
QLoRA fine-tuning, MLflow experiment tracking, vLLM serving, and staged promotion with quality checks at each gate.
THE PROBLEM
High-volume domain tasks need strong quality and lower unit cost without sacrificing reproducibility or rollback.
OUTCOME
Improved task F1 from 0.74 to 0.91 and reduced per-query inference cost by 38%.
REFERENCE ARCHITECTURE
DECISIONS
FAILURE CASE
Fine-tuning gains are task-specific; model promotion stayed gated by domain evaluation rather than benchmark performance alone.
SECURITY BOUNDARY
This case study exposes patterns, not employer architecture. It uses synthetic data, no client identifiers, no internal prompts, no proprietary datasets, and no production endpoints.