I fine-tuned a pretrained Qwen3.5-0.8B model with LoRA to assign one of 77 intent labels to a banking support message. On a held-out sample from the training dataset, exact-label accuracy rose from 35.7% to 85.3%. The selected adapter then scored 84.1% on the separate official test split. This is a small, measured post-training experiment, not a new language model trained from scratch.
The question
I wanted to learn fine-tuning through a task with a clear answer and a meaningful test. Banking77 supplies short customer messages and 77 intent labels, such as cardarrival and pendingtransfer. The model must return exactly one of those labels. This makes success easier to measure than judging whether generated prose sounds good.
I used Qwen3.5-0.8B as the pretrained base, Unsloth to run supervised fine-tuning on Kaggle, and a LoRA adapter rather than updating every model weight. The run trained 6.39 million parameters, about 0.74% of the model. The measured task is intent classification; these results do not measure general chat ability.
What the model saw
The dataset's official training split has 9,993 messages. I split it with seed 3407 into 8,993 training and 1,000 validation messages. The dataset's separate 3,076-message official test split was reserved for the final evaluation.
The training notebook checks that there are zero exact text matches between its training and validation rows and between its training rows and the official test rows. Its prompt lists the 77 label names from the official training split, but no official test message or answer is used to update weights or select a checkpoint. Validation loss was measured on 200 validation rows every 50 steps; the 300-row accuracy sample came from that same validation split, so it should not be described as an independent test.
The separate test notebook loads the saved 600-step adapter and performs inference. It has no training step. It verifies the source adapter and scores all 3,076 official test rows with the same prompt and greedy decoding used for validation. I looked at its test results only after choosing the 600-step adapter.
“Unseen” has a precise limit here: the official test messages were unseen during this fine-tuning and model selection. Banking77 is public, so I cannot prove that the original pretrained Qwen model never encountered the dataset. The exact-match check also cannot rule out close paraphrases across splits.
The experiment
The base model was prompted to return one label from the list. I measured it on the fixed 300-message validation sample before fine-tuning. Then I ran a 150-step pilot and a separate 600-step LoRA run under the same data split, seed, prompt, and decoding method. The longer run used an effective batch of four and saved the checkpoint with the lowest monitored validation loss. Its best checkpoint was at step 600.
Model/run Fixed validation accuracy Invalid labels Pretrained base, no fine-tuning 107/300 = 35.7% 29/300 = 9.7% 150-step LoRA pilot 207/300 = 69.0% 5/300 = 1.7% 600-step LoRA run 256/300 = 85.3% 1/300 = 0.3%
The 600-step run's best monitored validation loss was 0.0946. Its validation accuracy beat the shorter pilot by 16.3 percentage points on the same 300 examples. That is evidence that the 150-step pilot had not yet reached this run's best performance; it does not establish an optimal training length.
The held-out result
After selecting the adapter, I ran the official test evaluation once:
Official test measure Result Exact-label accuracy 2,587/3,076 = 84.1% Invalid labels 7/3,076 = 0.23% 95% Wilson interval for accuracy 82.8%–85.4%
The test score is 1.2 percentage points below the 300-row validation score. That gap alone does not suggest severe overfitting, though one split and one final test cannot rule it out. A few frequent errors confuse nearby intents: for example, cardarrival with carddeliveryestimate, and transfernotreceivedbyrecipient with balancenotupdatedafterbanktransfer. These are useful cases to inspect before changing the data or training recipe.
What I learned, and what comes next
The important result is not that a small model can emit a plausible label. It is that the training procedure produced a large improvement on a fixed validation sample, and the selected adapter retained similar accuracy on a separate official test. The experiment also exposed the boundary between validation, which guides development, and test, which estimates how the chosen system performs on held-out examples.
I would next audit the errors by intent and release a reproducible repo with the code, configuration, split seed, metrics, and prediction files. The Kaggle notebooks and outputs are currently private, so this post does not yet offer independently accessible run artifacts. The adapter and base weights do not need to be included to make the method and results reviewable.