Training a Small Model on My Own Home Network

I have been using LLMs for years without ever training one. So I spent a few weeks teaching a 0.6B model to answer questions about my own UniFi network, and learned that my measuring stick was broken, my control group was a lucky pair, and fine-tuning made hallucination worse instead of better.

Training a Small Model on My Home Network

(Or, me fighting a 0.6B model and my own measurement errors, in that order)

I have been using LLMs for a couple of years now - well, since ChatGPT was announced. I have not, until this summer, actually trained one.

Knowing fine-tuning exists and having done it are different things. So I went and did it, end to end: data curation, SFT, LoRA, preference tuning, RL, quantization. The whole ladder.

The immediate problem was picking something to train it on.

Most fine-tuning tutorials train on text. Summarize this, rewrite that, answer in the voice of a pirate. Which is fine, except that when you are done you have no idea if it worked. LLMs spit out text for a living - I wanted something gradeable.

So I trained mine on my own home network instead.

Why the network

I have a UDM Pro, a handful of APs and switches, a Prometheus stack on guiltyspark, and a NixOS fleet with Halo-themed hostnames because I am who I am. That is a pile of real, queryable data.

If I teach a small model to turn "what's on the guest wifi?" into a tool call, the grader can't* make stuff up. It has to return real data.

*yes, of course it could. The point is making it not.

Either the model named a real tool with valid arguments, or it did not. Either the call returned data from the live controller, or it threw an error. Ideally.

To be clear: the model is not the detector. Prometheus rules and unpoller already do detection, deterministically and testably. The model is just the language surface on top.

detection    Prometheus rules, Loki, unpoller, blackbox   (already built)
   |
   v
tools        16 typed verbs. Parameterized query catalog, NOT free PromQL.
   |
   v
SLM          llama.cpp + GBNF grammar -> output is always a valid call
   |
   v
policy       reads auto-run, writes only ever *proposed*, human confirms

The deliverable was never "a useful tool." My 27B already answers these questions better than anything I was going to train. The deliverable was the training.

The box

All of this runs on my old Gaming PC optimus - a Legion T7 with a 12700K and the RTX 3090 I moved over in August. (Yes, Optimus. Transformers are just as cool as Halo, right?)

Incidentally, I'm running Bazzite on this host, I've used quote a few immutable distros over the years, starting with CoreOS (Shoutout to Brian 'Redbeard').

The 3090 was already busy serving a 27B at Q4, which eats 17.9 of the 23.6 GB available. So before writing a line of training code I measured what would actually fit next to it:

Qwen3-0.6B bf16 peak VRAM with the 27B up
load + generate 1.19 GB fine
LoRA r=16, grad checkpointing 1.66 GB fits
full fine-tune backward >3 GB OOM

That table decided a lot. LoRA coexists with inference. A full fine-tune means stopping the 27B first. Desktop apps hold about 900 MB, so budget that too.

The training environment lives in its own Distrobox called netmon-train, deliberately separate from the llamacpp box. A torch upgrade should never be able to take the 27B down as collateral damage. I have been burned by shared Python environments enough times to stop pretending it is fine.

Rule zero: build the referee before the player

This is the part worth stealing.

Before any training, I built a second repo, netmon_evals, that exists only to grade. 208 questions, each with the tool call that correctly answers it. A tool contract. A scorer. A sha256 manifest so validate.py fails loudly if anyone edits the question set. (Anyone = future me.)

This, of course, is a recommended path - I didn't figure it out on my own. Learning from others is important!

Two rules, both written down early because I knew I would want to break them later:

  1. Never train on eval/. Check overlap before every single run.
  2. Never edit questions.jsonl in place. A changed set silently invalidates every score you have already recorded.

19 of the 208 are unanswerable or ambiguous on purpose. The scorer reports refusal recall and over-refusal separately, because a model that answers "which device did you mean?" to everything scores beautifully on the first one and is completely useless.

Scoring is on tool-call correctness, never answer text. Client counts around here change hourly. An answer-text eval would rot in a day.

Then I ran the floor and the ceiling:

Qwen3-0.6B (floor) Qwen3.8-27B (ceiling)
call accuracy 26.4% 81.7%
tool accuracy 34.1% 88.0%
refusal recall 36.8% 84.2%
over-refusals 21 6
208 questions in 47s 152s

About 55 points of headroom. That is the whole exercise: a number to move.

Worth noting, verify_gold.py executes all 208 gold calls against the live network, and all 208 pass. So a miss belongs to the model, not to my plumbing. That acceptance test took an afternoon and saved me from at least three wild-goose chases later.

Two things that make a tiny model viable

Constrained decoding, not prompting. GBNF grammar makes malformed output structurally impossible. You are not hoping for valid JSON; the sampler literally cannot emit anything else. Unconstrained, the 27B answers "is my internet up?" with four paragraphs about ping. Grammar on: {"tool":"get_health","args":{}} in 1.1 seconds. Zero unparseable outputs across 416 generations.

Tools return conclusions, not raw numbers. Never hand a model [143, 211, 98] and ask if that is unusual. Return this instead:

{"current": 211, "baseline_7d": 98, "delta_pct": 115, "verdict": "elevated"}

Every numeric comparison you move out of the weights is a hallucination you have deleted.

Small gotcha banked along the way: reasoning models and grammars do not mix. With thinking enabled, the grammar constrains the reasoning channel and content comes back empty. Set chat_template_kwargs: {"enable_thinking": false} and move on.

First training run

Hand-rolled SFT loop, no TRL - I wanted to see the masking logic with my own eyes. Qwen3-0.6B, full fine-tune, 688 examples, 3 epochs.

169 seconds. Val loss 1.5923 to 0.0524.

run call acc tool acc refusal recall
base 0.6B 22.6% 33.2% 31.6%
sft1 64.4% 74.0% 68.4%
27B ceiling 81.7% 88.0% 84.2%

Three quarters of the distance to a model 45 times its size, from 688 examples and under three minutes of compute. clarify went from 0.0% to 75.0%. Not bad for a 0.6B, imho.

One caveat, though. The base row there is re-scored through the same greedy decode path as the trained model, no grammar on either. The Phase 0 numbers used GBNF. Comparing across that line would have measured the grammar, not the training, and would have handed me a much nicer number than I earned.

The loss curve lies

Same data, same hyperparameters, one flag. --no-mask puts loss on the prompt tokens as well as the answer tokens.

masked unmasked
final val loss 0.0524 0.0561
call accuracy 64.4% 51.4%
refusal recall 68.4% 42.1%

A 7% difference in validation loss - the kind you would shrug at or never even see on a chart - is thirteen points of task accuracy. A quarter of the model's usefulness.

If I had been watching the loss curve I would have concluded both runs were fine. That is the entire argument for building the frozen eval before you write any training code. I got to learn it on run two instead of run forty, so I count that.

Composition beats volume

Two categories got worse after training. The data told me exactly why:

tool training examples what happened
list_wlans 5 7 misses, every one routed to get_network (100 examples)
list_port_forwards 4 misrouted to clarify and list_clients
get_sysinfo 4 routed to unsupported

My generator capped the dominant tools at 110 so propose_action and get_client would not swamp the set. Capping the head does nothing for the tail. list_wlans kept its five examples, and the model learned the higher-prior rule "named network thing goes to get_network" and cheerfully applied it to SSIDs.

688 examples bought 42 points. Nine missing examples cost two categories outright.

The ruler was broken

I eventually moved to Qwen3-4B with a merged LoRA at r32, which got to 86.5% on the frozen eval. That matches my 27B teacher at 6.75x fewer parameters, which felt great for about a day.

Then I went looking at the two categories still dragging, metrics at 53.8% and actions at 58.3%.

Here is what the shipped model produced for three action questions, and what the tool layer did with each:

"cut off the printer's internet"
  -> propose_action{verb: "block", target: "printer"}
  -> resolved: null, no error, status: ok

"unblock the kids iPad"
  -> propose_action{verb: "unblock", target: "kids iPad"}
  -> resolved: null, no error, status: ok

"bounce the switch that truenas is plugged into"
  -> propose_action{verb: "reconnect", target: "truenas"}
  -> resolved: truenas, no error, status: ok

The first two name nothing at all. The real printer is prusa-mk3-5, and my resolvers match by substring, so "printer" matches nothing. But propose_action attached resolved: null and returned normally. My dispatcher treats "no error key" as success. So the harness recorded a confident, well-formed, cleanly-executed instruction to act on a device that does not exist, and counted it toward a 98.1% execution rate. (Lovely.)

The third one is the cleanest example of this. truenas resolves. It is a real client. The call is well-formed and executes cleanly. It is also completely wrong, because the user asked about the switch the NAS is plugged into, and this proposal would kick the NAS off the wifi. Nothing anywhere in my stack could tell the difference.

Both paths return errors now. The check that made it safe to tighten them: all 189 executable gold calls still pass. So what the stricter layer rejects was genuinely broken, not merely unusual.

Re-measuring the same weights with the fixed denominator moved execution from 98.1% to 97.1%. Every previously published number for this project was flattering by exactly two phantom calls.

When a metric and a tool disagree about what "worked" means, fix the definition before you touch the model.

Confabulation

With that fixed, I built a second evaluation that needs no answer key at all. 74 questions the frozen eval does not contain, built from live inventory, scored only on things that are objectively checkable. Does a propose_action target resolve to a real device? Is a metric_query id in the catalog? Does a clarify assert something about the inventory that is false?

None of that can be improved by fitting the referee, which was the point. I had read those 208 eval questions far too many times.

Then I ran it across the whole family.

model confabulation unparseable invented a metric id
Qwen3-0.6B, untuned 9.5% 11 / 74 -
Qwen3-1.7B, untuned 5.4% 7 / 74 -
Qwen3-4B, untuned 5.4% 7 / 74 0
Qwen3-4B, fine-tuned 9.7% 0 5
Qwen3.8-27B teacher 0.0% 0 0

Read rows three and four together.

The untuned base never once invented a catalog id. Asked whether the radio link to the kitchen AP had been deteriorating, it produced no parseable call at all for seven of 74 questions, and otherwise stayed inside the catalog. My fine-tune removed every single parse failure, and produced five confident references to wlan_signal and wlan_link_quality, neither of which exists anywhere except in the model's own output. Both names are extremely plausible. Neither is in the catalog printed in its own prompt.

Fine-tuning did not reduce hallucination. It made hallucination well-formed.

And it is the same fact. My training set taught fluency in the output format and never once penalized invention, because no example in it ever needed to. So the model learned that a syntactically valid call is always available - true - and that its contents are always allowed. Which is not.

That inverts the intuition that the less-trained model is the riskier one. The base model's failures are visible: malformed output, an obvious error, a thing you notice. The fine-tune's failures are invisible: well-formed, executable, and wrong.

The 27B row matters too. It scores 0.0%, so this is not "every model does this." A big model handles it fine. The claim is narrower and more useful: fine-tuning a small model on data that never punished invention created a failure mode the base model did not have.

Fixing it was mostly a data problem, and it worked. A later revision matched the 27B's 0.0% while routing action-shaped questions better than the teacher does - 27 of 27 against 20 of 27. Which also rules out the boring "it just got timid" explanation.

Two seeds is not a control

I need to include this one, because the rest of the post reads too clean without it.

Every claim I made over those two days was measured against a control group of two seeds with a 0.5-point spread. I said out loud, more than once, that a spread that tight looked wrong. And then I kept iterating anyway.

Eventually I ran a third control seed. It moved the control from 86.3% to 87.0% and widened the spread from 0.5 to 2.4 points.

group seeds call accuracy range
control 3 87.0 +/- 1.2 86.1 - 88.5
v12 3 87.7 +/- 1.7 85.6 - 88.9
v13 5 87.9 +/- 2.2 85.6 - 89.9

My own run picker started returning the control as the winner and printing "tied within noise" for all three. On the headline number, five dataset revisions bought nothing measurable. The two-seed control was not a control, it was a lucky pair, and everything measured against its spread was measured against noise.

That third seed cost 45 minutes. It should have been the first thing I did.

What survives, with ranges still disjoint from the control's:

control v13
args given tool 92.8 +/- 0.6 94.5 +/- 0.8
metric_query 51.3 +/- 3.8 87.7 +/- 7.7
propose_action 54.5 +/- 9.1 76.4 +/- 4.5
get_health 57.1 +/- 0.0 77.1 +/- 14.3
probe confabulation 9.7% 0.0%

So. Net result: the work did not improve aggregate accuracy. It changed the failure mode. The model gets about the same proportion of questions right, and is far less likely to be confidently wrong in a way that reaches real hardware. Those are different properties, and I managed to conflate them for two solid days.

One more negative result

The plan for RL was to retry GRPO now that the model gets enough right for a k-sample filter to find prompts in the 20-80% success band, where group-relative advantage is actually non-zero. Two previous attempts had died on a saturated reward.

Before spending an hour on rollouts, I spent twenty minutes sampling k=8 at temperature 1.0 and printing a histogram:

success rate training set (240) hard set (90)
always right 220 (91.7%) 71 (78.9%)
in band 0.2 - 0.8 8 (3.3%) 5 (5.6%)
always wrong 1 8 (8.9%)

Thirteen usable prompts. GRPO is not blocked by my optimizer or my reward function, it is blocked by a policy too confident to disagree with itself. Eight prompts that are wrong all eight times have exactly as little gradient as eight that are right all eight times.

Run the cheap diagnostic before the expensive optimizer. That histogram would have killed both previous attempts for twenty minutes each.

What I learned

Build the referee before the player. A frozen eval, a hash to stop you editing it, and an acceptance test proving every question is answerable by your plumbing. Do the question set before you look at the training data - if you can't write two hundred questions about the thing, you don't know what the thing is. And run the gold answers through the live system once, so that when the model misses, the miss is the model's and not your wiring. That acceptance test took an afternoon. It was the cheapest insurance in the whole project.

Check what your data generator silently guarantees. One line of template code meant the correct answer was always a literal substring of the question, so a whole capability was never trained and no per-tool metric could possibly show it. The fix is a grep: generate the set once, and look for answers sitting inside the questions. Any silent guarantee in the generator shows up in the data before it ever shows up in a chart.

The loss curve lies. A difference you would not notice was a quarter of the model's task accuracy. So don't make decisions off it. Keep an eval that finishes in a couple of minutes, run it after every single change, and treat val loss as a sanity check for "is it learning something" - not as a score.

Fine-tuning for structured output can manufacture confident nonsense. It removes the visible failure and can add an invisible one. Accuracy cannot see it, so score for it directly: does the named target exist? Is the ID in the catalog? Does a refusal claim something false about the system? Build that check once, keep it, and run it against the untuned base too - the base's failures are ugly and obvious, and that contrast is most of the lesson.

Your control group needs as many seeds as your challenger. A two-seed control with a suspiciously tight spread is not a yardstick. I let one underwrite two days of conclusions. If a spread looks too small to be real, it is - run the extra seed. It costs less than one rollout hour, and it tells you whether your "win" is a pair of dice.

That is the list. No magic hyperparameters, no trick I can paste into a tweet. The model ended up matching the 27B on confabulation (0.0%), routing action-shaped questions slightly better, and landing right on top of the control on aggregate accuracy. Same number of correct answers, different kind of wrong - that is the takeaway. And the biggest wins were not in the model at all. They were in the plumbing around it.

Next, the obvious bit: find the ~13 prompts that actually leave GRPO something to optimize against, and see if preference tuning does anything the data revisions did not. If it does not, I will keep the 4B as-is. It already answers "what's on the guest wifi?" well enough, and that is a question I actually ask, from the couch, more than you would think.

The whole thing runs on one 3090 sharing a desktop. The lab notebook, the per-run analyses, and every wrong turn are in the repo - because the wrong turns were most of the value. If the eval harness or the grammar setup is useful to you (or it breaks in an interesting way), tell me.

All for now...

-BadgerOps