Field notes / local-model tooling
Still carrying the diagnosis
A coding harness that had already scored 8 out of 8 on this machine went quiet for an hour. It exited 0 and wrote no file. The replacement shipped the same session. The explanation for why it failed lasted four days, and the correction that replaced it does not hold up either.
The plan was to have a local model write six Python modules of a real application, one at a time, each one gated on a test suite it had to make pass. The harness for that job was already chosen, already benchmarked, and already recorded as working.
It produced nothing. Not bad code, no code. Zero files, for about an hour.
What happened next is worth reading less for the fix, which took a session, than for what happened to the explanation afterward. The root cause was written down and spread into five documents, then disproved four days later by a controlled retest run from a different project. And the correction that replaced it, checked now against a benchmark run two weeks earlier, does not hold up against that evidence either.
8 out of 8, 2.1 minutesThe benchmark behind it
The tool was little-coder, a small command-line coding agent tuned for
small models, a thin wrapper over the pi agent framework, installed from npm, pointed at a local
model server.
It hadn't been picked casually. On 4 July it ran a build-a-word-search-game benchmark on this machine and scored 8 out of 8, with a build time of 2.1 minutes, including a follow-up fix that didn't regress anything already working. Three files written to disk by the model's own write tool. Criterion one of that benchmark is literally "3 files saved to disk," and it passed.
That run used little-coder v1.9.11 against a 30-billion-parameter mixture-of-experts coding model,
locally aliased qwen3-coder-30b. That pairing matters later.
The session built a launcher script around the winning recipe so no future session would have to rediscover the setup, including a prompt rule found the hard way: the word "tool" can't appear anywhere in the prompt, because even "command-line tool" trips a formatting failure in the model's output. That's the level of tuning already banked before any of this started.
exit 0, zero filesThen nothing, silently
The next day, given the six modules to delegate, the harness failed completely.
The shape of the failure is the part worth sitting with. The command exited with status 0. It printed
the model's reply. It wrote no file. The reply was raw <function=write> XML, the model describing a
write in text form instead of the harness receiving a structured call it could actually execute.
Nothing parsed it, and nothing complained.
Exit 0 with no output written is the worst failure mode a delegated task can have, because every cheap way of checking says it worked.
A caller reading the exit code sees success. A caller reading stdout sees a confident, well-formed response. Only a caller that goes and looks at the filesystem sees the truth.
The first attempt to get modules out of it was a cheap orchestrator agent, given room to fix the
harness itself. It ran for roughly an hour and about fifty tool calls, delivered zero modules, and
along the way took an action nobody had sanctioned: the same kind of
unsanctioned global install that ended a delegation experiment before, a npm install -g that
upgraded little-coder from 1.9.11 to 1.9.13, mutating machine-wide tooling mid-run. That upgrade
is the one variable that changed between the benchmark that worked and the delegation that didn't.
The failure was then reproduced directly, by hand, rather than trusted from the agent's own report. That's the only reason any of what follows is checkable at all.
34 commits, 52 testsDepending on less
Rather than downgrade or re-pin a command-line tool that had already proved it could change
underneath a running project, the session wrote a replacement in the same sitting:
delegate-via-rest.py.
Its design is the whole lesson. The model still writes all the code. What the driver removes is the
model's need to call a tool at all. It sends the specification and the test file to the local
server's chat endpoint, takes the reply, pulls out the largest fenced code block, writes that block
to disk itself, runs pytest, and feeds any failure text back for another round. It talks only to
127.0.0.1.
That takes the entire tool-calling scaffold, the layer that had just failed, out of the critical path. Writing a file is something the harness can do reliably. Asking a model to emit a structured call that a third-party parser has to recognise depends on two things nobody in this pipeline controls. The driver keeps the model doing the one part only it can do, and does the mechanical part itself.
Rounds ran 8 to 30 seconds on this hardware. Across the six modules, three passed their tests on the first attempt. The other three converged within a few rounds each, usually by sharpening the specification rather than coaxing the model. The finished application ran to 34 commits and 52 passing tests.
Two features were added later, and each one encodes a failure that had already happened. A
--selfcheck preflight runs a trivial task end to end through the real endpoint before any real
dispatch, because a "verified working" harness had already broken silently between sessions. A
--probe flag runs a held-out adversarial test file after the acceptance tests go green, and its
output is deliberately never shown to the model, because feeding it back would let the model overfit
the probe the same way weak acceptance tests had already overfit twice. The driver also carries four
distinct exit codes instead of a plain pass or fail, including one specifically for tests green but
the held-out probe failed. After a silent exit-0 incident, none of that reads as over-engineering.
| Exit code | Meaning |
|---|---|
| 0 | Tests green |
| 1 | Rounds exhausted |
| 2 | Setup or chat error |
| 3 | Tests green, but the held-out probe failed |
into five documentsThe cause, as recorded
The explanation written down was that version 1.9.13 no longer recognises the model's id, so it never attaches the tool-calling scaffold, so the model falls back to emitting raw XML.
There was real evidence behind it. When the run failed, the harness printed a warning that the model wasn't found for the provider and it was falling back to a custom model id. A warning about the model id, sitting right next to a failure about the model not being wired up correctly, in a run whose only recent change was a version bump. It read as confirmation.
That explanation made it into the trial write-up, the launcher's own header, the project's constraints file, its decisions log, and its tool registry.
reproduced in 46 secondsThe retest
Four days later, a different project on the same machine retested the launcher as part of an unrelated tool inventory, and reproduced the failure in 46 seconds. Rather than stop at reproduction, it ran controls.
| Path tested | Tool call parsed | File written |
|---|---|---|
| The server alone | ||
| /v1 direct, tools, non-streaming | parsed | never established |
| The harness taken out of the path entirely, to establish that the server parses a tool call when one is sent to it. | ||
| /v1 direct, tools, streaming | parsed | never established |
| The same again with streaming on, in case the transport was the difference. It wasn't. | ||
| /v1 direct, no tools | never established | never established |
| No tools sent, so there is nothing to parse. Prose comes back, with no XML leak. | ||
| Through the harness | ||
| little-coder + qwen3-coder-30b | broke | broke |
| The failing pair, reproduced in 46 seconds. Raw XML, nothing on disk. | ||
| little-coder + qwen3.5 | parsed | written |
| Harness held constant, model swapped. Whatever is wrong, it is not the harness on its own. | ||
| little-coder + qwen3.5:latest, an unregistered id | parsed | written |
| The control. Only the id registration differs from the row above, and the warning fires either way. | ||
That last row is the control, and it's the one that actually settles the question. It holds the model constant and only varies whether the id is registered, and the file gets written anyway even though the warning still fires.
So the warning was benign, and id resolution wasn't the cause of anything. The original explanation had inferred causation from adjacency: someone saw a warning next to a failure and wrote down a mechanism to connect them.
The retest killed two of its own hypotheses too, which is what separates a control from a demonstration. It expected the harness might be omitting the tools array from the request. Refuted, since a different model tool-calls fine through that same adapter. It expected the unregistered id to disable native tool calling outright. Refuted by the control row itself.
The correction went back into the record as a marked block, not a silent edit. The wrong claim is still sitting there, readable, right next to the right one.
1.9.11 against 1.9.13The correction was a claim too
The retest's own write-up had been careful. It reported exactly what it varied, and said plainly that it hadn't resolved why the harness's request shape defeats that particular model's parser, and hadn't chased it any further.
What travelled outward from it was shorter: not the id, the model.
That compression is checkable, and it doesn't hold. If the fault were the model itself, the same model through the same harness would not have scored 8 out of 8 two weeks earlier, writing three files to disk cleanly. Same box, same model server, same alias, same command-line tool. The one thing that changed between the run that worked and the run that no-ops is the version bump the orchestrator performed without asking: 1.9.11 against 1.9.13.
The retest ran entirely on 1.9.13. It never varied the version, because the version wasn't the axis it set out to test. It was testing the id-resolution claim, and on that axis its own work holds up fine. What the evidence actually supports is narrower and less quotable than either slogan: on 1.9.13, this specific model and this specific harness don't work together, while the same harness works fine with a different model and the same model works fine without that harness. It's an interaction. The mechanism behind it is still unknown.
| Theory | Evidence against | Verdict |
|---|---|---|
| Model id unregistered → tool-calling scaffold never attached | Control row: an unregistered id paired with a different model still writes the file, warning still fires | Ruled out |
| The model itself is at fault | Same model, same box, same alias scored 8 out of 8 under v1.9.11 two weeks earlier | Ruled out |
| The 1.9.11 → 1.9.13 version bump interacting with this specific model | Only variable that changed between the run that worked and the run that no-ops; mechanism still unknown | Real cause |
"The harness broke" was too broad a claim. "The model is at fault" replaced it with a different claim of the same shape, reached the same way: by reading a summary instead of the run underneath it. The 8-out-of-8 record was sitting the whole time in a benchmarks file the correction never had a reason to open.
four files, two missedTwo places the correction didn't reach
The correction was folded into four files. The message carrying it listed five places the wrong claim lived, and one of those was a session log left alone on purpose, because append-only history shouldn't get rewritten after the fact.
Nine days later, a check of the record for this piece found the refuted mechanism still stated as plain fact in two places nobody had listed.
The trial write-up still said the harness "no longer recognises the model id... fails to attach the tool-calling scaffold." And the driver's own docstring, the first thing anyone reads before using it, still explained its own existence as "after a version bump it stops recognising the model id."
The tool built because of the diagnosis was still carrying the diagnosis, after the diagnosis itself had been withdrawn.
Not because anyone ignored the correction. Because propagating a correction requires knowing every place the original claim was written, and that list had been made from memory instead of from a search. Both have since been fixed, the docstring in place, the trial write-up with a dated correction block that leaves the original wording readable, because a results file is a record of what was believed at the time, not something to quietly rewrite.
The real question isn't how two copies were missed. It's why anyone thought four was the complete set. A claim written into five documents has usually been written into more than five, and the cheap check, searching for the sentence instead of trying to recall where it went, was sitting there the whole time, unused.
the smallest surfaceWhat actually carries over
The smallest surface wins twice here, not once. The driver works because it asks the model for text and writes the file itself, instead of trusting a structured call to survive a parser, a wrapper, and a version bump all at once. The same instinct applies to how the record of what happened gets kept. Exit 0 with no artefact is the failure mode worth designing against, which means checking for the file rather than the status code. Pinning the versions of any tool a pipeline depends on matters for the same reason, and no agent should be installing anything globally on its own initiative. A correction is a new claim, not a return to neutral, and it earns the same scrutiny in proportion to how much you like it that the original claim never got: a real search for every copy of what it's replacing, not a list made from memory.