When LM Studio and Ollama send different prompts for the same model

Same weight file, same messages, temperature 0: Ollama reported two more input tokens than LM Studio on phi-4, and the two runtimes returned different text. On Granite the counts matched. The cheapest check is one number you probably do not log.

A roll of punched paper tape splitting into two nearly identical strips whose hole patterns diverge in a few positions

Load the same model file into LM Studio and Ollama, send it the same messages at temperature 0, and the two runtimes can send different prompts. On phi-4 Q4_K_M they did: Ollama reported two more input tokens on every request we measured, and the two runtimes returned different text. On Granite 4.1 8B the counts matched exactly. Whether this bites you depends on the model file, and the cheapest way to find out is one number you are probably not logging.

What it looks like when you are not looking for it

Ask a model for a small amount of structured text at temperature 0 and the output should be deterministic. Same weights, same prompt, same temperature. That assumption sits underneath most local model comparisons, and it is the one that fails here.

On a four-row CSV prompt, LM Studio returned a header and four rows reading Alice,Clean the kitchen,2023-10-15 and similar. Ollama returned a header and four rows reading Alice,Prepare presentation,2023-11-15 and similar. Same shape, same length at sixty tokens each, not one row of content in common.

On an arithmetic probe asking for 18737 + 82415 + 97 × 3, which is 101443, LM Studio answered 18826 in three tokens with no working shown. Ollama answered 18737 + 82415 + 97 * 3 = 101249 in seventeen tokens: every term right, the final addition 194 below the correct value. Both wrong, and neither of them a sampling curiosity. They are what you get when the strings handed to the tokenizer are not the same string.

What the token counts show

Both runtimes report how many tokens they read. Same messages, same model file, same machine:

Promptphi-4 LM Studiophi-4 OllamaGranite LM StudioGranite Ollama
JSON invoice extract75777777
Short CSV artifact39414141
Fixed 60-line output61636363
Prefill probe19212121

Two tokens on every phi-4 prompt, and zero on every Granite prompt. The constant offset is the useful part: a fixed gap across four different prompts is structural rather than content, because anything in the messages would vary with the message. Granite is the control that keeps this honest, and it also limits the claim. A constant offset on one file and a zero on another is evidence that template selection can diverge between these runtimes, not that it always does. A single-model test would have produced a wrong answer in one direction or the other.

One thing the counts do not tell you: where the two tokens are, or what else changed with them. A matching count is not a matching string.

What we logged, and what we did not check

phi-4 Q4_K_M carries a chat template in its GGUF metadata, and it is ChatML with a specific separator, <|im_sep|>, between the role tag and the content. Ollama's import of that same file logged autodetected template chatml.

Asking Ollama which template it applied settles most of the mechanism. ollama show lmb-runtime-phi4 --template prints a ChatML variant that puts a newline where this GGUF puts <|im_sep|> after each role tag, and appends one more newline after each <|im_end|>. On these requests, every one of which carries a system and a user message in the same two-message shape, the swaps are one token for one token and invisible to the count, while the two added newlines, one after each closed turn, are exactly the constant +2 in the table. The same diff shows why a matched count is not a matched string: <|im_sep|> and a newline are different tokens even where the totals agree. What remains unlogged is the rendered prompt itself, so a second difference riding alongside the template is not excluded. The autodetection is also the import path: whether pulls from Ollama's own model library hit the same template is untested, and our number is for a file imported by hand.

Forcing one template, and the gap that fix opens

Putting the GGUF's own template into an Ollama Modelfile makes the counts line up: 75 against 75, 39 against 39, 61 against 61, 19 against 19. The remaining wrinkle is that Ollama's Modelfile takes a Go template and a GGUF's embedded Jinja will not parse in it, so the template has to be rewritten as an equivalent Go template that renders the same string:

FROM /path/to/model-Q4_K_M.gguf
TEMPLATE """{{- range .Messages }}<|im_start|>{{ .Role }}<|im_sep|>{{ .Content }}<|im_end|>{{ end }}<|im_start|>assistant<|im_sep|>"""
PARAMETER stop "<|im_start|>"
PARAMETER stop "<|im_end|>"

Those two PARAMETER lines are still worth writing, even though they turned out redundant here. The run behind the aligned counts set only TEMPLATE, and ollama show --modelfile on the result lists <|im_start|> and <|im_end|> stops anyway, the same two the autodetected model carries, so a TEMPLATE-only Modelfile did not leave the model stop-less, and on this file the defaults happened to be right. On a different file the same shortcut can leave a stop list that belongs to whatever the autodetector guessed, and a stop mismatch survives temperature 0 without showing up in any token count.

What aligning the prompt did and did not fix

With input tokens identical on all four phi-4 prompts, the answers were still not all the same:

  • The JSON invoice extract came back identical, 39 tokens on both sides and the same text throughout.
  • The fixed 60-line output agreed line for line through the first twenty lines and then stopped agreeing, finishing 370 tokens against 367.
  • The CSV prompt differed again, with no shared row, though Ollama's rows did change with its new prompt.
  • The arithmetic probe differed again: LM Studio repeated 18826 exactly, while Ollama's worked answer moved from 101249 to 101149, still wrong but changed by the template fix.

So the prompt was not the whole story. Aligning it demonstrably moved Ollama's outputs (new CSV rows, a different wrong sum) without moving them toward LM Studio's, and LM Studio's own answers stayed fixed across both runs. The stop list can be crossed off for this pair as well, since both models ended up with the same two stops. What remains untested is kernel and reduction order, batching, floating point accumulation, and a sampler difference that temperature 0 does not by itself exclude.

The speed question, such as it is

Any runtime comparison is contaminated by this until the prompt is matched, so it is worth separating the two. With the prompt aligned and output lengths held equal, only two of the ten cells still qualify: five prompts on two models, filtered to the rows where both sides generated the same number of tokens. Both columns are output tokens over the wall-clock seconds of the whole request:

ModelOutput tokens LM / OllamaLM Studio tok/sOllama tok/sLead
phi-4 Q4_K_M39 / 3922.3722.141.0%
Granite 4.1 8B30 / 3040.8735.3115.8%

Both favour LM Studio. Repeats inside a session spread by 0.4% to 0.6% on the short prompts and up to about 3.8% on the 60-line cell, so the 1.0% lead is about twice its own prompt's spread and inside the noise a longer prompt produces, while the 15.8% lead is well outside anything the Granite session shows. Two cells of three passes each are not a ranking.

Time to first token is steadier, and it is the one measurement that survived the alignment with the same sign on both models. A probe capped at max_tokens=1 reads under 40 input tokens: LM Studio at 51 ms against Ollama's 108 ms on phi-4, and 32 ms against 101 ms on Granite, so 57 ms and 69 ms in LM Studio's favour. Each is a single probe, and the probe adds fixed request overhead to both sides, so treat it as an upper bound on the difference rather than a clean prefill measurement.

Load time is a third number and it points the other way. A first Ollama request after a model switch carried 7.919 s of load time inside a 10.906 s request; the same work once resident took 2.99 s against LM Studio's 3.03 s. Ollama evicts on a keep-alive timer, so a comparison that does not discard warmup passes is timing the loader.

What to do with this

  • Log the input token count from both runtimes before believing a comparison. A constant offset is a template difference, it costs one field, and it catches the case above in a single request.
  • Log the rendered prompt when the counts disagree. A matching count is not a matching string, and the counts cannot tell you which part of the template moved.
  • Set the stop list whenever you set a template. Swapping one without the other is the easiest way to create a new difference while fixing an old one.
  • Keep output length next to every rate. A wall-clock number divided by a token count is only a rate if both sides generated the same amount of text.
  • Do not mix a decode-only counter with a whole-request rate. Ollama reports eval_count against eval_duration, which covers decode only, and prompt_eval_count separately for the prompt. LM Studio's OpenAI-compatible endpoint returns token counts, not a rate. The table above sidesteps this by measuring the same thing on both sides (output tokens over wall clock), and comparing Ollama's native figure against that kind of number favours Ollama by construction.
  • Discard warmup passes.

Caveats

  • Two models, one quantisation each, three kept repeats per cell, on a MacBook Pro M4 Pro with 64 GB. Both Q4_K_M, temperature 0, top_p 1. Three passes make a measurement repeatable, not a distribution.
  • The answer comparison is phi-4 only. Granite was not run through the same four-prompt test.
  • The two extra tokens are located from the template Ollama reports applying, not from a logged rendered prompt. What split the outputs after alignment is unexplained; the stop lists on both models were identical.