Measurement register

Tool-calling support matrix

Each row records one tested stack. The stack includes the model and its artifact. It also includes the server and decode mode. The matrix does not rank models or state verdicts.

Observed stacks

Capability matrix

How to read: Each ratio shows passed scenarios / all scenarios in that category.

pass model / response failure execution / server error not tested
n=1, no verdict: every published arm is a single run, so no cell carries a verdict; hatching marks this. identity status

Showing 32 stacks.

Model / quant / server Single call Parallel calls Streaming Tool choice Multi-turn Correctly declines
gemma3:12b (unverified artifact)
Q4_K_M / ollamapost-hoc (mapping)id: unresolveddetail
0/13not measurable 0/8not measurable 0/8not measurable 0/7not measurable 0/7not measurable 0/7not measurable
gemma3:4b (unverified artifact)
Q4_K_M / ollamapost-hoc (mapping)id: unresolveddetail
0/13not measurable 0/8not measurable 0/8not measurable 0/7not measurable 0/7not measurable 0/7not measurable
granite3.1-dense:8b (unverified artifact)
Q4_K_M / llama.cppgrammar (mapping)id: unresolveddetail
0/13 0/8 0/8 2/7 0/7 5/7
llama3-groq-tool-use:8b
Q4_0 / ollamapost-hoc (mapping)id: declareddetail
12/13 8/8 7/8 6/7 0/7 6/7
granite3.1-dense:8b
Q4_K_M / ollamapost-hoc (mapping)id: declareddetail
0/13 0/8 0/8 2/7 0/7 5/7
Model meta-llama/Meta-Llama-3.1-8B-Instruct
meta-llama/Meta-Llama-3.1-8B-Instruct
Q3_K_M / llama.cppgrammar (mapping)id: declareddetail
10/13 0/8 6/8 6/7 0/7 3/7
meta-llama/Meta-Llama-3.1-8B-Instruct
Q4_K_M / llama.cppgrammar (mapping)id: declareddetail
7/13 0/8 4/8 6/7 0/7 3/7
meta-llama/Meta-Llama-3.1-8B-Instruct
Q8_0 / llama.cppgrammar (mapping)id: declareddetail
7/13 0/8 4/8 6/7 0/7 3/7
meta-llama/Meta-Llama-3.1-8B-Instruct
4bit / MLX LMpost-hoc (run)id: declareddetail
0/13 0/8 0/8 2/7 0/7 5/7
llama3.1:8b
Q4_K_M / ollamapost-hoc (mapping)id: declareddetail
11/13 7/8 8/8 5/7 0/7 5/7
Model microsoft/Phi-4-mini-instruct
microsoft/Phi-4-mini-instruct
4bit / MLX LMpost-hoc (run)id: declareddetail
0/13 0/8 0/8 2/7 0/7 5/7
phi4-mini:latest
Q4_K_M / ollamapost-hoc (mapping)id: declareddetail
0/13 0/8 0/8 2/7 0/7 5/7
Model mistralai/Mistral-7B-Instruct-v0.3
mistralai/Mistral-7B-Instruct-v0.3
4bit / MLX LMpost-hoc (run)id: declareddetail
0/13 0/8 0/8 2/7 0/7 5/7
mistral:7b
Q4_K_M / ollamapost-hoc (mapping)id: declareddetail
11/13 8/8 6/8 6/7 0/7 5/7
hermes3:8b
Q4_0 / ollamapost-hoc (mapping)id: declareddetail
11/13 7/8 7/8 7/7 2/7 7/7
phi4-mini:latest (unverified artifact)
Q4_K_M / llama.cppgrammar (mapping)id: unresolveddetail
0/13 0/8 0/8 2/7 0/7 5/7
Model Qwen/Qwen2.5-1.5B-Instruct
Qwen/Qwen2.5-1.5B-Instruct
Q3_K_M / llama.cppgrammar (mapping)id: declareddetail
13/13 6/8 8/8 7/7 0/7 6/7
Qwen/Qwen2.5-1.5B-Instruct
Q4_K_M / llama.cppgrammar (mapping)id: declareddetail
13/13 8/8 8/8 6/7 0/7 7/7
Qwen/Qwen2.5-1.5B-Instruct
Q8_0 / llama.cppgrammar (mapping)id: declareddetail
13/13 8/8 8/8 7/7 0/7 6/7
Qwen/Qwen2.5-1.5B-Instruct
4bit / MLX LMpost-hoc (run)id: declareddetail
12/13 8/8 8/8 7/7 0/7 6/7
Model Qwen/Qwen2.5-7B-Instruct
Qwen/Qwen2.5-7B-Instruct
Q3_K_M / llama.cppgrammar (mapping)id: declareddetail
13/13 8/8 8/8 7/7 5/7 7/7
Qwen/Qwen2.5-7B-Instruct
Q4_K_M / llama.cppgrammar (mapping)id: declareddetail
13/13 8/8 8/8 7/7 2/7 7/7
Qwen/Qwen2.5-7B-Instruct
Q8_0 / llama.cppgrammar (mapping)id: declareddetail
13/13 8/8 8/8 7/7 4/7 7/7
Qwen/Qwen2.5-7B-Instruct
4bit / MLX LMpost-hoc (run)id: declareddetail
13/13 8/8 8/8 7/7 4/7 6/7
Qwen/Qwen2.5-7B-Instruct
8bit / MLX LMpost-hoc (run)id: declareddetail
0/13 0/8 0/8 2/7 0/7 5/7
qwen2.5:7b-instruct
Q4_K_M / ollamapost-hoc (mapping)id: declareddetail
11/13 8/8 8/8 7/7 6/7 6/7
qwen3:0.6b
Q4_K_M / ollamapost-hoc (mapping)id: declareddetail
7/13 5/8 4/8 6/7 0/7 5/7
qwen3:1.7b
Q4_K_M / ollamapost-hoc (mapping)id: declareddetail
11/13 8/8 7/8 6/7 4/7 6/7
qwen3:14b
Q4_K_M / ollamapost-hoc (mapping)id: declareddetail
12/13 8/8 8/8 7/7 5/7 7/7
qwen3:4b
Q4_K_M / ollamapost-hoc (mapping)id: declareddetail
13/13 8/8 8/8 7/7 1/7 6/7
qwen3:8b
Q4_K_M / ollamapost-hoc (mapping)id: declareddetail
12/13 8/8 7/8 7/7 4/7 7/7
watt-ai/watt-tool-8B
Q4_K_M / llama.cppgrammar (mapping)id: verifieddetail
0/13 0/8 0/8 2/7 0/7 5/7

Machine-readable observations: JSON and CSV. Scenario rasters are on the outcomes page; recorded metadata and transcripts are in the appendix.

One-run observations

Observed pass counts

Pass counts are shown in fixed stack order and withheld for rows with errors or skips. They are not scores or a ranking.

Figure 4. Observed pass-count strip plot Each dot shows one published stack observation from a single run. We plot a pass count only when every scenario has a pass or fail verdict. A separate panel lists rows with errors or skips. What this does not show: This figure does not show a model score distribution or rank models. It does not measure uncertainty or variation across repeated runs. Method and limitations - Data: JSON, CSV
Figure 4. Observed pass-count strip plot One pass-count dot per fully measurable published stack observation, with error-bearing or skipped rows listed separately and not assigned a pass count. Fully measurable observations 0 10 20 30 40 50 Not fully measurable - pass counts withheld gemma3:12b | Q4_K_M | ollama - 50 errors, 0 skipped; pass count not plotted gemma3:4b | Q4_K_M | ollama - 50 errors, 0 skipped; pass count not plotted meta-llama/Meta-Llama-3.1-8B-Instruct | Q3_K_M | llama.cpp - 7 errors, 0 skipped; pass count not plotted meta-llama/Meta-Llama-3.1-8B-Instruct | Q4_K_M | llama.cpp - 9 errors, 0 skipped; pass count not plotted meta-llama/Meta-Llama-3.1-8B-Instruct | Q8_0 | llama.cpp - 8 errors, 0 skipped; pass count not plotted granite3.1-dense:8b | Q4_K_M | llama.cpp: 7 passes phi4-mini:latest | Q4_K_M | llama.cpp: 7 passes llama3-groq-tool-use:8b | Q4_0 | ollama: 39 passes hermes3:8b | Q4_0 | ollama: 41 passes Qwen/Qwen2.5-1.5B-Instruct | Q3_K_M | llama.cpp: 40 passes Qwen/Qwen2.5-1.5B-Instruct | Q4_K_M | llama.cpp: 42 passes Qwen/Qwen2.5-1.5B-Instruct | Q8_0 | llama.cpp: 42 passes Qwen/Qwen2.5-1.5B-Instruct | 4bit | MLX LM: 41 passes Qwen/Qwen2.5-7B-Instruct | Q3_K_M | llama.cpp: 48 passes Qwen/Qwen2.5-7B-Instruct | Q4_K_M | llama.cpp: 45 passes Qwen/Qwen2.5-7B-Instruct | Q8_0 | llama.cpp: 47 passes Qwen/Qwen2.5-7B-Instruct | 4bit | MLX LM: 46 passes Qwen/Qwen2.5-7B-Instruct | 8bit | MLX LM: 7 passes qwen2.5:7b-instruct | Q4_K_M | ollama: 46 passes qwen3:0.6b | Q4_K_M | ollama: 27 passes qwen3:1.7b | Q4_K_M | ollama: 42 passes qwen3:14b | Q4_K_M | ollama: 47 passes qwen3:4b | Q4_K_M | ollama: 43 passes qwen3:8b | Q4_K_M | ollama: 45 passes granite3.1-dense:8b | Q4_K_M | ollama: 7 passes llama3.1:8b | Q4_K_M | ollama: 36 passes meta-llama/Meta-Llama-3.1-8B-Instruct | 4bit | MLX LM: 7 passes microsoft/Phi-4-mini-instruct | 4bit | MLX LM: 7 passes phi4-mini:latest | Q4_K_M | ollama: 7 passes mistral:7b | Q4_K_M | ollama: 36 passes mistralai/Mistral-7B-Instruct-v0.3 | 4bit | MLX LM: 7 passes watt-ai/watt-tool-8B | Q4_K_M | llama.cpp: 7 passes gemma3:12b | Q4_K_M | ollama: 50 errors, 0 skipped; pass count not plotted gemma3:4b | Q4_K_M | ollama: 50 errors, 0 skipped; pass count not plotted meta-llama/Meta-Llama-3.1-8B-Instruct | Q3_K_M | llama.cpp: 7 errors, 0 skipped; pass count not plotted meta-llama/Meta-Llama-3.1-8B-Instruct | Q4_K_M | llama.cpp: 9 errors, 0 skipped; pass count not plotted meta-llama/Meta-Llama-3.1-8B-Instruct | Q8_0 | llama.cpp: 8 errors, 0 skipped; pass count not plotted

Calibration notes

Method and limitations

Each cell measures one full stack. The stack combines a model, quant, server, and server version. A cell does not describe the model alone.

A failed observation applies only to the tested combination. It does not mean the weights are bad. The same weights can pass on one server and fail on another. When the evidence proves that difference, the cell includes a cause annotation.

When the result includes a transcript path, a failing observation links to the full request and response transcript. Legacy schema v1 results do not record transcript paths. Read the case studies in docs/case-studies/ for controlled comparisons.

The servers use different decode methods. llama.cpp compiles the supplied tool definitions into a GBNF grammar. It uses that grammar to constrain decoding. Ollama and MLX LM generate unconstrained text. They parse the tool call after decoding. Cross-band differences therefore reflect the full stack. Compare adjacent models only when they use the same server. Each row states whether the run recorded its decode mode or the site read the mode from the cited preset mapping. An unmapped preset has an unknown mode.

The site includes 50 distinct scenarios. Each published cell represents one run. Hatching marks that the cell has no verdict.

The case studies draw a verdict only after at least five runs per arm. The current case studies cover 90 runs across 18 quantization arms, and 40 runs across 8 arms for the peg-native anomaly.

Excluded rows

The quantization conclusion excludes Meta-Llama-3.1-8B-Instruct on llama.cpp (Q8_0, Q4_K_M, Q3_K_M). For this model, llama.cpp returns HTTP 500 on 7-9 of 50 scenarios per run ("does not match the expected peg-native format"). These results are server errors. They are not model failures and cannot be compared across arms. Read the peg-native case study.

All measurements used Apple M4 Max, 64GB; macOS 26.5.2.