Measurement register
Tool-calling support matrix
Each row records one tested stack. The stack includes the model and its artifact. It also includes the server and decode mode. The matrix does not rank models or state verdicts.
Observed stacks
Capability matrix
How to read: Each ratio shows passed scenarios / all scenarios in that category.
Showing 32 stacks.
| Model / quant / server | Single call | Parallel calls | Streaming | Tool choice | Multi-turn | Correctly declines |
|---|---|---|---|---|---|---|
| gemma3:12b (unverified artifact) | 0/13not measurable | 0/8not measurable | 0/8not measurable | 0/7not measurable | 0/7not measurable | 0/7not measurable |
| gemma3:4b (unverified artifact) | 0/13not measurable | 0/8not measurable | 0/8not measurable | 0/7not measurable | 0/7not measurable | 0/7not measurable |
| granite3.1-dense:8b (unverified artifact) | 0/13 | 0/8 | 0/8 | 2/7 | 0/7 | 5/7 |
| llama3-groq-tool-use:8b | 12/13 | 8/8 | 7/8 | 6/7 | 0/7 | 6/7 |
| granite3.1-dense:8b | 0/13 | 0/8 | 0/8 | 2/7 | 0/7 | 5/7 |
| Model meta-llama/Meta-Llama-3.1-8B-Instruct | ||||||
| meta-llama/Meta-Llama-3.1-8B-Instruct | 10/13 | 0/8 | 6/8 | 6/7 | 0/7 | 3/7 |
| meta-llama/Meta-Llama-3.1-8B-Instruct | 7/13 | 0/8 | 4/8 | 6/7 | 0/7 | 3/7 |
| meta-llama/Meta-Llama-3.1-8B-Instruct | 7/13 | 0/8 | 4/8 | 6/7 | 0/7 | 3/7 |
| meta-llama/Meta-Llama-3.1-8B-Instruct | 0/13 | 0/8 | 0/8 | 2/7 | 0/7 | 5/7 |
| llama3.1:8b | 11/13 | 7/8 | 8/8 | 5/7 | 0/7 | 5/7 |
| Model microsoft/Phi-4-mini-instruct | ||||||
| microsoft/Phi-4-mini-instruct | 0/13 | 0/8 | 0/8 | 2/7 | 0/7 | 5/7 |
| phi4-mini:latest | 0/13 | 0/8 | 0/8 | 2/7 | 0/7 | 5/7 |
| Model mistralai/Mistral-7B-Instruct-v0.3 | ||||||
| mistralai/Mistral-7B-Instruct-v0.3 | 0/13 | 0/8 | 0/8 | 2/7 | 0/7 | 5/7 |
| mistral:7b | 11/13 | 8/8 | 6/8 | 6/7 | 0/7 | 5/7 |
| hermes3:8b | 11/13 | 7/8 | 7/8 | 7/7 | 2/7 | 7/7 |
| phi4-mini:latest (unverified artifact) | 0/13 | 0/8 | 0/8 | 2/7 | 0/7 | 5/7 |
| Model Qwen/Qwen2.5-1.5B-Instruct | ||||||
| Qwen/Qwen2.5-1.5B-Instruct | 13/13 | 6/8 | 8/8 | 7/7 | 0/7 | 6/7 |
| Qwen/Qwen2.5-1.5B-Instruct | 13/13 | 8/8 | 8/8 | 6/7 | 0/7 | 7/7 |
| Qwen/Qwen2.5-1.5B-Instruct | 13/13 | 8/8 | 8/8 | 7/7 | 0/7 | 6/7 |
| Qwen/Qwen2.5-1.5B-Instruct | 12/13 | 8/8 | 8/8 | 7/7 | 0/7 | 6/7 |
| Model Qwen/Qwen2.5-7B-Instruct | ||||||
| Qwen/Qwen2.5-7B-Instruct | 13/13 | 8/8 | 8/8 | 7/7 | 5/7 | 7/7 |
| Qwen/Qwen2.5-7B-Instruct | 13/13 | 8/8 | 8/8 | 7/7 | 2/7 | 7/7 |
| Qwen/Qwen2.5-7B-Instruct | 13/13 | 8/8 | 8/8 | 7/7 | 4/7 | 7/7 |
| Qwen/Qwen2.5-7B-Instruct | 13/13 | 8/8 | 8/8 | 7/7 | 4/7 | 6/7 |
| Qwen/Qwen2.5-7B-Instruct | 0/13 | 0/8 | 0/8 | 2/7 | 0/7 | 5/7 |
| qwen2.5:7b-instruct | 11/13 | 8/8 | 8/8 | 7/7 | 6/7 | 6/7 |
| qwen3:0.6b | 7/13 | 5/8 | 4/8 | 6/7 | 0/7 | 5/7 |
| qwen3:1.7b | 11/13 | 8/8 | 7/8 | 6/7 | 4/7 | 6/7 |
| qwen3:14b | 12/13 | 8/8 | 8/8 | 7/7 | 5/7 | 7/7 |
| qwen3:4b | 13/13 | 8/8 | 8/8 | 7/7 | 1/7 | 6/7 |
| qwen3:8b | 12/13 | 8/8 | 7/8 | 7/7 | 4/7 | 7/7 |
| watt-ai/watt-tool-8B | 0/13 | 0/8 | 0/8 | 2/7 | 0/7 | 5/7 |
Machine-readable observations: JSON and CSV. Scenario rasters are on the outcomes page; recorded metadata and transcripts are in the appendix.
One-run observations
Observed pass counts
Pass counts are shown in fixed stack order and withheld for rows with errors or skips. They are not scores or a ranking.
Calibration notes
Method and limitations
Each cell measures one full stack. The stack combines a model, quant, server, and server version. A cell does not describe the model alone.
A failed observation applies only to the tested combination. It does not mean the weights are bad. The same weights can pass on one server and fail on another. When the evidence proves that difference, the cell includes a cause annotation.
When the result includes a transcript path, a failing observation links to the full request and response transcript. Legacy schema v1 results do not record transcript paths. Read the case studies in docs/case-studies/ for controlled comparisons.
The servers use different decode methods. llama.cpp compiles the supplied tool definitions into a GBNF grammar. It uses that grammar to constrain decoding. Ollama and MLX LM generate unconstrained text. They parse the tool call after decoding. Cross-band differences therefore reflect the full stack. Compare adjacent models only when they use the same server. Each row states whether the run recorded its decode mode or the site read the mode from the cited preset mapping. An unmapped preset has an unknown mode.
The site includes 50 distinct scenarios. Each published cell represents one run. Hatching marks that the cell has no verdict.
The case studies draw a verdict only after at least five runs per arm. The current case studies cover 90 runs across 18 quantization arms, and 40 runs across 8 arms for the peg-native anomaly.
Excluded rows
The quantization conclusion excludes Meta-Llama-3.1-8B-Instruct on llama.cpp (Q8_0, Q4_K_M, Q3_K_M). For this model, llama.cpp returns HTTP 500 on 7-9 of 50 scenarios per run ("does not match the expected peg-native format"). These results are server errors. They are not model failures and cannot be compared across arms. Read the peg-native case study.
All measurements used Apple M4 Max, 64GB; macOS 26.5.2.