Evidence layer
Observed outcomes
These rasters show the scenario-level evidence behind the matrix. They keep the same row order as the matrix. They do not rank observations from a single run.
Each mark shows one scenario outcome for one published stack. The data files include scenario descriptions and rationales. They also contain status text and evidence paths. They provide citation details. The appendix includes recorded stack metadata and transcript links.
Scenario evidence
The complete raster shows all 50 scenarios. The focused raster shows the seven multi-turn scenarios. The signature inventory groups identical 50-outcome vectors. Figure captions link to the complete text alternative in JSON and CSV.
Single call
Stacksingle-array-tagssingle-boolean-flagsingle-date-rangesingle-decimal-pricesingle-empty-argumentssingle-enum-formatsingle-integer-limitsingle-nested-optionssingle-optional-omittedsingle-optional-presentsingle-string-listsingle-weathersingle-zero-and-negative
gemma3:12b | Q4_K_M | ollama
gemma3:4b | Q4_K_M | ollama
granite3.1-dense:8b | Q4_K_M | llama.cpp
phi4-mini:latest | Q4_K_M | llama.cpp
llama3-groq-tool-use:8b | Q4_0 | ollama
hermes3:8b | Q4_0 | ollama
Qwen/Qwen2.5-1.5B-Instruct | Q3_K_M | llama.cpp
Qwen/Qwen2.5-1.5B-Instruct | Q4_K_M | llama.cpp
Qwen/Qwen2.5-1.5B-Instruct | Q8_0 | llama.cpp
Qwen/Qwen2.5-1.5B-Instruct | 4bit | MLX LM
Qwen/Qwen2.5-7B-Instruct | Q3_K_M | llama.cpp
Qwen/Qwen2.5-7B-Instruct | Q4_K_M | llama.cpp
Qwen/Qwen2.5-7B-Instruct | Q8_0 | llama.cpp
Qwen/Qwen2.5-7B-Instruct | 4bit | MLX LM
Qwen/Qwen2.5-7B-Instruct | 8bit | MLX LM
qwen2.5:7b-instruct | Q4_K_M | ollama
qwen3:0.6b | Q4_K_M | ollama
qwen3:1.7b | Q4_K_M | ollama
qwen3:14b | Q4_K_M | ollama
qwen3:4b | Q4_K_M | ollama
qwen3:8b | Q4_K_M | ollama
granite3.1-dense:8b | Q4_K_M | ollama
llama3.1:8b | Q4_K_M | ollama
meta-llama/Meta-Llama-3.1-8B-Instruct | Q3_K_M | llama.cpp
meta-llama/Meta-Llama-3.1-8B-Instruct | Q4_K_M | llama.cpp
meta-llama/Meta-Llama-3.1-8B-Instruct | Q8_0 | llama.cpp
meta-llama/Meta-Llama-3.1-8B-Instruct | 4bit | MLX LM
microsoft/Phi-4-mini-instruct | 4bit | MLX LM
phi4-mini:latest | Q4_K_M | ollama
mistral:7b | Q4_K_M | ollama
mistralai/Mistral-7B-Instruct-v0.3 | 4bit | MLX LM
watt-ai/watt-tool-8B | Q4_K_M | llama.cpp
Parallel calls
Stackparallel-city-timeparallel-different-toolsparallel-three-citiesparallel-three-different-toolsparallel-two-calculationsparallel-two-multiple-argsparallel-two-searchesparallel-weather-and-time
gemma3:12b | Q4_K_M | ollama
gemma3:4b | Q4_K_M | ollama
granite3.1-dense:8b | Q4_K_M | llama.cpp
phi4-mini:latest | Q4_K_M | llama.cpp
llama3-groq-tool-use:8b | Q4_0 | ollama
hermes3:8b | Q4_0 | ollama
Qwen/Qwen2.5-1.5B-Instruct | Q3_K_M | llama.cpp
Qwen/Qwen2.5-1.5B-Instruct | Q4_K_M | llama.cpp
Qwen/Qwen2.5-1.5B-Instruct | Q8_0 | llama.cpp
Qwen/Qwen2.5-1.5B-Instruct | 4bit | MLX LM
Qwen/Qwen2.5-7B-Instruct | Q3_K_M | llama.cpp
Qwen/Qwen2.5-7B-Instruct | Q4_K_M | llama.cpp
Qwen/Qwen2.5-7B-Instruct | Q8_0 | llama.cpp
Qwen/Qwen2.5-7B-Instruct | 4bit | MLX LM
Qwen/Qwen2.5-7B-Instruct | 8bit | MLX LM
qwen2.5:7b-instruct | Q4_K_M | ollama
qwen3:0.6b | Q4_K_M | ollama
qwen3:1.7b | Q4_K_M | ollama
qwen3:14b | Q4_K_M | ollama
qwen3:4b | Q4_K_M | ollama
qwen3:8b | Q4_K_M | ollama
granite3.1-dense:8b | Q4_K_M | ollama
llama3.1:8b | Q4_K_M | ollama
meta-llama/Meta-Llama-3.1-8B-Instruct | Q3_K_M | llama.cpp
meta-llama/Meta-Llama-3.1-8B-Instruct | Q4_K_M | llama.cpp
meta-llama/Meta-Llama-3.1-8B-Instruct | Q8_0 | llama.cpp
meta-llama/Meta-Llama-3.1-8B-Instruct | 4bit | MLX LM
microsoft/Phi-4-mini-instruct | 4bit | MLX LM
phi4-mini:latest | Q4_K_M | ollama
mistral:7b | Q4_K_M | ollama
mistralai/Mistral-7B-Instruct-v0.3 | 4bit | MLX LM
watt-ai/watt-tool-8B | Q4_K_M | llama.cpp
Streaming
Stackstreaming-array-argumentsstreaming-boolean-argumentstreaming-empty-argumentsstreaming-enum-argumentsstreaming-nested-argumentsstreaming-parallel-threestreaming-parallel-twostreaming-weather
gemma3:12b | Q4_K_M | ollama
gemma3:4b | Q4_K_M | ollama
granite3.1-dense:8b | Q4_K_M | llama.cpp
phi4-mini:latest | Q4_K_M | llama.cpp
llama3-groq-tool-use:8b | Q4_0 | ollama
hermes3:8b | Q4_0 | ollama
Qwen/Qwen2.5-1.5B-Instruct | Q3_K_M | llama.cpp
Qwen/Qwen2.5-1.5B-Instruct | Q4_K_M | llama.cpp
Qwen/Qwen2.5-1.5B-Instruct | Q8_0 | llama.cpp
Qwen/Qwen2.5-1.5B-Instruct | 4bit | MLX LM
Qwen/Qwen2.5-7B-Instruct | Q3_K_M | llama.cpp
Qwen/Qwen2.5-7B-Instruct | Q4_K_M | llama.cpp
Qwen/Qwen2.5-7B-Instruct | Q8_0 | llama.cpp
Qwen/Qwen2.5-7B-Instruct | 4bit | MLX LM
Qwen/Qwen2.5-7B-Instruct | 8bit | MLX LM
qwen2.5:7b-instruct | Q4_K_M | ollama
qwen3:0.6b | Q4_K_M | ollama
qwen3:1.7b | Q4_K_M | ollama
qwen3:14b | Q4_K_M | ollama
qwen3:4b | Q4_K_M | ollama
qwen3:8b | Q4_K_M | ollama
granite3.1-dense:8b | Q4_K_M | ollama
llama3.1:8b | Q4_K_M | ollama
meta-llama/Meta-Llama-3.1-8B-Instruct | Q3_K_M | llama.cpp
meta-llama/Meta-Llama-3.1-8B-Instruct | Q4_K_M | llama.cpp
meta-llama/Meta-Llama-3.1-8B-Instruct | Q8_0 | llama.cpp
meta-llama/Meta-Llama-3.1-8B-Instruct | 4bit | MLX LM
microsoft/Phi-4-mini-instruct | 4bit | MLX LM
phi4-mini:latest | Q4_K_M | ollama
mistral:7b | Q4_K_M | ollama
mistralai/Mistral-7B-Instruct-v0.3 | 4bit | MLX LM
watt-ai/watt-tool-8B | Q4_K_M | llama.cpp
Tool choice
Stacktool-choice-auto-calltool-choice-auto-no-calltool-choice-namedtool-choice-named-emptytool-choice-nonetool-choice-requiredtool-choice-required-empty
gemma3:12b | Q4_K_M | ollama
gemma3:4b | Q4_K_M | ollama
granite3.1-dense:8b | Q4_K_M | llama.cpp
phi4-mini:latest | Q4_K_M | llama.cpp
llama3-groq-tool-use:8b | Q4_0 | ollama
hermes3:8b | Q4_0 | ollama
Qwen/Qwen2.5-1.5B-Instruct | Q3_K_M | llama.cpp
Qwen/Qwen2.5-1.5B-Instruct | Q4_K_M | llama.cpp
Qwen/Qwen2.5-1.5B-Instruct | Q8_0 | llama.cpp
Qwen/Qwen2.5-1.5B-Instruct | 4bit | MLX LM
Qwen/Qwen2.5-7B-Instruct | Q3_K_M | llama.cpp
Qwen/Qwen2.5-7B-Instruct | Q4_K_M | llama.cpp
Qwen/Qwen2.5-7B-Instruct | Q8_0 | llama.cpp
Qwen/Qwen2.5-7B-Instruct | 4bit | MLX LM
Qwen/Qwen2.5-7B-Instruct | 8bit | MLX LM
qwen2.5:7b-instruct | Q4_K_M | ollama
qwen3:0.6b | Q4_K_M | ollama
qwen3:1.7b | Q4_K_M | ollama
qwen3:14b | Q4_K_M | ollama
qwen3:4b | Q4_K_M | ollama
qwen3:8b | Q4_K_M | ollama
granite3.1-dense:8b | Q4_K_M | ollama
llama3.1:8b | Q4_K_M | ollama
meta-llama/Meta-Llama-3.1-8B-Instruct | Q3_K_M | llama.cpp
meta-llama/Meta-Llama-3.1-8B-Instruct | Q4_K_M | llama.cpp
meta-llama/Meta-Llama-3.1-8B-Instruct | Q8_0 | llama.cpp
meta-llama/Meta-Llama-3.1-8B-Instruct | 4bit | MLX LM
microsoft/Phi-4-mini-instruct | 4bit | MLX LM
phi4-mini:latest | Q4_K_M | ollama
mistral:7b | Q4_K_M | ollama
mistralai/Mistral-7B-Instruct-v0.3 | 4bit | MLX LM
watt-ai/watt-tool-8B | Q4_K_M | llama.cpp
Multi-turn
Stackmulti-turn-calendar-followupmulti-turn-flight-bookingmulti-turn-inventory-ordermulti-turn-parallel-followupmulti-turn-routemulti-turn-ticket-updatemulti-turn-two-hop-chain
gemma3:12b | Q4_K_M | ollama
gemma3:4b | Q4_K_M | ollama
granite3.1-dense:8b | Q4_K_M | llama.cpp
phi4-mini:latest | Q4_K_M | llama.cpp
llama3-groq-tool-use:8b | Q4_0 | ollama
hermes3:8b | Q4_0 | ollama
Qwen/Qwen2.5-1.5B-Instruct | Q3_K_M | llama.cpp
Qwen/Qwen2.5-1.5B-Instruct | Q4_K_M | llama.cpp
Qwen/Qwen2.5-1.5B-Instruct | Q8_0 | llama.cpp
Qwen/Qwen2.5-1.5B-Instruct | 4bit | MLX LM
Qwen/Qwen2.5-7B-Instruct | Q3_K_M | llama.cpp
Qwen/Qwen2.5-7B-Instruct | Q4_K_M | llama.cpp
Qwen/Qwen2.5-7B-Instruct | Q8_0 | llama.cpp
Qwen/Qwen2.5-7B-Instruct | 4bit | MLX LM
Qwen/Qwen2.5-7B-Instruct | 8bit | MLX LM
qwen2.5:7b-instruct | Q4_K_M | ollama
qwen3:0.6b | Q4_K_M | ollama
qwen3:1.7b | Q4_K_M | ollama
qwen3:14b | Q4_K_M | ollama
qwen3:4b | Q4_K_M | ollama
qwen3:8b | Q4_K_M | ollama
granite3.1-dense:8b | Q4_K_M | ollama
llama3.1:8b | Q4_K_M | ollama
meta-llama/Meta-Llama-3.1-8B-Instruct | Q3_K_M | llama.cpp
meta-llama/Meta-Llama-3.1-8B-Instruct | Q4_K_M | llama.cpp
meta-llama/Meta-Llama-3.1-8B-Instruct | Q8_0 | llama.cpp
meta-llama/Meta-Llama-3.1-8B-Instruct | 4bit | MLX LM
microsoft/Phi-4-mini-instruct | 4bit | MLX LM
phi4-mini:latest | Q4_K_M | ollama
mistral:7b | Q4_K_M | ollama
mistralai/Mistral-7B-Instruct-v0.3 | 4bit | MLX LM
watt-ai/watt-tool-8B | Q4_K_M | llama.cpp
Correctly declines
Stacknegative-auto-arithmeticnegative-auto-knowledgenegative-greetingnegative-invalid-schemanegative-long-argumentnegative-plain-textnegative-unicode-argument
gemma3:12b | Q4_K_M | ollama
gemma3:4b | Q4_K_M | ollama
granite3.1-dense:8b | Q4_K_M | llama.cpp
phi4-mini:latest | Q4_K_M | llama.cpp
llama3-groq-tool-use:8b | Q4_0 | ollama
hermes3:8b | Q4_0 | ollama
Qwen/Qwen2.5-1.5B-Instruct | Q3_K_M | llama.cpp
Qwen/Qwen2.5-1.5B-Instruct | Q4_K_M | llama.cpp
Qwen/Qwen2.5-1.5B-Instruct | Q8_0 | llama.cpp
Qwen/Qwen2.5-1.5B-Instruct | 4bit | MLX LM
Qwen/Qwen2.5-7B-Instruct | Q3_K_M | llama.cpp
Qwen/Qwen2.5-7B-Instruct | Q4_K_M | llama.cpp
Qwen/Qwen2.5-7B-Instruct | Q8_0 | llama.cpp
Qwen/Qwen2.5-7B-Instruct | 4bit | MLX LM
Qwen/Qwen2.5-7B-Instruct | 8bit | MLX LM
qwen2.5:7b-instruct | Q4_K_M | ollama
qwen3:0.6b | Q4_K_M | ollama
qwen3:1.7b | Q4_K_M | ollama
qwen3:14b | Q4_K_M | ollama
qwen3:4b | Q4_K_M | ollama
qwen3:8b | Q4_K_M | ollama
granite3.1-dense:8b | Q4_K_M | ollama
llama3.1:8b | Q4_K_M | ollama
meta-llama/Meta-Llama-3.1-8B-Instruct | Q3_K_M | llama.cpp
meta-llama/Meta-Llama-3.1-8B-Instruct | Q4_K_M | llama.cpp
meta-llama/Meta-Llama-3.1-8B-Instruct | Q8_0 | llama.cpp
meta-llama/Meta-Llama-3.1-8B-Instruct | 4bit | MLX LM
microsoft/Phi-4-mini-instruct | 4bit | MLX LM
phi4-mini:latest | Q4_K_M | ollama
mistral:7b | Q4_K_M | ollama
mistralai/Mistral-7B-Instruct-v0.3 | 4bit | MLX LM
watt-ai/watt-tool-8B | Q4_K_M | llama.cpp