# qwen3.5-0.8b-mtp Concurrency Benchmark

Full sweep of concurrent agents against LM Studio (`LM Studio's local endpoint`, model `qwen3.5-0.8b-mtp`).

- **Harness**: `benchmark/benchmark.js` (agent loop, streaming, token-level timing), `benchmark/sweep.js` (sweep driver)
- **Config per run**: 60s hard timeout per agent, 12000-token budget per agent, temperature 0.9, agents run a multi-turn loop until token budget or timeout
- **Method**: 24 sequential runs at concurrency 1..24; results in `results/qwen3.5-0.8b-mtp/<N>/`, aggregate in `sweep-summary.{csv,json}` in the same folder
- **Warning — this model is served as a pure "thinking" model**: every request streams 100% of its output through `reasoning_content`. Probing confirmed `content: ''` and `usage.reasoning_tokens == completion_tokens` for a 300-token request, and `"thinking":{"type":"disabled"}` has no effect. So `reason_tok` is essentially all tokens, `content_tok` ≈ 0 except at the lowest concurrency (where the model finally finishes its reasoning preamble within the 60s window).

## Results (1–24 agents)

| agents | wall(s) | content_tok | reason_tok | total_tok | comb_tok/s | comb_all_tok/s | per-agent_tok/s | per-agent_all_tok/s | scale_eff | TTFT_mean | TTFT_max | ok | timeout | err |
|-------:|--------:|------------:|-----------:|----------:|-----------:|---------------:|----------------:|---------------------:|----------:|----------:|---------:|----|--------:|----:|
| 1 | 60.0 | 8623 | 3513 | 12136 | 144 | 202 | 235.3 | 208.1 | 100% | 21.8s | 21.8s | 0 | 1 | 0 |
| 2 | 60.0 | 3944 | 5798 | 9742 | 66 | 162 | 0.0 | 0.0 | 80% | 33.8s | 39.6s | 0 | 2 | 0 |
| 3 | 60.0 | 2312 | 7718 | 10030 | 39 | 167 | 0.0 | 0.0 | 83% | 36.7s | 42.7s | 0 | 3 | 0 |
| 4 | 60.0 | 3147 | 6980 | 10127 | 52 | 169 | 0.0 | 0.0 | 84% | 35.8s | 52.0s | 0 | 4 | 0 |
| 5 | 60.0 | 1017 | 9442 | 10459 | 17 | 174 | 0.0 | 0.0 | 86% | 43.9s | 45.1s | 0 | 5 | 0 |
| 6 | 60.0 | 298 | 10586 | 10884 | 5 | 181 | 0.0 | 0.0 | 90% | 53.8s | 57.7s | 0 | 6 | 0 |
| 7 | 60.0 | 518 | 10152 | 10670 | 9 | 178 | 0.0 | 0.0 | 88% | 40.3s | 40.3s | 0 | 7 | 0 |
| 8 | 60.0 | 178 | 10733 | 10911 | 3 | 182 | 0.0 | 0.0 | 90% | 51.7s | 51.7s | 0 | 8 | 0 |
| 9 | 60.0 | 840 | 10004 | 10844 | 14 | 181 | 0.0 | 0.0 | 90% | 19.4s | 19.4s | 0 | 9 | 0 |
| 10 | 60.0 | 0 | 10861 | 10861 | 0 | 181 | 0.0 | 0.0 | 90% | - | - | 0 | 10 | 0 |
| 11 | 60.0 | 0 | 10994 | 10994 | 0 | 183 | 0.0 | 0.0 | 91% | - | - | 0 | 11 | 0 |
| 12 | 60.0 | 233 | 10658 | 10891 | 4 | 181 | 0.0 | 0.0 | 90% | 42.7s | 42.7s | 0 | 12 | 0 |
| 13 | 60.0 | 0 | 11056 | 11056 | 0 | 184 | 0.0 | 0.0 | 91% | - | - | 0 | 13 | 0 |
| 14 | 60.0 | 0 | 11012 | 11012 | 0 | 183 | 0.0 | 0.0 | 91% | - | - | 0 | 14 | 0 |
| 15 | 60.0 | 0 | 10867 | 10867 | 0 | 181 | 0.0 | 0.0 | 90% | - | - | 0 | 15 | 0 |
| 16 | 60.0 | 0 | 10971 | 10971 | 0 | 183 | 0.0 | 0.0 | 91% | - | - | 0 | 16 | 0 |
| 17 | 60.0 | 0 | 10841 | 10841 | 0 | 181 | 0.0 | 0.0 | 90% | - | - | 0 | 17 | 0 |
| 18 | 60.0 | 0 | 10946 | 10946 | 0 | 182 | 0.0 | 0.0 | 90% | - | - | 0 | 18 | 0 |
| 19 | 60.0 | 0 | 10772 | 10772 | 0 | 179 | 0.0 | 0.0 | 89% | - | - | 0 | 19 | 0 |
| 20 | 60.0 | 0 | 10365 | 10365 | 0 | 173 | 0.0 | 0.0 | 86% | - | - | 0 | 20 | 0 |
| 21 | 60.0 | 0 | 10946 | 10946 | 0 | 182 | 0.0 | 0.0 | 90% | - | - | 0 | 21 | 0 |
| 22 | 60.0 | 0 | 11318 | 11318 | 0 | 189 | 0.0 | 0.0 | 94% | - | - | 0 | 22 | 0 |
| 23 | 60.0 | 0 | 10913 | 10913 | 0 | 182 | 0.0 | 0.0 | 90% | - | - | 0 | 23 | 0 |
| 24 | 60.0 | 0 | 10317 | 10317 | 0 | 172 | 0.0 | 0.0 | 85% | - | - | 0 | 24 | 0 |

*Note: `per-agent_tok/s` reads 0.0 at concurrency 2+ because the content-window metric has no visible content to time. `TTFT` is "-" where no visible content ever arrived in the 60s window.*

## Key findings

- **No visible output under load**: from concurrency ~10 up, agents produce **zero content tokens** in 60s — the entire budget goes to `reasoning_content` (all 10k+ tokens). Even the TTFT-to-first-content never arrives.
- **Massive reasoning preamble**: even solo, first visible content takes ~22s and only appears after ~3.5k reasoning tokens. Probing a fresh request returned `content: ''` with 100% reasoning.
- **Flat throughput — no concurrency scaling**: combined (all-token) throughput is ~180 tok/s from 2 to 24 agents (202 solo). Total per run ≈ one stream's worth (~10–11k tokens). This is the signature of **request serialization** — consistent with MTP (multi-token prediction) using a decode path that LM Studio cannot batch, so every request queues on one slot. Concurrency adds only latency (TTFT 20–53s), never throughput.
- **Errors**: 0 across all 24 runs — but the model is effectively unusable for the NPC/controller workload as served.

## Charts

Generated with `benchmark/charts.py` (matplotlib), saved as PNG in `results/qwen3.5-0.8b-mtp/charts/`:

| Chart | File | What it highlights |
|---|---|---|
| Dashboard (2x2 overview) | `dashboard_1_24.png` | Whole sweep in one view |
| Combined throughput | `combined_throughput.png` | Flat ~180 tok/s — no scaling |
| Per-agent throughput | `per_agent_throughput.png` | Collapses to 0 (no visible content) |
| Scaling efficiency | `scaling_efficiency.png` | 100% → ~85–94% — concurrency actively hurts |
| Time to first token | `time_to_first_token.png` | 20–54s, or never |
| Total tokens generated | `total_tokens_generated.png` | All-token stack is ~100% reasoning |
| Outcome breakdown | `outcome_breakdown.png` | Zero errors; all timeouts |
| Combined vs per-agent | `combined_vs_per_agent.png` | The concurrency trade-off on one chart |
| Content vs reasoning | `reasoning_vs_content.png` | ~100% reasoning share per run |
