# qwen3.5-0.8b-mtp Concurrency Benchmark — MTP Disabled

Full sweep of concurrent agents against LM Studio (`LM Studio's local endpoint`, model `qwen3.5-0.8b-mtp` with multi-token prediction **disabled** in the server config).

- **Harness**: `benchmark/benchmark.js` (agent loop, streaming, token-level timing), `benchmark/sweep.js` (sweep driver)
- **Config per run**: 60s hard timeout per agent, 12000-token budget per agent, temperature 0.9, agents run a multi-turn loop until token budget or timeout
- **Method**: 24 sequential runs at concurrency 1..24; results in `results/qwen3.5-0.8b-mtp (MTP Disabled)/<N>/`, aggregate in `sweep-summary.{csv,json}` in the same folder
- **Same model ID, MTP off**: MTP is a decoding optimization, not the model's reasoning mode — so output still streams mostly through `reasoning_content` (this model is served as a thinking model). The point of this run is the throughput comparison: with MTP off the server can now batch requests.

## Results (1–24 agents)

| agents | wall(s) | content_tok | reason_tok | total_tok | comb_tok/s | comb_all_tok/s | per-agent_tok/s | per-agent_all_tok/s | scale_eff | TTFT_mean | TTFT_max | ok | timeout | err |
|-------:|--------:|------------:|-----------:|----------:|-----------:|---------------:|----------------:|---------------------:|----------:|----------:|---------:|----|--------:|----:|
| 1 | 60.0 | 4202 | 10119 | 14321 | 70 | 239 | 243.7 | 491.5 | 100% | 12.0s | 12.0s | 0 | 1 | 0 |
| 2 | 60.0 | 7474 | 13428 | 20902 | 124 | 348 | 178.8 | 307.1 | 146% | 19.3s | 26.9s | 0 | 2 | 0 |
| 3 | 60.0 | 9350 | 16050 | 25400 | 156 | 423 | 100.2 | 155.1 | 177% | 23.8s | 40.2s | 0 | 3 | 0 |
| 4 | 60.0 | 16727 | 11559 | 28286 | 279 | 471 | 61.8 | 89.2 | 197% | 16.4s | 18.7s | 0 | 4 | 0 |
| 5 | 60.0 | 8705 | 18328 | 27033 | 145 | 450 | 39.5 | 48.6 | 188% | 21.8s | 25.7s | 0 | 5 | 0 |
| 6 | 60.0 | 11148 | 17662 | 28810 | 186 | 480 | 13.7 | 16.3 | 201% | 30.8s | 49.5s | 0 | 6 | 0 |
| 7 | 60.0 | 11581 | 18170 | 29751 | 193 | 496 | 0.0 | 0.0 | 208% | 33.1s | 43.2s | 0 | 7 | 0 |
| 8 | 60.0 | 11150 | 19823 | 30973 | 186 | 516 | 8.7 | 13.4 | 216% | 29.3s | 35.7s | 0 | 8 | 0 |
| 9 | 60.0 | 10113 | 24636 | 34749 | 168 | 579 | 64.2 | 75.8 | 242% | 32.0s | 55.2s | 3 | 6 | 0 |
| 10 | 60.0 | 9258 | 28371 | 37629 | 154 | 627 | 44.2 | 77.7 | 262% | 28.0s | 42.5s | 3 | 7 | 0 |
| 11 | 60.0 | 6169 | 32388 | 38557 | 103 | 642 | 43.4 | 76.7 | 269% | 34.2s | 44.2s | 3 | 8 | 0 |
| 12 | 60.0 | 6041 | 32946 | 38987 | 101 | 649 | 33.9 | 72.3 | 272% | 31.4s | 40.1s | 5 | 7 | 0 |
| 13 | 60.0 | 3099 | 35386 | 38485 | 52 | 641 | 24.8 | 65.0 | 268% | 37.3s | 43.1s | 7 | 6 | 0 |
| 14 | 60.0 | 2816 | 35670 | 38486 | 47 | 641 | 20.0 | 60.9 | 268% | 36.4s | 45.8s | 8 | 6 | 0 |
| 15 | 60.0 | 2335 | 36536 | 38871 | 39 | 647 | 12.6 | 60.4 | 271% | 32.0s | 41.3s | 11 | 4 | 0 |
| 16 | 60.0 | 1584 | 38288 | 39872 | 26 | 664 | 17.0 | 59.0 | 278% | 37.8s | 41.1s | 10 | 6 | 0 |
| 17 | 60.0 | 2404 | 37614 | 40018 | 40 | 667 | 16.1 | 56.6 | 279% | 34.2s | 42.1s | 11 | 6 | 0 |
| 18 | 60.0 | 1526 | 38674 | 40200 | 25 | 670 | 11.5 | 56.3 | 280% | 33.9s | 41.0s | 13 | 5 | 0 |
| 19 | 60.0 | 139 | 34971 | 35110 | 2 | 585 | 2.1 | 48.7 | 245% | 36.2s | 36.2s | 18 | 1 | 0 |
| 20 | 60.0 | 1034 | 34104 | 35138 | 17 | 585 | 2.0 | 46.4 | 245% | 12.3s | 12.3s | 19 | 1 | 0 |
| 21 | 60.0 | 515 | 35187 | 35702 | 9 | 595 | 1.8 | 47.9 | 249% | 23.0s | 23.0s | 20 | 1 | 0 |
| 22 | 37.1 | 0 | 29232 | 29232 | 0 | 787 | 0.0 | 36.6 | 329% | - | - | 22 | 0 | 0 |
| 23 | 37.4 | 0 | 29091 | 29091 | 0 | 779 | 0.0 | 34.7 | 326% | - | - | 23 | 0 | 0 |
| 24 | 60.0 | 94 | 34562 | 34656 | 2 | 577 | 1.4 | 40.6 | 241% | 33.6s | 33.6s | 23 | 1 | 0 |

## Key findings

- **MTP was the bottleneck — confirmed.** Combined throughput now scales with concurrency: 239 tok/s solo → **670 tok/s at 18 agents (2.8x)**, peaking at 787 tok/s at 22 (3.3x, though that run ended early). Compare the MTP-enabled run, which was flat at ~180 tok/s. Disabling MTP lets LM Studio batch/parallelize this model like the others.
- **Still a thinking model**: output remains ~70–99% `reasoning_content`. At low concurrency it now produces substantial visible content (up to 16.7k content tokens at 4 agents), but from ~12 agents up, visible content collapses toward zero as reasoning + queueing consume the 60s window.
- **Very slow TTFT**: 12s solo, 16–55s under load — the reasoning preamble dominates. Runs 22–23 produced zero visible content at all.
- **Errors**: 0 across all 24 runs.

## Compare: MTP on vs off (all-token combined tok/s)

| agents | MTP on | MTP off |
|-------:|-------:|--------:|
| 1 | 202 | 239 |
| 4 | 169 | 471 |
| 8 | 182 | 516 |
| 12 | 181 | 649 |
| 16 | 183 | 664 |
| 20 | 173 | 585 |
| 24 | 172 | 577 |

## Charts

Generated with `benchmark/charts.py` (matplotlib), saved as PNG in `results/qwen3.5-0.8b-mtp (MTP Disabled)/charts/`:

| Chart | File | What it highlights |
|---|---|---|
| Dashboard (2x2 overview) | `dashboard_1_24.png` | Whole sweep in one view |
| Combined throughput | `combined_throughput.png` | Scaling restored (vs flat with MTP on) |
| Per-agent throughput | `per_agent_throughput.png` | Per-agent slowdown under contention |
| Scaling efficiency | `scaling_efficiency.png` | 100% → ~280–330% |
| Time to first token | `time_to_first_token.png` | 12s–55s thinking latency |
| Total tokens generated | `total_tokens_generated.png` | Content vs reasoning stack per concurrency |
| Outcome breakdown | `outcome_breakdown.png` | Zero errors |
| Combined vs per-agent | `combined_vs_per_agent.png` | The concurrency trade-off on one chart |
| Content vs reasoning | `reasoning_vs_content.png` | Reasoning dominates every run |
