# qwen3-4b Concurrency Benchmark

Full sweep of concurrent agents against LM Studio (`LM Studio's local endpoint`, model `qwen3-4b`).

- **Harness**: `benchmark/benchmark.js` (agent loop, streaming, token-level timing), `benchmark/sweep.js` (sweep driver)
- **Config per run**: standard settings — 60s hard timeout, 12000-token total budget, `--output-cap 12000`, 32768-token context, temperature 0.9, agents run a multi-turn loop
- **Quants note**: this run is served with **K and V cache quants of Q4_0**.
- **Method**: 24 sequential runs at concurrency 1..24; results in `results/qwen3-4b/<N>/`, aggregate in `sweep-summary.{csv,json}` in the same folder
- **Reasoning tokens**: qwen3-4b streams hidden thinking via `reasoning_content`; its share explodes with concurrency (21% of tokens solo to **~95% at 22–24 agents**).

## Results (1–24 agents)

| agents | wall(s) | content_tok | reason_tok | total_tok | comb_tok/s | comb_all_tok/s | per-agent_tok/s | per-agent_all_tok/s | scale_eff | TTFT_mean | TTFT_max | ok | timeout | err |
|-------:|--------:|------------:|-----------:|----------:|-----------:|---------------:|----------------:|---------------------:|----------:|----------:|---------:|----|--------:|----:|
| 1 | 60.0 | 2833 | 773 | 3606 | 47 | 60 | 58.8 | 60.3 | 100% | 11.8s | 11.8s | 0 | 1 | 0 |
| 2 | 60.0 | 2970 | 2961 | 5931 | 49 | 99 | 44.0 | 49.6 | 165% | 27.0s | 41.0s | 0 | 2 | 0 |
| 3 | 60.0 | 6016 | 3077 | 9093 | 100 | 151 | 50.9 | 54.4 | 252% | 20.6s | 23.6s | 0 | 3 | 0 |
| 4 | 60.0 | 5941 | 4563 | 10504 | 99 | 175 | 45.2 | 49.2 | 292% | 27.4s | 41.3s | 0 | 4 | 0 |
| 5 | 60.0 | 5980 | 5630 | 11610 | 100 | 193 | 40.9 | 45.0 | 322% | 30.9s | 38.5s | 0 | 5 | 0 |
| 6 | 60.0 | 6136 | 6584 | 12720 | 102 | 212 | 31.5 | 41.4 | 353% | 27.5s | 29.8s | 0 | 6 | 0 |
| 7 | 60.0 | 6042 | 7552 | 13594 | 101 | 226 | 33.9 | 38.3 | 377% | 35.0s | 53.5s | 0 | 7 | 0 |
| 8 | 60.0 | 4780 | 8892 | 13672 | 80 | 228 | 29.8 | 34.6 | 380% | 40.1s | 54.6s | 0 | 8 | 0 |
| 9 | 60.0 | 3976 | 11047 | 15023 | 66 | 250 | 25.9 | 34.1 | 417% | 43.2s | 56.6s | 0 | 9 | 0 |
| 10 | 60.0 | 3967 | 11271 | 15238 | 66 | 254 | 25.4 | 32.9 | 423% | 44.6s | 55.1s | 0 | 10 | 0 |
| 11 | 60.0 | 4106 | 12210 | 16316 | 68 | 272 | 25.0 | 31.6 | 453% | 45.1s | 59.5s | 0 | 11 | 0 |
| 12 | 60.0 | 4092 | 13222 | 17314 | 68 | 288 | 17.8 | 31.3 | 480% | 41.1s | 54.3s | 0 | 12 | 0 |
| 13 | 60.0 | 3839 | 13846 | 17685 | 64 | 295 | 21.5 | 30.4 | 492% | 46.4s | 58.4s | 0 | 13 | 0 |
| 14 | 60.0 | 4724 | 13271 | 17995 | 79 | 300 | 21.8 | 29.4 | 500% | 44.7s | 55.6s | 0 | 14 | 0 |
| 15 | 60.0 | 3382 | 14967 | 18349 | 56 | 306 | 16.2 | 28.4 | 510% | 46.1s | 53.8s | 0 | 15 | 0 |
| 16 | 60.0 | 3269 | 16194 | 19463 | 54 | 324 | 18.5 | 27.4 | 540% | 49.0s | 59.6s | 0 | 16 | 0 |
| 17 | 60.0 | 2563 | 15523 | 18086 | 43 | 301 | 10.7 | 24.6 | 502% | 46.1s | 55.6s | 0 | 17 | 0 |
| 18 | 60.0 | 1373 | 17309 | 18682 | 23 | 311 | 9.6 | 23.6 | 518% | 52.0s | 59.1s | 0 | 18 | 0 |
| 19 | 60.0 | 2286 | 16817 | 19103 | 38 | 318 | 11.7 | 22.9 | 530% | 49.8s | 58.8s | 0 | 19 | 0 |
| 20 | 60.0 | 1700 | 17376 | 19076 | 28 | 318 | 10.0 | 22.6 | 530% | 51.6s | 59.1s | 0 | 20 | 0 |
| 21 | 60.0 | 1371 | 17363 | 18734 | 23 | 312 | 9.4 | 22.1 | 520% | 52.4s | 59.9s | 0 | 21 | 0 |
| 22 | 60.0 | 1049 | 18469 | 19518 | 17 | 325 | 7.9 | 21.4 | 542% | 54.0s | 59.0s | 0 | 22 | 0 |
| 23 | 60.0 | 509 | 18869 | 19378 | 8 | 323 | 4.4 | 21.1 | 538% | 55.7s | 60.0s | 0 | 23 | 0 |
| 24 | 60.0 | 993 | 19396 | 20389 | 17 | 340 | 5.0 | 20.6 | 567% | 51.6s | 57.2s | 0 | 24 | 0 |

## Key findings

- **The worst latency profile of any model tested**: TTFT is **11.8s solo and 40–56s under load**, with max values up to 60.0s — at high concurrency agents often wait the *entire* 60s window before producing a single token. The 4b params + Q4_0 KV + thinking prefill makes every request start extremely slowly.
- **Very slow absolute speed**: only **60 tok/s solo** (47 visible), and combined all-token throughput peaks at just **340 tok/s at 24 agents (5.7x)** — near the bottom of the field.
- **Reasoning dominates almost everything at high concurrency**: visible content collapses to 8–17 tok/s at 18–24 agents (~95% of output is hidden `reasoning_content`).
- **Errors**: 0 across all 24 runs — but as a latency-sensitive or content-heavy workload, this model is a poor fit.

## Charts

Generated with `benchmark/charts.py` (matplotlib), saved as PNG in `results/qwen3-4b/charts/`:

| Chart | File | What it highlights |
|---|---|---|
| Dashboard (2x2 overview) | `dashboard_1_24.png` | Whole sweep in one view |
| Combined throughput | `combined_throughput.png` | Content vs content+reasoning curves |
| Per-agent throughput | `per_agent_throughput.png` | Per-agent slowdown under contention |
| Scaling efficiency | `scaling_efficiency.png` | 100% → 567% |
| Time to first token | `time_to_first_token.png` | 12s solo → ~56s — worst of any model |
| Total tokens generated | `total_tokens_generated.png` | Content vs reasoning stack per concurrency |
| Outcome breakdown | `outcome_breakdown.png` | Zero errors |
| Combined vs per-agent | `combined_vs_per_agent.png` | The concurrency trade-off on one chart |
| Content vs reasoning | `reasoning_vs_content.png` | ~95% reasoning at high concurrency |
