# qwen2.5-0.5b-instruct Concurrency Benchmark

Full sweep of concurrent agents against LM Studio (`LM Studio's local endpoint`, model `qwen2.5-0.5b-instruct`).

- **Harness**: `benchmark/benchmark.js` (agent loop, streaming, token-level timing), `benchmark/sweep.js` (sweep driver)
- **Config per run**: standard settings — 60s hard timeout, 12000-token total budget, `--output-cap 12000`, 32768-token context, temperature 0.9, agents run a multi-turn loop
- **Quants note**: this run is served with **K and V cache at F16** (full precision — no KV quantization).
- **Method**: 24 sequential runs at concurrency 1..24; results in `results/qwen2.5-0.5b-instruct/<N>/`, aggregate in `sweep-summary.{csv,json}` in the same folder
- **Reasoning tokens**: 0 across every run — a clean non-thinking model; all output is visible content.

## Results (1–24 agents)

| agents | wall(s) | content_tok | reason_tok | total_tok | comb_tok/s | comb_all_tok/s | per-agent_tok/s | per-agent_all_tok/s | scale_eff | TTFT_mean | TTFT_max | ok | timeout | err |
|-------:|--------:|------------:|-----------:|----------:|-----------:|---------------:|----------------:|---------------------:|----------:|----------:|---------:|----|--------:|----:|
| 1 | 33.9 | 12000 | 0 | 12000 | 354 | 354 | 354.8 | 354.8 | 100% | 55ms | 55ms | 1 | 0 | 0 |
| 2 | 45.6 | 24000 | 0 | 24000 | 527 | 527 | 264.1 | 268.0 | 149% | 76ms | 84ms | 2 | 0 | 0 |
| 3 | 60.0 | 30077 | 0 | 30077 | 501 | 501 | 472.2* | 489.6* | 142% | 9.7s | 9.7s | 0 | 3 | 0 |
| 4 | 60.0 | 32129 | 0 | 32129 | 535 | 535 | 166.9 | 178.3 | 151% | 8.9s | 8.9s | 0 | 4 | 0 |
| 5 | 60.0 | 32476 | 0 | 32476 | 541 | 541 | 161.4 | 169.9 | 153% | 6.4s | 6.4s | 0 | 5 | 0 |
| 6 | 60.0 | 32443 | 0 | 32443 | 540 | 540 | 134.0 | 139.0 | 153% | 8.8s | 8.8s | 0 | 6 | 0 |
| 7 | 60.0 | 32882 | 0 | 32882 | 548 | 548 | 199.1 | 201.8 | 155% | 11.0s | 11.0s | 0 | 7 | 0 |
| 8 | 60.0 | 33382 | 0 | 33382 | 556 | 556 | 85.9 | 89.2 | 157% | 696ms | 723ms | 0 | 8 | 0 |
| 9 | 60.1 | 34893 | 0 | 34893 | 581 | 581 | 146.1 | 148.6 | 164% | 9.9s | 10.0s | 0 | 9 | 0 |
| 10 | 60.1 | 36567 | 0 | 36567 | 609 | 609 | 117.4 | 119.0 | 172% | 10.4s | 10.4s | 0 | 10 | 0 |
| 11 | 60.0 | 37939 | 0 | 37939 | 632 | 632 | 100.9 | 105.3 | 179% | 10.7s | 10.7s | 0 | 11 | 0 |
| 12 | 60.1 | 38531 | 0 | 38531 | 642 | 642 | 73.6 | 76.4 | 181% | 5.7s | 5.8s | 0 | 12 | 0 |
| 13 | 60.0 | 37386 | 0 | 37386 | 623 | 623 | 73.5 | 74.5 | 176% | 14.4s | 14.5s | 0 | 13 | 0 |
| 14 | 60.0 | 38049 | 0 | 38049 | 634 | 634 | 75.9 | 76.8 | 179% | 12.6s | 12.6s | 0 | 14 | 0 |
| 15 | 60.1 | 43862 | 0 | 43862 | 730 | 730 | 68.1 | 70.7 | 206% | 13.9s | 14.0s | 0 | 15 | 0 |
| 16 | 60.1 | 56214 | 0 | 56214 | 936 | 936 | 69.9 | 71.9 | 264% | 2.0s | 2.0s | 0 | 16 | 0 |
| 17 | 60.1 | 54694 | 0 | 54694 | 911 | 911 | 81.9 | 85.0 | 257% | 5.2s | 5.2s | 0 | 17 | 0 |
| 18 | 60.1 | 51477 | 0 | 51477 | 857 | 857 | 79.2 | 82.1 | 242% | 9.9s | 9.9s | 0 | 18 | 0 |
| 19 | 60.1 | 51674 | 0 | 51674 | 860 | 860 | 60.2 | 61.9 | 243% | 9.4s | 9.5s | 0 | 19 | 0 |
| 20 | 60.1 | 56706 | 0 | 56706 | 944 | 944 | 71.7 | 74.2 | 267% | 4.4s | 4.5s | 0 | 20 | 0 |
| 21 | 60.1 | 54171 | 0 | 54171 | 902 | 902 | 70.0 | 72.4 | 255% | 10.1s | 10.1s | 0 | 21 | 0 |
| 22 | 60.1 | 51814 | 0 | 51814 | 863 | 863 | 64.6 | 67.4 | 244% | 8.2s | 8.3s | 0 | 22 | 0 |
| 23 | 60.1 | 52938 | 0 | 52938 | 881 | 881 | 74.0 | 77.3 | 249% | 11.0s | 11.1s | 0 | 23 | 0 |
| 24 | 60.1 | 52334 | 0 | 52334 | 871 | 871 | 73.4 | 77.0 | 246% | 12.3s | 12.3s | 0 | 24 | 0 |

*\* burst-timing artifacts — see methodology caveat.*

## Key findings

- **Fastest absolute throughput of any model tested**: **944 tok/s at 20 agents** (and 936 at 16, 871 at 24) — well ahead of the previous best (falcon3 839, qwen-deepseek 829, minicpm 983* measured all-token). The tiny 0.5b model plus **F16 KV** makes it a throughput monster: single-agent is **354 tok/s** (also the fastest solo of the whole set).
- **Solo run finished early on budget**: concurrency 1 hit the 12k budget in 33.9s, concurrency 2 in 45.6s — the only runs that didn't use the full 60s, because the model out-runs the budget.
- **Low relative scaling (100% → 267%)** — but that's a good thing here: the model is already near capacity at 2 agents (527 tok/s), so concurrency adds little relative headroom. The mid-range (3–14) grows slowly, then a jump at 15–20.
- **Excellent latency**: TTFT 55ms solo, mostly 2–14s under load with no pathological spikes.
- **All visible content**, 0 reasoning tokens.
- **Errors**: 0 across all 24 runs.

## Charts

Generated with `benchmark/charts.py` (matplotlib), saved as PNG in `results/qwen2.5-0.5b-instruct/charts/`:

| Chart | File | What it highlights |
|---|---|---|
| Dashboard (2x2 overview) | `dashboard_1_24.png` | Whole sweep in one view |
| Combined throughput | `combined_throughput.png` | 354 → 944 tok/s — fastest of all models |
| Per-agent throughput | `per_agent_throughput.png` | Per-agent slowdown under contention |
| Scaling efficiency | `scaling_efficiency.png` | 100% → 267% |
| Time to first token | `time_to_first_token.png` | 55ms solo, mostly 2–14s under load |
| Total tokens generated | `total_tokens_generated.png` | ~32k–57k tokens/60s |
| Outcome breakdown | `outcome_breakdown.png` | Zero errors |
| Combined vs per-agent | `combined_vs_per_agent.png` | The concurrency trade-off on one chart |
