# qwen2.5-1.5b Concurrency Benchmark

Full sweep of concurrent agents against LM Studio (`LM Studio's local endpoint`, model `qwen2.5-1.5b`).

- **Harness**: `benchmark/benchmark.js` (agent loop, streaming, token-level timing), `benchmark/sweep.js` (sweep driver)
- **Config per run**: 60s hard timeout per agent, 12000-token budget per agent, temperature 0.9, agents run a multi-turn loop until token budget or timeout
- **Method**: 24 sequential runs at concurrency 1..24; results in `results/qwen2.5-1.5b/<N>/`, aggregate in `sweep-summary.{csv,json}` in the same folder
- **Reasoning tokens**: 0 across every run — qwen2.5-1.5b is a clean non-thinking model; all output is visible content.

## Results (1–24 agents)

| agents | wall(s) | content_tok | reason_tok | total_tok | comb_tok/s | comb_all_tok/s | per-agent_tok/s | per-agent_all_tok/s | scale_eff | TTFT_mean | TTFT_max | ok | timeout | err |
|-------:|--------:|------------:|-----------:|----------:|-----------:|---------------:|----------------:|---------------------:|----------:|----------:|---------:|----|--------:|----:|
| 1 | 60.0 | 10404 | 0 | 10404 | 173 | 173 | 173.7 | 173.7 | 100% | 108ms | 108ms | 0 | 1 | 0 |
| 2 | 60.0 | 17603 | 0 | 17603 | 293 | 293 | 222.6 | 223.1 | 169% | 108ms | 111ms | 0 | 2 | 0 |
| 3 | 60.0 | 19106 | 0 | 19106 | 318 | 318 | 2947.5* | 2959.4* | 184% | 11.0s | 11.0s | 0 | 3 | 0 |
| 4 | 60.0 | 23185 | 0 | 23185 | 386 | 386 | 201.0 | 201.8 | 223% | 10.8s | 10.8s | 0 | 4 | 0 |
| 5 | 60.0 | 23088 | 0 | 23088 | 385 | 385 | 2756.1* | 2756.7* | 223% | 12.9s | 13.0s | 0 | 5 | 0 |
| 6 | 60.0 | 25347 | 0 | 25347 | 422 | 422 | 2503.8* | 2504.2* | 244% | 14.6s | 14.6s | 0 | 6 | 0 |
| 7 | 60.0 | 23636 | 0 | 23636 | 394 | 394 | 1205.5* | 1205.5* | 228% | 16.1s | 16.1s | 0 | 7 | 0 |
| 8 | 60.0 | 25334 | 0 | 25334 | 422 | 422 | 498.2* | 499.4* | 244% | 13.7s | 13.8s | 0 | 8 | 0 |
| 9 | 60.0 | 29086 | 0 | 29086 | 484 | 484 | 73.0 | 73.7 | 280% | 15.4s | 15.4s | 0 | 9 | 0 |
| 10 | 60.0 | 32245 | 0 | 32245 | 537 | 537 | 58.2 | 65.5 | 310% | 427ms | 457ms | 1 | 9 | 0 |
| 11 | 60.0 | 31875 | 0 | 31875 | 531 | 531 | 53.7 | 60.8 | 307% | 729ms | 748ms | 1 | 10 | 0 |
| 12 | 60.0 | 31067 | 0 | 31067 | 517 | 517 | 48.6 | 57.5 | 299% | 523ms | 567ms | 4 | 8 | 0 |
| 13 | 60.1 | 30341 | 0 | 30341 | 505 | 505 | 45.2 | 55.1 | 292% | 711ms | 761ms | 2 | 11 | 0 |
| 14 | 60.1 | 31634 | 0 | 31634 | 527 | 527 | 45.6 | 55.2 | 305% | 794ms | 891ms | 4 | 10 | 0 |
| 15 | 60.0 | 31487 | 0 | 31487 | 524 | 524 | 40.1 | 48.0 | 303% | 1.3s | 1.4s | 4 | 11 | 0 |
| 16 | 60.0 | 34284 | 0 | 34284 | 571 | 571 | 46.1 | 57.1 | 330% | 675ms | 725ms | 5 | 11 | 0 |
| 17 | 60.0 | 31246 | 0 | 31246 | 520 | 520 | 39.0 | 48.2 | 301% | 3.2s | 3.2s | 5 | 12 | 0 |
| 18 | 60.1 | 33595 | 0 | 33595 | 559 | 559 | 40.0 | 46.5 | 323% | 2.3s | 2.4s | 7 | 11 | 0 |
| 19 | 60.0 | 32812 | 0 | 32812 | 546 | 546 | 34.7 | 43.3 | 316% | 1.6s | 1.8s | 5 | 14 | 0 |
| 20 | 60.1 | 34622 | 0 | 34622 | 577 | 577 | 38.4 | 46.1 | 334% | 626ms | 716ms | 7 | 13 | 0 |
| 21 | 60.1 | 34535 | 0 | 34535 | 575 | 575 | 34.6 | 42.4 | 332% | 3.5s | 3.6s | 6 | 15 | 0 |
| 22 | 60.1 | 30720 | 0 | 30720 | 511 | 511 | 29.1 | 37.2 | 295% | 432ms | 631ms | 5 | 17 | 0 |
| 23 | 60.1 | 33437 | 0 | 33437 | 557 | 557 | 30.9 | 39.1 | 322% | 1.9s | 2.0s | 7 | 16 | 0 |
| 24 | 60.1 | 33615 | 0 | 33615 | 560 | 560 | 29.6 | 37.1 | 324% | 1.2s | 1.5s | 8 | 16 | 0 |

*\* per-agent tok/s figures marked with * are burst-timing artifacts: when agents sit in the queue for ~11–16s then generate in a short burst, the per-agent metric is inflated. Trust the combined throughput and TTFT columns.*

## Key findings

- **Clean, all-content output**: 0 reasoning tokens everywhere — every token is visible content.
- **Good scaling**: combined throughput 173 tok/s solo → **577 tok/s at 20 agents (3.3x)**, plateauing around 520–577 from 10–24 agents.
- **Best latency of the lot at low concurrency**: TTFT just 108ms at 1–2 agents. It degrades in the mid-range (11–16s at 3–9 agents — likely server batching transitions) then recovers to sub-second at 10+ agents. This erratic profile mirrors what we saw with the other models.
- **Per-agent speed**: ~30–46 tok/s at high concurrency (vs 173 solo) — the usual slot-contention erosion.
- **Errors**: 0 across all 24 runs.

## Charts

Generated with `benchmark/charts.py` (matplotlib), saved as PNG in `results/qwen2.5-1.5b/charts/`:

| Chart | File | What it highlights |
|---|---|---|
| Dashboard (2x2 overview) | `dashboard_1_24.png` | Whole sweep in one view |
| Combined throughput | `combined_throughput.png` | 173 → 577 tok/s scaling, plateau at 10+ |
| Per-agent throughput | `per_agent_throughput.png` | Per-agent slowdown (burst artifact noted) |
| Scaling efficiency | `scaling_efficiency.png` | 100% → 334% |
| Time to first token | `time_to_first_token.png` | 108ms solo, mid-range spikes, then recovery |
| Total tokens generated | `total_tokens_generated.png` | ~10k → 34k tokens/60s |
| Outcome breakdown | `outcome_breakdown.png` | Zero errors |
| Combined vs per-agent | `combined_vs_per_agent.png` | The concurrency trade-off on one chart |
