# qwen2.5-0.5b-instruct_x2 Concurrency Benchmark (2 parallel instances)

Dual-instance sweep of concurrent agents against LM Studio (`LM Studio's local endpoint`), running **two model instances in parallel**:

- **Instance A**: `qwen2.5-0.5b-instruct`
- **Instance B**: `qwen2.5-0.5b-instruct-2`

- **Harness**: `benchmark/benchmark.js` (per instance, streaming multi-turn loop) driven by `benchmark/sweep_x2.js` which launches both instances concurrently at each level and merges the results.
- **Config per run**: 60s hard timeout, 12000-token total budget per agent, `--output-cap 12000`, 32768-token context, temperature 0.9.
- **Method**: at each level, **N agents run on instance A and N on instance B simultaneously** → total concurrency 2..48 (even levels). Results in `results/qwen2.5-0.5b-instruct_x2/<N>/` (with `a/` and `b/` per-instance subdirs + merged `results.json`), aggregate in `sweep-summary.{csv,json}` in the same folder.
- **Note**: this test runs **2 models in LM Studio in parallel** (two API handler slots sharing the GPU) — unlike all single-model tests.
- **Reasoning tokens**: 0 across every run — a clean non-thinking model; all output is visible content.

## Results (total concurrency 2–48)

| agents | wall(s) | content_tok | reason_tok | total_tok | comb_tok/s | comb_all_tok/s | per-agent_tok/s | per-agent_all_tok/s | scale_eff | TTFT_mean | TTFT_max | ok | timeout | err |
|-------:|--------:|------------:|-----------:|----------:|-----------:|---------------:|----------------:|---------------------:|----------:|----------:|---------:|----|--------:|----:|
| 2 | 60.0 | 21558 | 0 | 21558 | 359 | 359 | 205.6 | 205.6 | 100% | 7.6s | 11.1s | 0 | 2 | 0 |
| 4 | 60.0 | 34264 | 0 | 34264 | 571 | 571 | 151.1 | 168.8 | 159% | 111ms | 129ms | 0 | 4 | 0 |
| 6 | 60.0 | 50226 | 0 | 50226 | 836 | 836 | 583.9* | 599.6* | 233% | 2.0s | 2.1s | 0 | 6 | 0 |
| 8 | 60.0 | 43370 | 0 | 43370 | 722 | 722 | 225.1* | 229.8* | 201% | 16.2s | 17.3s | 0 | 8 | 0 |
| 10 | 60.0 | 42075 | 0 | 42075 | 701 | 701 | 131.3 | 135.2 | 195% | 15.2s | 15.5s | 0 | 10 | 0 |
| 12 | 60.0 | 46594 | 0 | 46594 | 776 | 776 | 113.2 | 116.4 | 216% | 11.8s | 13.3s | 0 | 12 | 0 |
| 14 | 60.0 | 50885 | 0 | 50885 | 847 | 847 | 131.4 | 133.2 | 236% | 13.7s | 14.9s | 0 | 14 | 0 |
| 16 | 60.0 | 49695 | 0 | 49695 | 828 | 828 | 88.1 | 89.0 | 231% | 16.2s | 18.4s | 0 | 16 | 0 |
| 18 | 60.0 | 53576 | 0 | 53576 | 892 | 892 | 113.4 | 114.4 | 248% | 19.4s | 20.1s | 0 | 18 | 0 |
| 20 | 60.0 | 55550 | 0 | 55550 | 925 | 925 | 95.2 | 96.8 | 258% | 19.7s | 21.0s | 0 | 20 | 0 |
| 22 | 60.0 | 61222 | 0 | 61222 | 1020 | 1020 | 98.5 | 99.5 | 284% | 21.0s | 23.1s | 0 | 22 | 0 |
| 24 | 60.1 | 51204 | 0 | 51204 | 853 | 853 | 88.5 | 89.2 | 238% | 27.2s | 28.2s | 0 | 24 | 0 |
| 26 | 60.1 | 60695 | 0 | 60695 | 1011 | 1011 | 84.9 | 85.5 | 282% | 18.9s | 21.0s | 0 | 26 | 0 |
| 28 | 60.1 | 57712 | 0 | 57712 | 961 | 961 | 85.4 | 86.3 | 268% | 26.1s | 26.7s | 0 | 28 | 0 |
| 30 | 60.1 | 64897 | 0 | 64897 | 1081 | 1081 | 76.9 | 77.4 | 301% | 25.0s | 26.3s | 0 | 30 | 0 |
| 32 | 60.1 | 86878 | 0 | 86878 | 1446 | 1446 | 71.6 | 72.9 | 403% | 10.5s | 15.4s | 0 | 32 | 0 |
| 34 | 60.1 | 95454 | 0 | 95454 | 1589 | 1589 | 75.0 | 76.9 | 443% | 7.8s | 9.5s | 0 | 34 | 0 |
| 36 | 60.1 | 85474 | 0 | 85474 | 1423 | 1423 | 53.9 | 56.0 | 396% | 13.5s | 14.2s | 0 | 36 | 0 |
| 38 | 60.1 | 109008 | 0 | 109008 | 1815 | 1815 | 66.4 | 68.5 | 506% | 3.6s | 3.9s | 0 | 38 | 0 |
| 40 | 60.1 | 95521 | 0 | 95521 | 1590 | 1590 | 66.3 | 68.9 | 443% | 13.9s | 15.9s | 0 | 40 | 0 |
| 42 | 60.1 | 93382 | 0 | 93382 | 1555 | 1555 | 63.2 | 66.4 | 433% | 13.2s | 14.5s | 0 | 42 | 0 |
| 44 | 60.1 | 93846 | 0 | 93846 | 1562 | 1562 | 65.4 | 68.6 | 435% | 13.6s | 15.2s | 0 | 44 | 0 |
| 46 | 60.1 | 84457 | 0 | 84457 | 1406 | 1406 | 55.5 | 58.7 | 392% | 19.0s | 20.1s | 0 | 46 | 0 |
| 48 | 60.1 | 96869 | 0 | 96869 | 1612 | 1612 | 65.3 | 69.3 | 449% | 13.0s | 13.8s | 0 | 48 | 0 |

*\* burst-timing artifacts — see methodology caveat.*

## Key findings

- **Two instances roughly double throughput**: combined output reaches **1815 tok/s at 38 total agents (5.1x the 2-agent baseline)** and ~1400–1610 from 40–48 — roughly 2x the ~900–944 tok/s ceiling of the single-instance run. Running two API-handler slots in parallel works.
- **Steeper scaling than single-instance**: single-instance plateaued at ~267% (relative to its solo); the dual run keeps climbing to 506% at 38 — the extra slot gives real headroom.
- **Solo-baseline caveat**: the 2-agent row (1 per instance) is 359 tok/s, similar to the single-instance 354 — each slot independently sustains its solo speed, and combining them adds up.
- **Noisy mid-range, then a jump at 30–34**: throughput swings 701–1020 between 8–30, then jumps to ~1446–1815 at 32–38 — LM Studio's per-slot batching transitions.
- **TTFT**: 111ms–27s, erratic like the single-instance run; no pathological outliers.
- **All visible content**, 0 reasoning tokens.
- **Errors**: 0 across all 24 levels.

## Compare: single vs dual instance (comb_tok/s)

| total agents | single-instance | dual-instance (x2) |
|-------------:|----------------:|-------------------:|
| 2 | 527 | 359 |
| 8 | 556 | 722 |
| 16 | 936 | 828 |
| 20 | 944 | 925 |
| 32 | ~870 | 1446 |
| 38 | ~870 | **1815** |
| 48 | ~871 | 1612 |

## Charts

Generated with `benchmark/charts.py` (matplotlib), saved as PNG in `results/qwen2.5-0.5b-instruct_x2/charts/`:

| Chart | File | What it highlights |
|---|---|---|
| Dashboard (2x2 overview) | `dashboard_1_24.png` | Whole sweep in one view |
| Combined throughput | `combined_throughput.png` | 359 → 1815 tok/s — 2x single-instance |
| Per-agent throughput | `per_agent_throughput.png` | Per-agent slowdown under contention |
| Scaling efficiency | `scaling_efficiency.png` | 100% → 506% peak |
| Time to first token | `time_to_first_token.png` | 111ms–27s erratic latency |
| Total tokens generated | `total_tokens_generated.png` | ~22k → 109k tokens/60s |
| Outcome breakdown | `outcome_breakdown.png` | Zero errors |
| Combined vs per-agent | `combined_vs_per_agent.png` | The concurrency trade-off on one chart |
