# minicpm5-1b-claude-opus-fable5-v2-thinking_x2 Concurrency Benchmark (2 parallel instances)

Dual-instance sweep of concurrent agents against LM Studio (`LM Studio's local endpoint`), running **two model instances in parallel**:

- **Instance A**: `minicpm5-1b-claude-opus-fable5-v2-thinking`
- **Instance B**: `minicpm5-1b-claude-opus-fable5-v2-thinking-2`

- **Harness**: `benchmark/benchmark.js` (per instance, streaming multi-turn loop) driven by `benchmark/sweep_x2.js` which launches both instances concurrently at each level and merges the results.
- **Config per run**: 60s hard timeout, 12000-token total budget per agent, `--output-cap 12000`, 32768-token context, temperature 0.9.
- **Method**: at each level, **N agents run on instance A and N on instance B simultaneously** → total concurrency 2..48 (even levels). Results in `results/minicpm5-1b-claude-opus-fable5-v2-thinking_x2/<N>/` (with `a/` and `b/` per-instance subdirs + merged `results.json`), aggregate in `sweep-summary.{csv,json}` in the same folder.
- **Note**: this test runs **2 models in LM Studio in parallel** (two API handler slots sharing the GPU) — unlike all single-model tests.
- **Reasoning tokens**: counted as before — this is a thinking model; ~25–55% of output is hidden `reasoning_content` (share climbs with load).

## Results (total concurrency 2–48)

| agents | wall(s) | content_tok | reason_tok | total_tok | comb_tok/s | comb_all_tok/s | per-agent_tok/s | per-agent_all_tok/s | scale_eff | TTFT_mean | TTFT_max | ok | timeout | err |
|-------:|--------:|------------:|-----------:|----------:|-----------:|---------------:|----------------:|---------------------:|----------:|----------:|---------:|----|--------:|----:|
| 2 | 60.0 | 5957 | 10003 | 15960 | 99 | 266 | 117.9 | 311.0 | 100% | 2.4s | 2.9s | 0 | 2 | 0 |
| 4 | 60.0 | 18640 | 10849 | 29489 | 310 | 491 | 148.9 | 180.6 | 185% | 13.7s | 40.8s | 0 | 4 | 0 |
| 6 | 60.0 | 18742 | 16109 | 34851 | 312 | 580 | 114.5 | 175.4 | 218% | 18.3s | 45.0s | 0 | 6 | 0 |
| 8 | 60.0 | 19367 | 19022 | 38389 | 323 | 639 | 69.4 | 114.0 | 240% | 15.5s | 53.0s | 0 | 8 | 0 |
| 10 | 60.0 | 27462 | 12396 | 39858 | 457 | 664 | 72.5 | 103.9 | 250% | 9.4s | 18.1s | 0 | 10 | 0 |
| 12 | 60.0 | 26050 | 18731 | 44781 | 434 | 746 | 80.6 | 108.9 | 280% | 12.3s | 59.9s | 0 | 12 | 0 |
| 14 | 60.0 | 26611 | 17207 | 43818 | 443 | 730 | 67.2 | 94.7 | 274% | 21.2s | 43.5s | 0 | 14 | 0 |
| 16 | 60.0 | 30262 | 11768 | 42030 | 504 | 700 | 66.2 | 79.3 | 263% | 23.4s | 33.0s | 0 | 16 | 0 |
| 18 | 60.0 | 30662 | 16506 | 47168 | 511 | 786 | 65.9 | 79.5 | 295% | 24.3s | 54.7s | 0 | 18 | 0 |
| 20 | 60.0 | 36144 | 14200 | 50344 | 602 | 838 | 60.1 | 74.4 | 315% | 23.3s | 32.2s | 0 | 20 | 0 |
| 22 | 60.0 | 35373 | 15247 | 50620 | 589 | 843 | 57.6 | 70.5 | 317% | 26.9s | 38.2s | 0 | 22 | 0 |
| 24 | 60.0 | 34024 | 15861 | 49885 | 567 | 831 | 53.3 | 66.3 | 312% | 29.5s | 41.4s | 0 | 24 | 0 |
| 26 | 60.0 | 38187 | 15880 | 54067 | 636 | 900 | 54.3 | 64.4 | 338% | 27.8s | 45.9s | 0 | 26 | 0 |
| 28 | 60.0 | 37126 | 16458 | 53584 | 618 | 892 | 54.0 | 64.8 | 335% | 31.9s | 52.2s | 0 | 28 | 0 |
| 30 | 60.0 | 42308 | 17074 | 59382 | 705 | 989 | 51.0 | 60.5 | 372% | 30.0s | 44.4s | 0 | 30 | 0 |
| 32 | 60.0 | 34699 | 22610 | 57309 | 578 | 954 | 42.0 | 52.4 | 359% | 32.6s | 57.9s | 0 | 32 | 0 |
| 34 | 60.0 | 40074 | 19247 | 59321 | 667 | 988 | 43.6 | 53.9 | 371% | 32.0s | 39.5s | 0 | 34 | 0 |
| 36 | 60.0 | 43907 | 16403 | 60310 | 731 | 1004 | 56.8 | 58.9 | 377% | 32.9s | 48.3s | 0 | 36 | 0 |
| 38 | 60.0 | 37917 | 20791 | 58708 | 631 | 978 | 40.2 | 50.0 | 368% | 33.4s | 45.8s | 0 | 38 | 0 |
| 40 | 60.0 | 48223 | 45189 | 93412 | 803 | 1556 | 43.1 | 64.8 | 585% | 17.4s | 56.9s | 0 | 40 | 0 |
| 42 | 60.0 | 43992 | 35332 | 79324 | 733 | 1321 | 40.6 | 58.4 | 497% | 22.3s | 57.3s | 0 | 42 | 0 |
| 44 | 60.1 | 49306 | 52456 | 101762 | 821 | 1695 | 34.4 | 54.5 | 637% | 14.6s | 48.6s | 0 | 44 | 0 |
| 46 | 60.0 | 41365 | 46836 | 88201 | 689 | 1469 | 29.8 | 48.1 | 552% | 22.7s | 52.6s | 0 | 46 | 0 |
| 48 | 60.1 | 50144 | 51793 | 101937 | 835 | 1698 | 38.5 | 55.0 | 638% | 14.3s | 50.6s | 0 | 48 | 0 |

## Key findings

- **Two instances roughly double throughput again**: combined all-token output reaches **1698 tok/s at 48 total agents (6.4x the 2-agent baseline)** — vs the single-instance minicpm peak of ~983 tok/s. Visible content throughput also tops out at 835 tok/s at 48, well above the single-instance ~537.
- **Late-game jump**: throughput holds ~830–1000 between 20–38 agents, then jumps to **1321–1698 at 40–48** — the two slots' batching kicks in hard at the top end.
- **Reasoning share climbs with load**: ~63% of tokens at total-2, settling to ~25–55% in the mid-range, then back up to ~48–52% at 44–48.
- **TTFT is poor (thinking model)**: 2.4s at total-2, mostly 14–33s under load with max values up to 59.9s. Latency-critical use remains a bad fit.
- **Errors**: 0 across all 24 levels.

## Compare: single vs dual instance (comb_all_tok/s)

| total agents | single-instance | dual-instance (x2) |
|-------------:|----------------:|-------------------:|
| 2 | 227 (all-token) | 266 |
| 24 | 983 (peak, all-token) | 831 |
| 40 | ~980 | 1556 |
| 48 | ~980 | **1698** |

## Charts

Generated with `benchmark/charts.py` (matplotlib), saved as PNG in `results/minicpm5-1b-claude-opus-fable5-v2-thinking_x2/charts/`:

| Chart | File | What it highlights |
|---|---|---|
| Dashboard (2x2 overview) | `dashboard_1_24.png` | Whole sweep in one view |
| Combined throughput | `combined_throughput.png` | 266 → 1698 tok/s — 2x single-instance |
| Per-agent throughput | `per_agent_throughput.png` | Per-agent slowdown under contention |
| Scaling efficiency | `scaling_efficiency.png` | 100% → 638% peak |
| Time to first token | `time_to_first_token.png` | 2.4s → ~33s thinking latency |
| Total tokens generated | `total_tokens_generated.png` | ~16k → 102k tokens/60s (content+reasoning) |
| Outcome breakdown | `outcome_breakdown.png` | Zero errors |
| Combined vs per-agent | `combined_vs_per_agent.png` | The concurrency trade-off on one chart |
| Content vs reasoning | `reasoning_vs_content.png` | Hidden reasoning share per level |
