# minicpm5-1b-claude-opus-fable5-v2-thinking Concurrency Benchmark

Full sweep of concurrent agents against LM Studio (`LM Studio's local endpoint`, model `minicpm5-1b-claude-opus-fable5-v2-thinking`).

- **Harness**: `benchmark/benchmark.js` (agent loop, streaming, token-level timing), `benchmark/sweep.js` (sweep driver)
- **Config per run**: 60s hard timeout per agent, 12000-token budget per agent, temperature 0.9, agents run a multi-turn loop until token budget or timeout
- **Method**: 24 sequential runs at concurrency 1..24; results in `results/minicpm5-1b-claude-opus-fable5-v2-thinking/<N>/`, aggregate in `sweep-summary.{csv,json}` in the same folder
- **Reasoning tokens**: this model streams hidden thinking via `reasoning_content`. The harness now counts these separately — `content_tok` is visible output, `reason_tok` is hidden reasoning, `total_tok` is both. `comb_all_tok/s` and `per-agent_all_tok/s` include reasoning.

## Results (1–24 agents)

| agents | wall(s) | content_tok | reason_tok | total_tok | comb_tok/s | comb_all_tok/s | per-agent_tok/s | per-agent_all_tok/s | scale_eff | TTFT_mean | TTFT_max | ok | timeout | err |
|-------:|--------:|------------:|-----------:|----------:|-----------:|---------------:|----------------:|---------------------:|----------:|----------:|---------:|----|--------:|----:|
| 1 | 23.9 | 3049 | 2363 | 5412 | 128 | 227 | 229.6 | 230.5 | 100% | 1.5s | 1.5s | 1 | 0 | 0 |
| 2 | 60.0 | 10658 | 11673 | 22331 | 178 | 372 | 198.4 | 222.1 | 164% | 12.0s | 21.6s | 0 | 2 | 0 |
| 3 | 60.0 | 12869 | 11592 | 24461 | 214 | 407 | 175.0 | 184.0 | 179% | 11.5s | 22.2s | 1 | 2 | 0 |
| 4 | 60.0 | 17141 | 9759 | 26900 | 285 | 448 | 151.6 | 172.2 | 197% | 6.0s | 9.9s | 1 | 3 | 0 |
| 5 | 60.0 | 19867 | 9610 | 29477 | 331 | 491 | 161.0 | 151.1 | 216% | 10.3s | 15.6s | 1 | 4 | 0 |
| 6 | 60.0 | 18674 | 11556 | 30230 | 311 | 503 | 88.0 | 98.6 | 222% | 17.8s | 47.9s | 0 | 6 | 0 |
| 7 | 60.0 | 18497 | 10165 | 28662 | 308 | 477 | 88.9 | 92.9 | 210% | 21.9s | 52.2s | 0 | 7 | 0 |
| 8 | 60.0 | 20539 | 10343 | 30882 | 342 | 514 | 70.6 | 75.6 | 226% | 14.3s | 15.7s | 0 | 8 | 0 |
| 9 | 60.0 | 17577 | 14467 | 32044 | 293 | 534 | 71.9 | 92.2 | 235% | 16.8s | 18.1s | 0 | 9 | 0 |
| 10 | 60.0 | 23646 | 9135 | 32781 | 394 | 546 | 74.4 | 85.8 | 241% | 17.9s | 29.4s | 0 | 10 | 0 |
| 11 | 60.0 | 24071 | 10816 | 34887 | 401 | 581 | 63.8 | 71.2 | 256% | 18.3s | 22.5s | 2 | 9 | 0 |
| 12 | 60.0 | 27078 | 17687 | 44765 | 451 | 746 | 69.4 | 85.7 | 329% | 5.4s | 26.2s | 6 | 6 | 0 |
| 13 | 60.0 | 29505 | 13223 | 42728 | 491 | 712 | 68.6 | 77.5 | 314% | 9.5s | 18.3s | 4 | 9 | 0 |
| 14 | 60.0 | 27163 | 19524 | 46687 | 452 | 778 | 68.5 | 81.4 | 343% | 12.3s | 40.0s | 5 | 9 | 0 |
| 15 | 60.0 | 27313 | 19168 | 46481 | 455 | 774 | 50.0 | 61.7 | 341% | 7.1s | 11.2s | 8 | 7 | 0 |
| 16 | 60.0 | 31414 | 18972 | 50386 | 523 | 839 | 55.8 | 68.7 | 370% | 5.6s | 10.7s | 8 | 8 | 0 |
| 17 | 60.0 | 30167 | 19960 | 50127 | 502 | 835 | 54.1 | 66.9 | 368% | 5.8s | 10.7s | 7 | 10 | 0 |
| 18 | 60.0 | 29237 | 21092 | 50329 | 487 | 838 | 52.5 | 64.9 | 369% | 7.9s | 11.3s | 8 | 10 | 0 |
| 19 | 60.0 | 32259 | 19714 | 51973 | 537 | 866 | 48.4 | 59.0 | 381% | 9.3s | 24.2s | 12 | 7 | 0 |
| 20 | 60.0 | 28395 | 26889 | 55284 | 473 | 921 | 53.2 | 62.4 | 406% | 6.5s | 22.1s | 8 | 12 | 0 |
| 21 | 60.0 | 28847 | 24785 | 53632 | 480 | 893 | 48.0 | 58.1 | 393% | 7.8s | 11.2s | 12 | 9 | 0 |
| 22 | 60.0 | 28178 | 28955 | 57133 | 469 | 952 | 51.5 | 61.0 | 419% | 7.1s | 22.8s | 12 | 10 | 0 |
| 23 | 60.0 | 27084 | 28745 | 55829 | 451 | 930 | 48.6 | 55.1 | 410% | 11.5s | 22.4s | 12 | 11 | 0 |
| 24 | 60.0 | 28180 | 30838 | 59018 | 469 | 983 | 43.2 | 56.7 | 433% | 7.0s | 13.2s | 12 | 12 | 0 |

## Key findings

- **Reasoning dominates**: this "thinking" model spends a large share of its output on hidden reasoning — ~43–52% of all tokens across runs (e.g. 30,838 of 59,018 at 24 agents). Visible content is only about half of what it actually generates.
- **All-in throughput is high**: counting reasoning, combined throughput reaches **983 tok/s at 24 agents (4.3x single-agent)** — well above llama-3.2-1b-mini-agent's 652 tok/s. But visible content throughput peaks much lower (~537 tok/s at 19) because of the reasoning overhead.
- **Single-agent is slow to start and short**: the concurrency-1 run finished in ~24s (agent ended naturally) at 128 content / 227 all tok/s — slower than the non-thinking model's 196 tok/s.
- **TTFT is the weak point**: 1.5s–52s across the sweep (mean usually 6–18s under load) — the reasoning prefill dominates first-token latency. Not suitable for latency-critical controllers at high concurrency.
- **Errors**: 0 across all 24 runs.

## Charts

Generated with `benchmark/charts.py` (matplotlib), saved as PNG in `results/minicpm5-1b-claude-opus-fable5-v2-thinking/charts/`:

| Chart | File | What it highlights |
|---|---|---|
| Dashboard (2x2 overview) | `dashboard_1_24.png` | Whole sweep in one view |
| Combined throughput | `combined_throughput.png` | Content vs content+reasoning curves |
| Per-agent throughput | `per_agent_throughput.png` | Per-agent slowdown under contention |
| Scaling efficiency | `scaling_efficiency.png` | 100% → 433% total, per-agent share collapses |
| Time to first token | `time_to_first_token.png` | 1.5s–52s reasoning-prefill latency |
| Total tokens generated | `total_tokens_generated.png` | Content vs reasoning stack per concurrency |
| Outcome breakdown | `outcome_breakdown.png` | Zero errors; early natural stops at high concurrency |
| Combined vs per-agent | `combined_vs_per_agent.png` | The concurrency trade-off on one chart |
| Content vs reasoning | `reasoning_vs_content.png` | Hidden reasoning budget share per run |
