# mistralai/ministral-3-3b-instruct-2512 Concurrency Benchmark

Full sweep of concurrent agents against LM Studio (`LM Studio's local endpoint`, model `mistralai/ministral-3-3b-instruct-2512`).

- **Harness**: `benchmark/benchmark.js` (agent loop, streaming, token-level timing), `benchmark/sweep.js` (sweep driver)
- **Config per run**: standard settings — 60s hard timeout, 12000-token total budget, `--output-cap 12000`, 32768-token context, temperature 0.9, agents run a multi-turn loop
- **Quants note**: this run is served with **K and V cache quants of Q4_0**, matching the `unsloth/ministral-3-3b-instruct-2512` run so the two builds are directly comparable.
- **Method**: 24 sequential runs at concurrency 1..24; results in `results/mistralai-ministral-3-3b-instruct-2512/<N>/`, aggregate in `sweep-summary.{csv,json}` in the same folder
- **Reasoning tokens**: 0 across every run — a clean non-thinking model; all output is visible content.

## Results (1–24 agents)

| agents | wall(s) | content_tok | reason_tok | total_tok | comb_tok/s | comb_all_tok/s | per-agent_tok/s | per-agent_all_tok/s | scale_eff | TTFT_mean | TTFT_max | ok | timeout | err |
|-------:|--------:|------------:|-----------:|----------:|-----------:|---------------:|----------------:|---------------------:|----------:|----------:|---------:|----|--------:|----:|
| 1 | 60.0 | 4248 | 0 | 4248 | 71 | 71 | 174.4 | 174.4 | 100% | 92ms | 92ms | 0 | 1 | 0 |
| 2 | 60.0 | 7026 | 0 | 7026 | 117 | 117 | 116.1 | 116.1 | 165% | 112ms | 125ms | 0 | 2 | 0 |
| 3 | 60.0 | 9730 | 0 | 9730 | 162 | 162 | 77.0 | 77.5 | 228% | 3.5s | 3.5s | 0 | 3 | 0 |
| 4 | 60.0 | 12100 | 0 | 12100 | 202 | 202 | 72.5 | 72.6 | 285% | 3.9s | 3.9s | 0 | 4 | 0 |
| 5 | 60.0 | 13500 | 0 | 13500 | 225 | 225 | 51.1 | 51.1 | 317% | 6.0s | 6.1s | 0 | 5 | 0 |
| 6 | 60.0 | 14961 | 0 | 14961 | 249 | 249 | 48.4 | 48.4 | 351% | 7.0s | 7.0s | 0 | 6 | 0 |
| 7 | 60.0 | 16443 | 0 | 16443 | 274 | 274 | 46.7 | 46.7 | 386% | 7.6s | 7.6s | 0 | 7 | 0 |
| 8 | 60.0 | 17130 | 0 | 17130 | 285 | 285 | 41.8 | 41.8 | 401% | 8.7s | 8.7s | 0 | 8 | 0 |
| 9 | 60.0 | 18175 | 0 | 18175 | 303 | 303 | 40.6 | 40.6 | 427% | 9.2s | 9.2s | 0 | 9 | 0 |
| 10 | 60.0 | 19413 | 0 | 19413 | 323 | 323 | 38.6 | 38.6 | 455% | 9.7s | 9.7s | 0 | 10 | 0 |
| 11 | 60.0 | 19277 | 0 | 19277 | 321 | 321 | 38.4 | 38.4 | 452% | 11.4s | 11.5s | 0 | 11 | 0 |
| 12 | 60.0 | 21114 | 0 | 21114 | 352 | 352 | 37.3 | 37.3 | 496% | 11.6s | 11.7s | 0 | 12 | 0 |
| 13 | 60.0 | 21976 | 0 | 21976 | 366 | 366 | 36.1 | 36.1 | 515% | 12.1s | 12.1s | 0 | 13 | 0 |
| 14 | 60.0 | 22169 | 0 | 22169 | 369 | 369 | 33.6 | 33.6 | 520% | 12.6s | 12.6s | 0 | 14 | 0 |
| 15 | 60.0 | 23092 | 0 | 23092 | 385 | 385 | 32.6 | 32.6 | 542% | 12.8s | 12.9s | 0 | 15 | 0 |
| 16 | 60.0 | 23822 | 0 | 23822 | 397 | 397 | 32.0 | 32.0 | 559% | 13.4s | 13.5s | 0 | 16 | 0 |
| 17 | 60.1 | 22402 | 0 | 22402 | 373 | 373 | 28.7 | 28.7 | 525% | 14.0s | 14.1s | 0 | 17 | 0 |
| 18 | 60.0 | 23572 | 0 | 23572 | 393 | 393 | 27.5 | 27.5 | 554% | 12.4s | 12.5s | 0 | 18 | 0 |
| 19 | 60.0 | 23481 | 0 | 23481 | 391 | 391 | 26.5 | 26.5 | 551% | 13.3s | 13.4s | 0 | 19 | 0 |
| 20 | 60.0 | 23724 | 0 | 23724 | 395 | 395 | 26.2 | 26.2 | 556% | 14.7s | 14.8s | 0 | 20 | 0 |
| 21 | 60.0 | 23393 | 0 | 23393 | 390 | 390 | 25.7 | 25.7 | 549% | 16.6s | 16.8s | 0 | 21 | 0 |
| 22 | 60.0 | 24238 | 0 | 24238 | 404 | 404 | 24.8 | 24.8 | 569% | 15.6s | 15.7s | 0 | 22 | 0 |
| 23 | 60.0 | 23529 | 0 | 23529 | 392 | 392 | 23.4 | 23.4 | 552% | 16.3s | 16.5s | 0 | 23 | 0 |
| 24 | 60.0 | 24026 | 0 | 24026 | 400 | 400 | 23.1 | 23.1 | 563% | 16.5s | 16.8s | 0 | 24 | 0 |

## Key findings

- **Clean, all-content output**: 0 reasoning tokens everywhere.
- **Much better behaved than the unsloth build at the same KV Q4_0**: solo throughput is **71 tok/s vs 47** for `unsloth/ministral-3-3b-instruct-2512` (same 4-bit KV), and the scaling curve is smooth and monotonic (71 → 400 tok/s) instead of unsloth's erratic 106–365 swings.
- **Peak 404 tok/s at 22 agents (5.7x)**; essentially flat at ~390–400 from 16–24. Near-identical absolute peak to the unsloth build (~410), so the two builds converge at high concurrency even though mistralai wins solo.
- **TTFT**: 92ms solo → ~16.5s at 24; smooth growth, no outliers.
- **Per-agent speed**: ~23–25 tok/s at high concurrency (vs ~174 solo burst metric).
- **Errors**: 0 across all 24 runs.

## Compare with unsloth/ministral-3-3b-instruct-2512 (both KV Q4_0)

| concurrency | unsloth comb_tok/s | mistralai comb_tok/s |
|------------:|-------------------:|---------------------:|
| 1 | 47 | 71 |
| 8 | 135 | 285 |
| 16 | 375 | 397 |
| 22 | 394 | **404** |
| 24 | 400 | 400 |

The mistralai build is faster at every concurrency below ~24 and far more stable; at the top end the two converge near the server's capacity.

## Charts

Generated with `benchmark/charts.py` (matplotlib), saved as PNG in `results/mistralai-ministral-3-3b-instruct-2512/charts/`:

| Chart | File | What it highlights |
|---|---|---|
| Dashboard (2x2 overview) | `dashboard_1_24.png` | Whole sweep in one view |
| Combined throughput | `combined_throughput.png` | Smooth 71 → 404 tok/s scaling |
| Per-agent throughput | `per_agent_throughput.png` | Per-agent slowdown under contention |
| Scaling efficiency | `scaling_efficiency.png` | 100% → 569% peak |
| Time to first token | `time_to_first_token.png` | 92ms → ~16.5s latency growth |
| Total tokens generated | `total_tokens_generated.png` | ~4k → 24k tokens/60s |
| Outcome breakdown | `outcome_breakdown.png` | Zero errors |
| Combined vs per-agent | `combined_vs_per_agent.png` | The concurrency trade-off on one chart |
