# unsloth/ministral-3-3b-instruct-2512 Concurrency Benchmark

Full sweep of concurrent agents against LM Studio (`LM Studio's local endpoint`, model `unsloth/ministral-3-3b-instruct-2512`).

- **Harness**: `benchmark/benchmark.js` (agent loop, streaming, token-level timing), `benchmark/sweep.js` (sweep driver)
- **Config per run**: standard settings — 60s hard timeout, 12000-token total budget, `--output-cap 12000`, 32768-token context, temperature 0.9, agents run a multi-turn loop
- **Quants note**: this run is served with **K and V cache quants of Q4_0** — the KV cache (attention key/value) is quantized to 4-bit Q4_0, which trades a little accuracy for lower memory/bandwidth. Single-agent speed here is the lowest of any model tested, which is consistent with Q4_0 KV quants on this architecture being slow rather than a speedup.
- **Method**: 24 sequential runs at concurrency 1..24; results in `results/unsloth-ministral-3-3b-instruct-2512/<N>/`, aggregate in `sweep-summary.{csv,json}` in the same folder
- **Reasoning tokens**: 0 across every run — a clean non-thinking model; all output is visible content.

## Results (1–24 agents)

| agents | wall(s) | content_tok | reason_tok | total_tok | comb_tok/s | comb_all_tok/s | per-agent_tok/s | per-agent_all_tok/s | scale_eff | TTFT_mean | TTFT_max | ok | timeout | err |
|-------:|--------:|------------:|-----------:|----------:|-----------:|---------------:|----------------:|---------------------:|----------:|----------:|---------:|----|--------:|----:|
| 1 | 60.0 | 2835 | 0 | 2835 | 47 | 47 | 65.5 | 65.5 | 100% | 115ms | 115ms | 0 | 1 | 0 |
| 2 | 60.0 | 4799 | 0 | 4799 | 80 | 80 | 40.7 | 40.7 | 170% | 112ms | 127ms | 0 | 2 | 0 |
| 3 | 60.0 | 10015 | 0 | 10015 | 167 | 167 | 88.5 | 88.5 | 355% | 2.6s | 2.6s | 0 | 3 | 0 |
| 4 | 60.0 | 11075 | 0 | 11075 | 184 | 184 | 54.1 | 54.1 | 391% | 5.4s | 5.4s | 0 | 4 | 0 |
| 5 | 60.0 | 6387 | 0 | 6387 | 106 | 106 | 23.6 | 23.6 | 226% | 5.8s | 5.8s | 0 | 5 | 0 |
| 6 | 60.0 | 7190 | 0 | 7190 | 120 | 120 | 21.2 | 21.2 | 255% | 3.5s | 3.5s | 0 | 6 | 0 |
| 7 | 60.0 | 8407 | 0 | 8407 | 140 | 140 | 21.4 | 21.4 | 298% | 3.8s | 3.8s | 0 | 7 | 0 |
| 8 | 60.0 | 8127 | 0 | 8127 | 135 | 135 | 18.7 | 18.7 | 287% | 5.7s | 5.7s | 0 | 8 | 0 |
| 9 | 60.0 | 17849 | 0 | 17849 | 297 | 297 | 37.1 | 37.1 | 632% | 5.3s | 5.4s | 0 | 9 | 0 |
| 10 | 60.0 | 17841 | 0 | 17841 | 297 | 297 | 38.0 | 38.0 | 632% | 11.4s | 11.4s | 0 | 10 | 0 |
| 11 | 60.0 | 18324 | 0 | 18324 | 305 | 305 | 36.1 | 36.1 | 649% | 11.1s | 11.1s | 0 | 11 | 0 |
| 12 | 60.0 | 15193 | 0 | 15193 | 253 | 253 | 37.8 | 37.8 | 538% | 26.5s | 26.5s | 0 | 12 | 0 |
| 13 | 60.0 | 21925 | 0 | 21925 | 365 | 365 | 33.5 | 33.5 | 777% | 9.5s | 9.6s | 0 | 13 | 0 |
| 14 | 60.0 | 21584 | 0 | 21584 | 359 | 359 | 33.9 | 33.9 | 764% | 12.9s | 12.9s | 0 | 14 | 0 |
| 15 | 60.0 | 22363 | 0 | 22363 | 372 | 372 | 31.8 | 31.8 | 791% | 12.9s | 12.9s | 0 | 15 | 0 |
| 16 | 60.0 | 22537 | 0 | 22537 | 375 | 375 | 30.8 | 30.8 | 798% | 14.2s | 14.3s | 0 | 16 | 0 |
| 17 | 60.0 | 22218 | 0 | 22218 | 370 | 370 | 27.8 | 27.8 | 787% | 13.0s | 13.1s | 0 | 17 | 0 |
| 18 | 60.0 | 22455 | 0 | 22455 | 374 | 374 | 26.8 | 26.8 | 796% | 13.4s | 13.5s | 0 | 18 | 0 |
| 19 | 60.0 | 22909 | 0 | 22909 | 382 | 382 | 25.9 | 25.9 | 813% | 13.3s | 13.5s | 0 | 19 | 0 |
| 20 | 60.0 | 22237 | 0 | 22237 | 370 | 370 | 25.5 | 25.5 | 787% | 16.5s | 16.6s | 0 | 20 | 0 |
| 21 | 60.0 | 23800 | 0 | 23800 | 396 | 396 | 24.7 | 24.7 | 843% | 14.1s | 14.3s | 0 | 21 | 0 |
| 22 | 60.0 | 23630 | 0 | 23630 | 394 | 394 | 24.3 | 24.3 | 838% | 15.7s | 15.9s | 0 | 22 | 0 |
| 23 | 60.0 | 24605 | 0 | 24605 | 410 | 410 | 23.8 | 23.8 | 872% | 15.0s | 15.2s | 0 | 23 | 0 |
| 24 | 60.0 | 24005 | 0 | 24005 | 400 | 400 | 23.6 | 23.6 | 851% | 17.7s | 17.9s | 0 | 24 | 0 |

## Key findings

- **Clean, all-content output**: 0 reasoning tokens everywhere.
- **Slowest single-agent of every model tested**: only **47 tok/s at 1 agent** — notably slower than granite-4.1-3b (79), chronos (88), and llama-3.2-3b (95). With the **Q4_0 KV quant**, the KV cache path is 4-bit; on this setup that quant appears to hurt rather than help single-stream speed.
- **Giant relative scaling (highest recorded)**: combined climbs to **410 tok/s at 23 agents (8.7x)** — the largest multiplier of any run — but only because the solo baseline is so tiny. Absolute peak (~375–410 tok/s from 13–24) is still mid-pack vs other models.
- **Noisy mid-range**: big swings between 5–8 agents (106–184) and the 9–13 transition (297 → 253 → 365), showing erratic server batching behavior for this model.
- **TTFT**: 112ms solo → 17.7s at 24 agents; a ~26.5s outlier at concurrency 12.
- **Per-agent speed**: ~24–34 tok/s at high concurrency (vs 65 solo).
- **Errors**: 0 across all 24 runs.

## Charts

Generated with `benchmark/charts.py` (matplotlib), saved as PNG in `results/unsloth-ministral-3-3b-instruct-2512/charts/`:

| Chart | File | What it highlights |
|---|---|---|
| Dashboard (2x2 overview) | `dashboard_1_24.png` | Whole sweep in one view |
| Combined throughput | `combined_throughput.png` | 47 → 410 tok/s, highest relative scaling |
| Per-agent throughput | `per_agent_throughput.png` | Per-agent slowdown under contention |
| Scaling efficiency | `scaling_efficiency.png` | 100% → 872% peak |
| Time to first token | `time_to_first_token.png` | 112ms → ~18s latency growth |
| Total tokens generated | `total_tokens_generated.png` | ~3k → 24k tokens/60s |
| Outcome breakdown | `outcome_breakdown.png` | Zero errors |
| Combined vs per-agent | `combined_vs_per_agent.png` | The concurrency trade-off on one chart |
