# granite-4.1-3b Concurrency Benchmark

Full sweep of concurrent agents against LM Studio (`LM Studio's local endpoint`, model `granite-4.1-3b`).

- **Harness**: `benchmark/benchmark.js` (agent loop, streaming, token-level timing), `benchmark/sweep.js` (sweep driver)
- **Config per run**: standard settings — 60s hard timeout, 12000-token total budget, `--output-cap 12000`, 32768-token context, temperature 0.9, agents run a multi-turn loop
- **Method**: 24 sequential runs at concurrency 1..24; results in `results/granite-4.1-3b/<N>/`, aggregate in `sweep-summary.{csv,json}` in the same folder
- **Reasoning tokens**: 0 across every run — a clean non-thinking model; all output is visible content.

## Results (1–24 agents)

| agents | wall(s) | content_tok | reason_tok | total_tok | comb_tok/s | comb_all_tok/s | per-agent_tok/s | per-agent_all_tok/s | scale_eff | TTFT_mean | TTFT_max | ok | timeout | err |
|-------:|--------:|------------:|-----------:|----------:|-----------:|---------------:|----------------:|---------------------:|----------:|----------:|---------:|----|--------:|----:|
| 1 | 60.0 | 4749 | 0 | 4749 | 79 | 79 | 84.6 | 85.3 | 100% | 191ms | 191ms | 0 | 1 | 0 |
| 2 | 60.0 | 7660 | 0 | 7660 | 128 | 128 | 73.5 | 76.4 | 162% | 156ms | 164ms | 0 | 2 | 0 |
| 3 | 60.1 | 9511 | 0 | 9511 | 158 | 158 | 64.8 | 68.5 | 200% | 2.8s | 2.8s | 0 | 3 | 0 |
| 4 | 60.0 | 11894 | 0 | 11894 | 198 | 198 | 65.0 | 65.6 | 251% | 3.1s | 3.2s | 0 | 4 | 0 |
| 5 | 60.0 | 12693 | 0 | 12693 | 211 | 211 | 63.4 | 64.1 | 267% | 4.5s | 4.5s | 0 | 5 | 0 |
| 6 | 60.0 | 14137 | 0 | 14137 | 235 | 235 | 69.4 | 69.4 | 297% | 5.7s | 5.7s | 0 | 6 | 0 |
| 7 | 60.0 | 14417 | 0 | 14417 | 240 | 240 | 52.8 | 52.8 | 304% | 8.9s | 9.0s | 0 | 7 | 0 |
| 8 | 60.0 | 16199 | 0 | 16199 | 270 | 270 | 57.7 | 57.7 | 342% | 8.8s | 8.8s | 0 | 8 | 0 |
| 9 | 60.0 | 16717 | 0 | 16717 | 278 | 278 | 53.6 | 53.6 | 352% | 12.7s | 12.7s | 0 | 9 | 0 |
| 10 | 60.0 | 14791 | 0 | 14791 | 246 | 246 | 35.3 | 35.3 | 311% | 13.0s | 13.0s | 0 | 10 | 0 |
| 11 | 60.0 | 19682 | 0 | 19682 | 328 | 328 | 42.8 | 42.9 | 415% | 12.6s | 12.6s | 0 | 11 | 0 |
| 12 | 60.0 | 14036 | 0 | 14036 | 234 | 234 | 24.2 | 24.2 | 296% | 8.5s | 8.6s | 0 | 12 | 0 |
| 13 | 60.0 | 15316 | 0 | 15316 | 255 | 255 | 21.3 | 21.3 | 323% | 3.3s | 3.5s | 0 | 13 | 0 |
| 14 | 60.0 | 20696 | 0 | 20696 | 345 | 345 | 40.8 | 40.8 | 437% | 17.2s | 17.2s | 0 | 14 | 0 |
| 15 | 60.0 | 20653 | 0 | 20653 | 344 | 344 | 35.7 | 35.7 | 435% | 19.4s | 19.4s | 0 | 15 | 0 |
| 16 | 60.0 | 21755 | 0 | 21755 | 362 | 362 | 34.8 | 34.8 | 458% | 19.1s | 19.2s | 0 | 16 | 0 |
| 17 | 60.0 | 20303 | 0 | 20303 | 338 | 338 | 30.8 | 30.8 | 428% | 20.3s | 20.4s | 0 | 17 | 0 |
| 18 | 60.0 | 21551 | 0 | 21551 | 359 | 359 | 30.6 | 30.6 | 454% | 19.7s | 19.8s | 0 | 18 | 0 |
| 19 | 60.0 | 18241 | 0 | 18241 | 304 | 304 | 22.6 | 22.6 | 385% | 17.4s | 17.9s | 0 | 19 | 0 |
| 20 | 60.0 | 16037 | 0 | 16037 | 267 | 267 | 22.5 | 22.5 | 338% | 24.3s | 24.8s | 0 | 20 | 0 |
| 21 | 60.0 | 19491 | 0 | 19491 | 325 | 325 | 22.2 | 22.2 | 411% | 18.2s | 18.6s | 0 | 21 | 0 |
| 22 | 60.0 | 18958 | 0 | 18958 | 316 | 316 | 21.9 | 21.9 | 400% | 20.7s | 21.3s | 0 | 22 | 0 |
| 23 | 60.0 | 20125 | 0 | 20125 | 335 | 335 | 21.7 | 21.7 | 424% | 19.7s | 20.2s | 0 | 23 | 0 |
| 24 | 60.0 | 19826 | 0 | 19826 | 330 | 330 | 21.4 | 21.4 | 418% | 21.5s | 22.0s | 0 | 24 | 0 |

## Key findings

- **Clean, all-content output**: 0 reasoning tokens everywhere.
- **Slowest solo of all models tested**: only **79 tok/s at 1 agent** (even slower than chronos's 88). Combined peaks at **362 tok/s at 16 agents (4.6x)** but the curve is **noisy/erratic** — dips at 10 (246), 12 (234), 19–20 (304/267) amid the climb. Granite shows the least consistent scaling of the non-thinking models.
- **TTFT climbs the hardest**: 191ms solo rising to **~20–24s at 20+ agents**, with the worst values of the clean models.
- **Per-agent speed**: ~21–22 tok/s at 24 agents (vs 85 solo).
- **Errors**: 0 across all 24 runs.

## Charts

Generated with `benchmark/charts.py` (matplotlib), saved as PNG in `results/granite-4.1-3b/charts/`:

| Chart | File | What it highlights |
|---|---|---|
| Dashboard (2x2 overview) | `dashboard_1_24.png` | Whole sweep in one view |
| Combined throughput | `combined_throughput.png` | Noisy 79 → 362 tok/s climb |
| Per-agent throughput | `per_agent_throughput.png` | Per-agent slowdown under contention |
| Scaling efficiency | `scaling_efficiency.png` | 100% → 458% peak |
| Time to first token | `time_to_first_token.png` | 156ms → ~24s latency growth |
| Total tokens generated | `total_tokens_generated.png` | ~5k → 22k tokens/60s |
| Outcome breakdown | `outcome_breakdown.png` | Zero errors |
| Combined vs per-agent | `combined_vs_per_agent.png` | The concurrency trade-off on one chart |
