# llama-3.2-3b-instruct Concurrency Benchmark

Full sweep of concurrent agents against LM Studio (`LM Studio's local endpoint`, model `llama-3.2-3b-instruct`).

- **Harness**: `benchmark/benchmark.js` (agent loop, streaming, token-level timing), `benchmark/sweep.js` (sweep driver)
- **Config per run**: standard settings — 60s hard timeout, 12000-token total budget, `--output-cap 12000`, 32768-token context, temperature 0.9, agents run a multi-turn loop
- **Method**: 24 sequential runs at concurrency 1..24; results in `results/llama-3.2-3b-instruct/<N>/`, aggregate in `sweep-summary.{csv,json}` in the same folder
- **Reasoning tokens**: 0 across every run — a clean non-thinking model; all output is visible content.

## Results (1–24 agents)

| agents | wall(s) | content_tok | reason_tok | total_tok | comb_tok/s | comb_all_tok/s | per-agent_tok/s | per-agent_all_tok/s | scale_eff | TTFT_mean | TTFT_max | ok | timeout | err |
|-------:|--------:|------------:|-----------:|----------:|-----------:|---------------:|----------------:|---------------------:|----------:|----------:|---------:|----|--------:|----:|
| 1 | 60.0 | 5696 | 0 | 5696 | 95 | 95 | 96.1 | 97.6 | 100% | 117ms | 117ms | 0 | 1 | 0 |
| 2 | 60.0 | 9131 | 0 | 9131 | 152 | 152 | 103.1 | 105.1 | 160% | 118ms | 128ms | 0 | 2 | 0 |
| 3 | 60.0 | 10791 | 0 | 10791 | 180 | 180 | 84.1 | 85.9 | 189% | 2.3s | 2.3s | 0 | 3 | 0 |
| 4 | 60.0 | 12879 | 0 | 12879 | 215 | 215 | 98.0 | 98.7 | 226% | 2.0s | 2.0s | 0 | 4 | 0 |
| 5 | 60.0 | 13437 | 0 | 13437 | 224 | 224 | 76.6 | 76.9 | 236% | 5.4s | 5.4s | 0 | 5 | 0 |
| 6 | 60.0 | 14751 | 0 | 14751 | 246 | 246 | 73.0 | 73.0 | 259% | 5.4s | 5.4s | 0 | 6 | 0 |
| 7 | 60.0 | 15621 | 0 | 15621 | 260 | 260 | 76.8 | 76.8 | 274% | 5.8s | 5.8s | 0 | 7 | 0 |
| 8 | 60.0 | 16191 | 0 | 16191 | 270 | 270 | 54.6 | 54.6 | 284% | 8.7s | 8.7s | 0 | 8 | 0 |
| 9 | 60.0 | 17569 | 0 | 17569 | 293 | 293 | 53.0 | 53.0 | 308% | 9.6s | 9.7s | 0 | 9 | 0 |
| 10 | 60.0 | 18737 | 0 | 18737 | 312 | 312 | 56.5 | 56.5 | 328% | 11.2s | 11.3s | 0 | 10 | 0 |
| 11 | 60.0 | 19319 | 0 | 19319 | 322 | 322 | 52.1 | 52.1 | 339% | 11.3s | 11.3s | 0 | 11 | 0 |
| 12 | 60.0 | 19959 | 0 | 19959 | 332 | 332 | 42.5 | 42.5 | 349% | 12.1s | 12.2s | 0 | 12 | 0 |
| 13 | 60.0 | 20456 | 0 | 20456 | 341 | 341 | 46.4 | 46.4 | 359% | 12.9s | 13.0s | 0 | 13 | 0 |
| 14 | 60.0 | 21708 | 0 | 21708 | 362 | 362 | 45.6 | 45.6 | 381% | 12.8s | 12.8s | 0 | 14 | 0 |
| 15 | 60.0 | 21301 | 0 | 21301 | 355 | 355 | 35.0 | 35.0 | 374% | 14.2s | 14.4s | 0 | 15 | 0 |
| 16 | 60.0 | 22483 | 0 | 22483 | 374 | 374 | 34.9 | 34.9 | 394% | 15.0s | 15.0s | 0 | 16 | 0 |
| 17 | 60.0 | 20385 | 0 | 20385 | 339 | 339 | 29.2 | 29.2 | 357% | 16.2s | 16.3s | 0 | 17 | 0 |
| 18 | 60.0 | 22609 | 0 | 22609 | 377 | 377 | 28.1 | 28.1 | 397% | 13.2s | 13.4s | 0 | 18 | 0 |
| 19 | 60.0 | 22781 | 0 | 22781 | 379 | 379 | 27.3 | 27.3 | 399% | 13.8s | 13.9s | 0 | 19 | 0 |
| 20 | 60.0 | 22689 | 0 | 22689 | 378 | 378 | 25.9 | 25.9 | 398% | 15.8s | 16.0s | 0 | 20 | 0 |
| 21 | 60.0 | 23067 | 0 | 23067 | 384 | 384 | 25.2 | 25.2 | 404% | 15.6s | 15.8s | 0 | 21 | 0 |
| 22 | 60.1 | 23090 | 0 | 23090 | 385 | 385 | 24.3 | 24.3 | 405% | 16.7s | 16.9s | 0 | 22 | 0 |
| 23 | 60.0 | 24114 | 0 | 24114 | 402 | 402 | 24.0 | 24.0 | 423% | 15.5s | 15.7s | 0 | 23 | 0 |
| 24 | 60.0 | 24255 | 0 | 24255 | 404 | 404 | 23.1 | 23.1 | 425% | 16.1s | 16.3s | 0 | 24 | 0 |

## Key findings

- **Clean, all-content output**: 0 reasoning tokens everywhere.
- **Steady, efficient scaling**: combined throughput climbs nearly monotonically from **95 tok/s solo to 404 tok/s at 24 agents (4.25x)** — a smooth, well-behaved curve with no rolloff and no stall (unlike falcon3's rolloff or the thinking models' swings).
- **Slower single-agent than the 1b siblings**: 95 tok/s solo vs 196 for llama-3.2-1b — the extra params cost per-token speed.
- **TTFT grows steadily with load**: 117ms solo, rising smoothly to 12–17s at 14+ agents as requests queue. No erratic 50s+ spikes, but high concurrency clearly serializes starts.
- **Per-agent speed**: ~23–25 tok/s at 24 agents (vs 95 solo).
- **Errors**: 0 across all 24 runs.

## Charts

Generated with `benchmark/charts.py` (matplotlib), saved as PNG in `results/llama-3.2-3b-instruct/charts/`:

| Chart | File | What it highlights |
|---|---|---|
| Dashboard (2x2 overview) | `dashboard_1_24.png` | Whole sweep in one view |
| Combined throughput | `combined_throughput.png` | Smooth 95 → 404 tok/s scaling |
| Per-agent throughput | `per_agent_throughput.png` | Per-agent slowdown under contention |
| Scaling efficiency | `scaling_efficiency.png` | 100% → 425% |
| Time to first token | `time_to_first_token.png` | 117ms → ~16s latency growth |
| Total tokens generated | `total_tokens_generated.png` | ~6k → 24k tokens/60s |
| Outcome breakdown | `outcome_breakdown.png` | Zero errors |
| Combined vs per-agent | `combined_vs_per_agent.png` | The concurrency trade-off on one chart |
