# llama-3.2-1b-mini-agent Concurrency Benchmark

Full sweep of concurrent agents against LM Studio (`LM Studio's local endpoint`, model `llama-3.2-1b-mini-agent`).

- **Harness**: `benchmark/benchmark.js` (agent loop, streaming, token-level timing), `benchmark/sweep.js` (sweep driver)
- **Config per run**: 60s hard timeout per agent, 12000-token budget per agent, temperature 0.9, agents run a multi-turn loop until token budget or timeout
- **Method**: 24 sequential runs at concurrency 1..24; results in `results/llama-3-2-1b-mini-agent/<N>/`, aggregate in `sweep-summary.{csv,json}` in the same folder
- **Note**: `llama-3.2-1b-mini-agent` emits no reasoning tokens (0 `reasoning_content`), so "total tokens" == content tokens.

## Results (1–24 agents)

| agents | wall(s) | total_tok | combined_tok/s | per-agent_tok/s | scale_eff | per-agent_rel | TTFT_mean | TTFT_max | ok | timeout | err |
|-------:|--------:|----------:|---------------:|----------------:|----------:|--------------:|----------:|---------:|----|--------:|----:|
| 1 | 60.0 | 11741 | 196 | 209.3 | 100% | 107% | 880ms | 880ms | 0 | 1 | 0 |
| 2 | 60.0 | 18908 | 315 | 171.4 | 161% | 87% | 188ms | 194ms | 0 | 2 | 0 |
| 3 | 60.0 | 20642 | 344 | 161.3 | 176% | 82% | 10391ms | 10402ms | 0 | 3 | 0 |
| 4 | 60.0 | 25816 | 430 | 169.2 | 219% | 86% | 8754ms | 8763ms | 0 | 4 | 0 |
| 5 | 60.0 | 28081 | 468 | 126.1 | 239% | 64% | 9397ms | 9431ms | 0 | 5 | 0 |
| 6 | 60.0 | 27983 | 466 | 107.8 | 238% | 55% | 10394ms | 10410ms | 0 | 6 | 0 |
| 7 | 60.0 | 30879 | 514 | 81.7 | 262% | 42% | 3001ms | 3022ms | 0 | 7 | 0 |
| 8 | 60.0 | 32819 | 547 | 81.3 | 279% | 41% | 782ms | 819ms | 1 | 7 | 0 |
| 9 | 60.1 | 30729 | 512 | 70.1 | 261% | 36% | 555ms | 578ms | 2 | 7 | 0 |
| 10 | 60.0 | 33683 | 561 | 73.3 | 286% | 37% | 281ms | 295ms | 3 | 7 | 0 |
| 11 | 60.0 | 36650 | 610 | 80.0 | 311% | 41% | 570ms | 594ms | 5 | 6 | 0 |
| 12 | 60.1 | 36701 | 611 | 71.0 | 312% | 36% | 1010ms | 1040ms | 4 | 8 | 0 |
| 13 | 60.1 | 38681 | 644 | 72.4 | 329% | 37% | 1454ms | 1477ms | 6 | 7 | 0 |
| 14 | 60.0 | 36525 | 608 | 64.0 | 310% | 33% | 1520ms | 1564ms | 4 | 10 | 0 |
| 15 | 60.1 | 37062 | 617 | 59.9 | 315% | 31% | 854ms | 908ms | 5 | 10 | 0 |
| 16 | 60.1 | 38106 | 634 | 61.2 | 323% | 31% | 2178ms | 2229ms | 6 | 10 | 0 |
| 17 | 60.1 | 33519 | 558 | 52.2 | 285% | 27% | 258ms | 308ms | 7 | 10 | 0 |
| 18 | 60.0 | 33750 | 562 | 43.0 | 287% | 22% | 1104ms | 1229ms | 6 | 12 | 0 |
| 19 | 60.1 | 38547 | 642 | 54.8 | 328% | 28% | 716ms | 794ms | 7 | 12 | 0 |
| 20 | 60.0 | 37080 | 618 | 47.4 | 315% | 24% | 638ms | 751ms | 8 | 12 | 0 |
| 21 | 60.1 | 37177 | 619 | 46.0 | 316% | 23% | 591ms | 650ms | 10 | 11 | 0 |
| 22 | 60.1 | 36271 | 604 | 45.5 | 308% | 23% | 676ms | 761ms | 9 | 13 | 0 |
| 23 | 60.0 | 35518 | 591 | 38.1 | 302% | 19% | 1695ms | 1900ms | 9 | 14 | 0 |
| 24 | 60.1 | 39161 | 652 | 45.0 | 333% | 23% | 674ms | 791ms | 12 | 12 | 0 |

## Key findings

- **Throughput**: combined tok/s rises from 196 (1 agent) to a plateau of ~600–650 tok/s at 13+ agents; peak is 652 tok/s at 24 agents (3.3x single-agent). Adding agents past ~13–16 buys almost nothing.
- **Per-agent speed**: erodes from 209 tok/s alone to 38–55 tok/s at 17–24 agents (19–27% of single-agent throughput) due to slot contention.
- **TTFT**: highly variable (188ms to ~10.4s) across runs — reflects LM Studio batching/prefill under load, not a clean linear queue.
- **Errors**: 0 across all 24 runs.
- **Caveat**: at higher concurrency some agents end early via an "empty-turn" stop (model returns a turn with no content while competing for a slot); these are counted as `ok` with roughly 1.1–1.9k tokens rather than reaching the 12k budget.

## Charts

Generated with `benchmark/charts.py` (matplotlib), saved as PNG in `results/llama-3-2-1b-mini-agent/charts/`:

| Chart | File | What it highlights |
|---|---|---|
| Dashboard (2x2 overview) | `dashboard_1_24.png` | Whole sweep in one view |
| Combined throughput | `combined_throughput.png` | Sublinear scaling, plateau at 13+ agents |
| Per-agent throughput | `per_agent_throughput.png` | ~78% per-agent slowdown from slot contention |
| Scaling efficiency | `scaling_efficiency.png` | 100% → 333% total, while per-agent share collapses |
| Time to first token | `time_to_first_token.png` | Erratic 0.2s–10.4s batching/prefill latency |
| Total tokens generated | `total_tokens_generated.png` | Server's ~30–39k tokens/60s aggregate ceiling |
| Outcome breakdown | `outcome_breakdown.png` | Zero errors; empty-turn early stops at high concurrency |
| Combined vs per-agent | `combined_vs_per_agent.png` | The concurrency trade-off on one chart |
