# google/gemma-4-e2b Concurrency Benchmark

Full sweep of concurrent agents against LM Studio (`LM Studio's local endpoint`, model `google/gemma-4-e2b`).

- **Harness**: `benchmark/benchmark.js` (agent loop, streaming, token-level timing), `benchmark/sweep.js` (sweep driver)
- **Config per run**: standard settings — 60s hard timeout, 12000-token total budget, `--output-cap 12000`, 32768-token context, temperature 0.9, agents run a multi-turn loop
- **Quants note**: this run is served with **K and V cache quants of Q8_0** (8-bit KV cache).
- **Method**: 24 sequential runs at concurrency 1..24; results in `results/google-gemma-4-e2b/<N>/`, aggregate in `sweep-summary.{csv,json}` in the same folder
- **Reasoning tokens**: gemma-4-e2b streams hidden thinking via `reasoning_content`; its share explodes with concurrency (19% of tokens solo to **~99% at 21–24 agents**).

## Results (1–24 agents)

| agents | wall(s) | content_tok | reason_tok | total_tok | comb_tok/s | comb_all_tok/s | per-agent_tok/s | per-agent_all_tok/s | scale_eff | TTFT_mean | TTFT_max | ok | timeout | err |
|-------:|--------:|------------:|-----------:|----------:|-----------:|---------------:|----------------:|---------------------:|----------:|----------:|---------:|----|--------:|----:|
| 1 | 60.0 | 3487 | 814 | 4301 | 58 | 72 | 72.3 | 72.3 | 100% | 11.8s | 11.8s | 0 | 1 | 0 |
| 2 | 60.0 | 5383 | 1761 | 7144 | 90 | 119 | 58.6 | 59.8 | 165% | 14.1s | 14.2s | 0 | 2 | 0 |
| 3 | 60.0 | 5982 | 2528 | 8510 | 100 | 142 | 46.3 | 48.3 | 197% | 16.9s | 18.2s | 0 | 3 | 0 |
| 4 | 60.0 | 6571 | 3776 | 10347 | 109 | 172 | 42.4 | 43.7 | 239% | 21.3s | 27.8s | 0 | 4 | 0 |
| 5 | 60.0 | 6118 | 4256 | 10374 | 102 | 173 | 34.3 | 35.2 | 240% | 24.3s | 25.9s | 0 | 5 | 0 |
| 6 | 60.0 | 6799 | 5062 | 11861 | 113 | 198 | 32.5 | 33.8 | 275% | 25.1s | 26.0s | 0 | 6 | 0 |
| 7 | 60.0 | 6152 | 5894 | 12046 | 102 | 201 | 28.1 | 29.3 | 279% | 28.7s | 31.5s | 0 | 7 | 0 |
| 8 | 60.0 | 6179 | 6749 | 12928 | 103 | 215 | 26.4 | 27.6 | 299% | 30.8s | 40.7s | 0 | 8 | 0 |
| 9 | 60.0 | 5702 | 7779 | 13481 | 95 | 225 | 24.4 | 25.7 | 312% | 34.1s | 42.8s | 0 | 9 | 0 |
| 10 | 60.0 | 5945 | 8513 | 14458 | 99 | 241 | 23.5 | 24.9 | 335% | 34.7s | 41.7s | 0 | 10 | 0 |
| 11 | 60.0 | 5246 | 9195 | 14441 | 87 | 241 | 21.6 | 22.6 | 335% | 37.9s | 42.9s | 0 | 11 | 0 |
| 12 | 60.0 | 4741 | 10389 | 15130 | 79 | 252 | 20.8 | 21.8 | 350% | 41.0s | 48.7s | 0 | 12 | 0 |
| 13 | 60.0 | 4411 | 10753 | 15164 | 73 | 253 | 19.1 | 20.2 | 351% | 42.3s | 54.3s | 0 | 13 | 0 |
| 14 | 60.0 | 4732 | 11449 | 16181 | 79 | 270 | 19.1 | 20.1 | 375% | 42.3s | 53.0s | 0 | 14 | 0 |
| 15 | 60.0 | 3693 | 12505 | 16198 | 62 | 270 | 16.7 | 18.7 | 375% | 45.3s | 55.9s | 0 | 15 | 0 |
| 16 | 60.0 | 3387 | 13521 | 16908 | 56 | 282 | 17.4 | 18.4 | 392% | 47.8s | 56.5s | 0 | 16 | 0 |
| 17 | 60.0 | 1869 | 14538 | 16407 | 31 | 273 | 14.2 | 16.8 | 379% | 52.2s | 58.9s | 0 | 17 | 0 |
| 18 | 60.0 | 2311 | 14858 | 17169 | 38 | 286 | 14.0 | 16.7 | 397% | 50.8s | 58.4s | 0 | 18 | 0 |
| 19 | 60.0 | 1544 | 15592 | 17136 | 26 | 285 | 10.5 | 15.9 | 396% | 52.2s | 57.1s | 0 | 19 | 0 |
| 20 | 60.0 | 852 | 16546 | 17398 | 14 | 290 | 8.9 | 15.3 | 403% | 55.1s | 59.2s | 0 | 20 | 0 |
| 21 | 60.0 | 157 | 16352 | 16509 | 3 | 275 | 3.7 | 13.9 | 382% | 57.8s | 59.7s | 0 | 21 | 0 |
| 22 | 60.0 | 348 | 16728 | 17076 | 6 | 284 | 3.2 | 13.7 | 394% | 54.7s | 59.8s | 0 | 22 | 0 |
| 23 | 60.0 | 95 | 17255 | 17350 | 2 | 289 | 0.6 | 13.4 | 401% | 52.5s | 52.5s | 0 | 23 | 0 |
| 24 | 60.0 | 235 | 17372 | 17607 | 4 | 293 | 0.5 | 13.1 | 407% | 41.4s | 41.4s | 0 | 24 | 0 |

## Key findings

- **One of the slowest profiles tested**: only **72 tok/s solo** (58 visible), and combined all-token throughput peaks at just **293 tok/s at 24 agents (4.1x)** — near the bottom of the field alongside qwen3-4b.
- **Near-total reasoning collapse under load**: visible content drops to 2–6 tok/s at 20–24 agents (~99% of output is hidden `reasoning_content`). At 21–24 agents agents produce essentially zero usable content in the 60s window.
- **Extreme TTFT — the second-worst after qwen3-4b**: **11.8s solo**, rising steadily to **41–58s under load**, with max values up to 59.8s. Agents routinely wait almost the whole run for their first token.
- **Steady but slow scaling**: all-token throughput climbs monotonically 72 → 293 without rolloff, but absolute numbers are low.
- **Errors**: 0 across all 24 runs.

## Charts

Generated with `benchmark/charts.py` (matplotlib), saved as PNG in `results/google-gemma-4-e2b/charts/`:

| Chart | File | What it highlights |
|---|---|---|
| Dashboard (2x2 overview) | `dashboard_1_24.png` | Whole sweep in one view |
| Combined throughput | `combined_throughput.png` | Content vs content+reasoning curves |
| Per-agent throughput | `per_agent_throughput.png` | Per-agent slowdown under contention |
| Scaling efficiency | `scaling_efficiency.png` | 100% → 407% |
| Time to first token | `time_to_first_token.png` | 12s solo → ~58s — among the worst |
| Total tokens generated | `total_tokens_generated.png` | Content vs reasoning stack per concurrency |
| Outcome breakdown | `outcome_breakdown.png` | Zero errors |
| Combined vs per-agent | `combined_vs_per_agent.png` | The concurrency trade-off on one chart |
| Content vs reasoning | `reasoning_vs_content.png` | ~99% reasoning at high concurrency |
