# gemma-2b-it-smashed Concurrency Benchmark

Full sweep of concurrent agents against LM Studio (`LM Studio's local endpoint`, model `gemma-2b-it-smashed`).

- **Harness**: `benchmark/benchmark.js` (agent loop, streaming, token-level timing), `benchmark/sweep.js` (sweep driver)
- **Config per run**: same limits as the falcon3 run — **8k context window**, `--output-cap 2000` (max output per single request) with history trimmed to a rolling window (system + task + latest output + continue prompt), 60s hard timeout, 12000-token total budget (safety net), temperature 0.9, agents run a multi-turn loop
- **Method**: 24 sequential runs at concurrency 1..24; results in `results/gemma-2b-it-smashed/<N>/`, aggregate in `sweep-summary.{csv,json}` in the same folder
- **Reasoning tokens**: 0 across every run — gemma-2b-it-smashed is a clean non-thinking model; all output is visible content.

## Results (1–24 agents)

| agents | wall(s) | content_tok | reason_tok | total_tok | comb_tok/s | comb_all_tok/s | per-agent_tok/s | per-agent_all_tok/s | scale_eff | TTFT_mean | TTFT_max | ok | timeout | err |
|-------:|--------:|------------:|-----------:|----------:|-----------:|---------------:|----------------:|---------------------:|----------:|----------:|---------:|----|--------:|----:|
| 1 | 60.0 | 7289 | 0 | 7289 | 121 | 121 | 132.4 | 134.4 | 100% | 94ms | 94ms | 0 | 1 | 0 |
| 2 | 60.0 | 11374 | 0 | 11374 | 189 | 189 | 99.1 | 106.3 | 156% | 61ms | 68ms | 0 | 2 | 0 |
| 3 | 60.0 | 15109 | 0 | 15109 | 252 | 252 | 89.9 | 94.6 | 208% | 763ms | 776ms | 0 | 3 | 0 |
| 4 | 60.0 | 16773 | 0 | 16773 | 279 | 279 | 76.8 | 80.6 | 231% | 1.1s | 1.1s | 0 | 4 | 0 |
| 5 | 60.0 | 18909 | 0 | 18909 | 315 | 315 | 75.7 | 78.4 | 260% | 843ms | 877ms | 0 | 5 | 0 |
| 6 | 60.1 | 19950 | 0 | 19950 | 332 | 332 | 62.2 | 64.2 | 274% | 2.0s | 2.0s | 0 | 6 | 0 |
| 7 | 60.0 | 21772 | 0 | 21772 | 363 | 363 | 59.1 | 60.7 | 300% | 1.5s | 1.5s | 0 | 7 | 0 |
| 8 | 60.0 | 23116 | 0 | 23116 | 385 | 385 | 51.8 | 53.3 | 318% | 1.7s | 1.7s | 0 | 8 | 0 |
| 9 | 60.0 | 29323 | 0 | 29323 | 488 | 488 | 59.6 | 63.6 | 403% | 614ms | 702ms | 0 | 9 | 0 |
| 10 | 60.0 | 28116 | 0 | 28116 | 468 | 468 | 54.6 | 57.1 | 387% | 1.5s | 1.6s | 1 | 9 | 0 |
| 11 | 60.0 | 30608 | 0 | 30608 | 510 | 510 | 53.4 | 57.0 | 421% | 359ms | 454ms | 1 | 10 | 0 |
| 12 | 60.0 | 29305 | 0 | 29305 | 488 | 488 | 52.9 | 55.9 | 403% | 1.0s | 1.1s | 2 | 10 | 0 |
| 13 | 60.0 | 30756 | 0 | 30756 | 512 | 512 | 57.6 | 61.2 | 423% | 783ms | 872ms | 3 | 10 | 0 |
| 14 | 60.1 | 31170 | 0 | 31170 | 519 | 519 | 49.8 | 54.6 | 429% | 1.8s | 1.9s | 2 | 12 | 0 |
| 15 | 60.0 | 30180 | 0 | 30180 | 503 | 503 | 46.9 | 51.4 | 416% | 1.5s | 1.6s | 4 | 11 | 0 |
| 16 | 60.0 | 31634 | 0 | 31634 | 527 | 527 | 42.5 | 47.3 | 436% | 794ms | 878ms | 2 | 14 | 0 |
| 17 | 60.0 | 32331 | 0 | 32331 | 538 | 538 | 43.4 | 48.0 | 445% | 609ms | 697ms | 4 | 13 | 0 |
| 18 | 60.1 | 32084 | 0 | 32084 | 534 | 534 | 43.0 | 47.8 | 441% | 613ms | 758ms | 4 | 14 | 0 |
| 19 | 60.1 | 31700 | 0 | 31700 | 528 | 528 | 43.0 | 48.3 | 436% | 788ms | 947ms | 4 | 15 | 0 |
| 20 | 60.1 | 31317 | 0 | 31317 | 521 | 521 | 42.6 | 49.3 | 431% | 1.0s | 1.2s | 4 | 16 | 0 |
| 21 | 60.1 | 29364 | 0 | 29364 | 489 | 489 | 39.7 | 45.7 | 404% | 2.2s | 2.3s | 4 | 17 | 0 |
| 22 | 60.1 | 30572 | 0 | 30572 | 509 | 509 | 40.9 | 47.3 | 421% | 955ms | 1.1s | 3 | 19 | 0 |
| 23 | 60.1 | 31852 | 0 | 31852 | 530 | 530 | 43.5 | 50.2 | 438% | 434ms | 581ms | 5 | 18 | 0 |
| 24 | 60.0 | 29957 | 0 | 29957 | 499 | 499 | 55.6 | 61.3 | 412% | 2.2s | 2.4s | 6 | 18 | 0 |

*Every run lasted the full 60s; `timeout` is the intended 60s cap, not an error.*

## Key findings

- **Clean, all-content output**: 0 reasoning tokens everywhere.
- **Steady scaling, modest ceiling**: combined throughput rises from 121 tok/s solo to **538 tok/s at 17 agents (4.5x)**, holding ~490–538 from 9–24. A solid middle-of-the-pack profile.
- **Best latency of the smaller models**: TTFT stays in the **61ms–2.4s** range across the entire sweep — no thinking-model spikes, and better than most.
- **Per-agent speed**: ~43–50 tok/s at high concurrency (vs 121 solo).
- **Errors**: 0 across all 24 runs.

## Charts

Generated with `benchmark/charts.py` (matplotlib), saved as PNG in `results/gemma-2b-it-smashed/charts/`:

| Chart | File | What it highlights |
|---|---|---|
| Dashboard (2x2 overview) | `dashboard_1_24.png` | Whole sweep in one view |
| Combined throughput | `combined_throughput.png` | 121 → 538 tok/s scaling |
| Per-agent throughput | `per_agent_throughput.png` | Per-agent slowdown under contention |
| Scaling efficiency | `scaling_efficiency.png` | 100% → 445% peak |
| Time to first token | `time_to_first_token.png` | 61ms–2.4s latency profile |
| Total tokens generated | `total_tokens_generated.png` | ~7k → 32k tokens/60s |
| Outcome breakdown | `outcome_breakdown.png` | Zero errors |
| Combined vs per-agent | `combined_vs_per_agent.png` | The concurrency trade-off on one chart |
