# gemma-3-1b-it-glm-4.7-flash-heretic-uncensored-thinking_gguf Concurrency Benchmark

Full sweep of concurrent agents against LM Studio (`LM Studio's local endpoint`, model `gemma-3-1b-it-glm-4.7-flash-heretic-uncensored-thinking_gguf`).

- **Harness**: `benchmark/benchmark.js` (agent loop, streaming, token-level timing), `benchmark/sweep.js` (sweep driver)
- **Config per run**: 60s hard timeout per agent, 12000-token budget per agent, temperature 0.9, agents run a multi-turn loop until token budget or timeout
- **Method**: 24 sequential runs at concurrency 1..24; results in `results/gemma-3-1b-it-glm-4.7-flash-heretic-uncensored-thinking_gguf/<N>/`, aggregate in `sweep-summary.{csv,json}` in the same folder
- **Reasoning tokens**: unlike minicpm (which used `reasoning_content`), this model streams its thinking inline as `<think>...</think>` tags inside the content. The harness detects these tags and counts the enclosed tokens as reasoning (`reason_tok`); the tags themselves are stripped from the visible `content_tok` stream and never fed back into the continuation context.

## Results (1–24 agents)

| agents | wall(s) | content_tok | reason_tok | total_tok | comb_tok/s | comb_all_tok/s | per-agent_tok/s | per-agent_all_tok/s | scale_eff | TTFT_mean | TTFT_max | ok | timeout | err |
|-------:|--------:|------------:|-----------:|----------:|-----------:|---------------:|----------------:|---------------------:|----------:|----------:|---------:|----|--------:|----:|
| 1 | 60.0 | 8922 | 1354 | 10276 | 149 | 171 | 162.2 | 188.2 | 100% | 86ms | 86ms | 0 | 1 | 0 |
| 2 | 60.0 | 14359 | 4393 | 18752 | 239 | 312 | 217.1 | 284.4 | 182% | 102ms | 114ms | 0 | 2 | 0 |
| 3 | 60.0 | 17370 | 5827 | 23197 | 289 | 386 | 118.4 | 147.3 | 226% | 1.4s | 1.4s | 0 | 3 | 0 |
| 4 | 60.0 | 20495 | 4922 | 25417 | 341 | 423 | 57.0 | 69.4 | 247% | 2.3s | 2.3s | 0 | 4 | 0 |
| 5 | 60.0 | 19786 | 5792 | 25578 | 330 | 426 | 55.6 | 79.4 | 249% | 2.7s | 2.7s | 0 | 5 | 0 |
| 6 | 60.0 | 18635 | 9526 | 28161 | 310 | 469 | 48.8 | 66.5 | 274% | 2.5s | 2.5s | 0 | 6 | 0 |
| 7 | 60.0 | 19724 | 8669 | 28393 | 329 | 473 | 38.4 | 59.7 | 277% | 3.2s | 3.2s | 0 | 7 | 0 |
| 8 | 60.0 | 19550 | 8763 | 28313 | 326 | 472 | 36.3 | 45.4 | 276% | 4.1s | 4.2s | 0 | 8 | 0 |
| 9 | 60.0 | 26857 | 7207 | 34064 | 447 | 567 | 62.9 | 74.9 | 332% | 3.7s | 3.7s | 0 | 9 | 0 |
| 10 | 60.0 | 21787 | 14632 | 36419 | 363 | 607 | 58.6 | 81.6 | 355% | 4.4s | 4.4s | 1 | 9 | 0 |
| 11 | 60.0 | 23557 | 14515 | 38072 | 392 | 634 | 48.2 | 78.0 | 371% | 2.9s | 2.9s | 0 | 11 | 0 |
| 12 | 60.0 | 24200 | 14343 | 38543 | 403 | 642 | 54.5 | 73.0 | 375% | 3.5s | 3.5s | 1 | 11 | 0 |
| 13 | 60.0 | 24937 | 15489 | 40426 | 415 | 673 | 55.1 | 76.7 | 394% | 3.8s | 3.8s | 1 | 12 | 0 |
| 14 | 60.0 | 27272 | 12930 | 40202 | 454 | 670 | 47.5 | 70.2 | 392% | 4.1s | 4.1s | 0 | 14 | 0 |
| 15 | 60.0 | 24444 | 16038 | 40482 | 407 | 674 | 69.9 | 69.5 | 394% | 4.3s | 4.3s | 4 | 11 | 0 |
| 16 | 60.0 | 24868 | 16176 | 41044 | 414 | 684 | 61.0 | 63.5 | 400% | 3.6s | 3.6s | 4 | 12 | 0 |
| 17 | 60.0 | 22070 | 19933 | 42003 | 368 | 700 | 51.2 | 65.1 | 409% | 4.0s | 4.0s | 4 | 13 | 0 |
| 18 | 60.0 | 24557 | 17007 | 41564 | 409 | 692 | 76.8 | 64.0 | 405% | 4.7s | 4.7s | 7 | 11 | 0 |
| 19 | 60.0 | 25391 | 16368 | 41759 | 423 | 695 | 73.2 | 62.0 | 406% | 4.1s | 4.1s | 8 | 11 | 0 |
| 20 | 60.0 | 23776 | 17549 | 41325 | 396 | 688 | 72.3 | 59.5 | 402% | 4.2s | 4.2s | 10 | 10 | 0 |
| 21 | 60.0 | 18108 | 20089 | 38197 | 302 | 636 | 76.9 | 51.3 | 372% | 4.6s | 4.6s | 15 | 6 | 0 |
| 22 | 60.0 | 22247 | 19914 | 42161 | 370 | 702 | 54.6 | 51.7 | 411% | 2.5s | 2.5s | 13 | 9 | 0 |
| 23 | 60.0 | 16252 | 22703 | 38955 | 271 | 649 | 57.3 | 47.8 | 380% | 3.6s | 3.6s | 17 | 6 | 0 |
| 24 | 60.0 | 22958 | 20217 | 43175 | 382 | 719 | 55.2 | 55.4 | 420% | 2.9s | 2.9s | 15 | 9 | 0 |

## Key findings

- **Reasonable latency**: TTFT stays in the 86ms–4.7s range across the whole sweep (single agent: 86ms) — dramatically better than minicpm5's 1.5s–52s. This is the friendliest "thinking" profile of the three models tested.
- **All-in throughput**: counting reasoning, combined throughput reaches **719 tok/s at 24 agents (4.2x single-agent)**, mid-way between llama-3.2 (652) and minicpm (983). Visible content tops out around ~454 tok/s.
- **Thinking share grows with load**: reasoning is only ~13% of tokens at concurrency 1, but climbs to ~45–58% at 17–24 — the model thinks proportionally more when slots contend.
- **Per-agent speed**: falls from ~162–217 tok/s at low concurrency to ~36–77 at high concurrency — same slot-contention pattern as the other models.
- **Errors**: 0 across all 24 runs.

## Charts

Generated with `benchmark/charts.py` (matplotlib), saved as PNG in `results/gemma-3-1b-it-glm-4.7-flash-heretic-uncensored-thinking_gguf/charts/`:

| Chart | File | What it highlights |
|---|---|---|
| Dashboard (2x2 overview) | `dashboard_1_24.png` | Whole sweep in one view |
| Combined throughput | `combined_throughput.png` | Content vs content+reasoning curves |
| Per-agent throughput | `per_agent_throughput.png` | Per-agent slowdown under contention |
| Scaling efficiency | `scaling_efficiency.png` | 100% → 420% total, per-agent share collapses |
| Time to first token | `time_to_first_token.png` | 86ms–4.7s latency profile |
| Total tokens generated | `total_tokens_generated.png` | Content vs reasoning stack per concurrency |
| Outcome breakdown | `outcome_breakdown.png` | Zero errors; natural stops at high concurrency |
| Combined vs per-agent | `combined_vs_per_agent.png` | The concurrency trade-off on one chart |
| Content vs reasoning | `reasoning_vs_content.png` | Hidden `<think>` reasoning share per run |
