Gemma 2B IT (smashed)

model page · concurrency sweep 1–24 agents

Gemma 2B IT (smashed)

server model id: gemma-2b-it-smashed
non-thinkingKV Q8_0ctx 8k2B params

Middle of the pack and steady — 538 tok/s peak, and the best latency of the smaller models (61ms–2.4s).

Key findings

  • Clean, all-content output: 0 reasoning tokens everywhere.
  • Steady scaling, modest ceiling: combined throughput rises from 121 tok/s solo to 538 tok/s at 17 agents (4.5x), holding ~490–538 from 9–24. A solid middle-of-the-pack profile.
  • Best latency of the smaller models: TTFT stays in the 61ms–2.4s range across the entire sweep — no thinking-model spikes, and better than most.
  • Per-agent speed: ~43–50 tok/s at high concurrency (vs 121 solo).
  • Errors: 0 across all 24 runs.

Sweep — total concurrency 1–24

agentswall (s)contentreasontotalcomb tok/scomb allper-agentper-agent allscale %TTFT meanTTFT maxoktimeouterr
160.07,28907,289121121132.4134.4100%94ms94ms010
260.011,374011,37418918999.1106.3156%61ms68ms020
360.015,109015,10925225289.994.6208%763ms776ms030
460.016,773016,77327927976.880.6231%1.1s1.1s040
560.018,909018,90931531575.778.4260%843ms877ms050
660.119,950019,95033233262.264.2274%2.0s2.0s060
760.021,772021,77236336359.160.7300%1.5s1.5s070
860.023,116023,11638538551.853.3318%1.7s1.7s080
960.029,323029,32348848859.663.6403%614ms702ms090
1060.028,116028,11646846854.657.1387%1.5s1.6s190
1160.130,608030,60851051053.457.0421%359ms454ms1100
1260.129,305029,30548848852.955.9403%1.0s1.1s2100
1360.030,756030,75651251257.661.2423%783ms872ms3100
1460.131,170031,17051951949.854.6429%1.8s1.9s2120
1560.130,180030,18050350346.951.4416%1.5s1.6s4110
1660.031,634031,63452752742.547.3436%794ms878ms2140
17 peak60.132,331032,33153853843.448.0445%609ms697ms4130
1860.132,084032,08453453443.047.8441%613ms758ms4140
1960.131,700031,70052852843.048.3436%788ms947ms4150
2060.131,317031,31752152142.649.3431%1.0s1.2s4160
2160.129,364029,36448948939.745.7404%2.2s2.3s4170
2260.130,572030,57250950940.947.3421%955ms1.1s3190
2360.131,852031,85253053043.550.2438%434ms581ms5180
2460.029,957029,95749949955.661.3412%2.2s2.4s6180

* burst artifact — agents queued ~11–16s then generated in a short burst; the per-agent timer inflates the number. Trust combined throughput and TTFT columns. Peak row = highest combined (all-token) throughput.

Charts

Whole sweep at a glance — combined throughput, scaling efficiency, per-agent throughput and outcomes across every concurrency level.
Whole sweep at a glance — combined throughput, scaling efficiency, per-agent throughput and outcomes across every concurrency level.
Combined throughput (all tokens) vs concurrent agents. Annotated peak with the concurrency level where it occurred.
Combined throughput (all tokens) vs concurrent agents. Annotated peak with the concurrency level where it occurred.
Per-agent throughput erosion as slots contend. Red stars mark burst-timing artifacts (queued ≥ 10s then generated in a burst — inflated values, not real throughput).
Per-agent throughput erosion as slots contend. Red stars mark burst-timing artifacts (queued ≥ 10s then generated in a burst — inflated values, not real throughput).
Scaling efficiency relative to the single-agent baseline (100%).
Scaling efficiency relative to the single-agent baseline (100%).
Time to first token — mean (solid) and max (dashed). Log scale. Gaps mean no visible content ever arrived within the 60s window.
Time to first token — mean (solid) and max (dashed). Log scale. Gaps mean no visible content ever arrived within the 60s window.
Total tokens per 60s run — all tokens vs visible content only.
Total tokens per 60s run — all tokens vs visible content only.
How each run of N agents finished: completed, timed out (60s budget), or errored.
How each run of N agents finished: completed, timed out (60s budget), or errored.
The concurrency trade-off on one chart — combined throughput rises while per-agent speed falls.
The concurrency trade-off on one chart — combined throughput rises while per-agent speed falls.

Raw data & downloads