Granite 4.1 3B

model page · concurrency sweep 1–24 agents

Granite 4.1 3B

server model id: granite-4.1-3b
non-thinkingKV Q4_0ctx 32k3B params

Slowest solo of all (79 tok/s) and the noisiest scaling curve, but TTFT is the real weak point — 20–24s at 20+ agents.

Key findings

  • Clean, all-content output: 0 reasoning tokens everywhere.
  • Slowest solo of all models tested: only 79 tok/s at 1 agent (even slower than chronos's 88). Combined peaks at 362 tok/s at 16 agents (4.6x) but the curve is noisy/erratic — dips at 10 (246), 12 (234), 19–20 (304/267) amid the climb. Granite shows the least consistent scaling of the non-thinking models.
  • TTFT climbs the hardest: 191ms solo rising to ~20–24s at 20+ agents, with the worst values of the clean models.
  • Per-agent speed: ~21–22 tok/s at 24 agents (vs 85 solo).
  • Errors: 0 across all 24 runs.

Sweep — total concurrency 1–24

agentswall (s)contentreasontotalcomb tok/scomb allper-agentper-agent allscale %TTFT meanTTFT maxoktimeouterr
160.04,74904,749797984.685.3100%191ms191ms010
260.07,66007,66012812873.576.4162%156ms164ms020
360.19,51109,51115815864.868.5200%2.8s2.8s030
460.011,894011,89419819865.065.6251%3.1s3.2s040
560.012,693012,69321121163.464.1267%4.5s4.5s050
660.014,137014,13723523569.469.4297%5.7s5.7s060
760.014,417014,41724024052.852.8304%8.9s9.0s070
860.016,199016,19927027057.757.7342%8.8s8.8s080
960.016,717016,71727827853.653.6352%12.7s12.7s090
1060.014,791014,79124624635.335.3311%13.0s13.0s0100
1160.019,682019,68232832842.842.9415%12.6s12.6s0110
1260.014,036014,03623423424.224.2296%8.5s8.6s0120
1360.015,316015,31625525521.321.3323%3.3s3.5s0130
1460.020,696020,69634534540.840.8437%17.2s17.2s0140
1560.020,653020,65334434435.735.7435%19.4s19.4s0150
16 peak60.021,755021,75536236234.834.8458%19.1s19.2s0160
1760.020,303020,30333833830.830.8428%20.3s20.4s0170
1860.021,551021,55135935930.630.6454%19.7s19.8s0180
1960.018,241018,24130430422.622.6385%17.4s17.9s0190
2060.016,037016,03726726722.522.5338%24.3s24.8s0200
2160.019,491019,49132532522.222.2411%18.2s18.6s0210
2260.018,958018,95831631621.921.9400%20.7s21.3s0220
2360.020,125020,12533533521.721.7424%19.7s20.2s0230
2460.019,826019,82633033021.421.4418%21.5s22.0s0240

* burst artifact — agents queued ~11–16s then generated in a short burst; the per-agent timer inflates the number. Trust combined throughput and TTFT columns. Peak row = highest combined (all-token) throughput.

Charts

Whole sweep at a glance — combined throughput, scaling efficiency, per-agent throughput and outcomes across every concurrency level.
Whole sweep at a glance — combined throughput, scaling efficiency, per-agent throughput and outcomes across every concurrency level.
Combined throughput (all tokens) vs concurrent agents. Annotated peak with the concurrency level where it occurred.
Combined throughput (all tokens) vs concurrent agents. Annotated peak with the concurrency level where it occurred.
Per-agent throughput erosion as slots contend. Red stars mark burst-timing artifacts (queued ≥ 10s then generated in a burst — inflated values, not real throughput).
Per-agent throughput erosion as slots contend. Red stars mark burst-timing artifacts (queued ≥ 10s then generated in a burst — inflated values, not real throughput).
Scaling efficiency relative to the single-agent baseline (100%).
Scaling efficiency relative to the single-agent baseline (100%).
Time to first token — mean (solid) and max (dashed). Log scale. Gaps mean no visible content ever arrived within the 60s window.
Time to first token — mean (solid) and max (dashed). Log scale. Gaps mean no visible content ever arrived within the 60s window.
Total tokens per 60s run — all tokens vs visible content only.
Total tokens per 60s run — all tokens vs visible content only.
How each run of N agents finished: completed, timed out (60s budget), or errored.
How each run of N agents finished: completed, timed out (60s budget), or errored.
The concurrency trade-off on one chart — combined throughput rises while per-agent speed falls.
The concurrency trade-off on one chart — combined throughput rises while per-agent speed falls.

Raw data & downloads