Falcon3 1B Instruct
server model id:
falcon3-1b-instructnon-thinkingKV F16ctx 8k1B params
Fastest pure-content scorer — 839 tok/s at 13 agents, then rolls off as the 8k context saturates.
Key findings
- Clean, all-content output: 0 reasoning tokens everywhere.
- Strong scaling, with a peak and then a rolloff: combined throughput climbs from 187 tok/s solo to 839 tok/s at 13 agents (4.5x), holds ~730–840 through 18, then *declines* to 522 at 24 — unlike the other models, which plateaued. Under heavy contention with a small context/output cap, throughput rolls off.
- Latency stays sane: TTFT 72ms–2.7s across the whole sweep (no 10s+ spikes like the thinking models).
- Per-agent speed: falls from ~200 tok/s solo to ~24 at 24 agents.
- Errors: 0 across all 24 runs.
Sweep — total concurrency 1–24
| agents | wall (s) | content | reason | total | comb tok/s | comb all | per-agent | per-agent all | scale % | TTFT mean | TTFT max | ok | timeout | err |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 60.0 | 11,214 | 0 | 11,214 | 187 | 187 | 202.5 | 206.8 | 100% | 1.0s | 1.0s | 0 | 1 | 0 |
| 2 | 60.0 | 16,684 | 0 | 16,684 | 278 | 278 | 146.8 | 159.2 | 149% | 72ms | 91ms | 0 | 2 | 0 |
| 3 | 60.0 | 20,889 | 0 | 20,889 | 348 | 348 | 126.6 | 135.8 | 186% | 1.2s | 1.2s | 0 | 3 | 0 |
| 4 | 60.0 | 24,587 | 0 | 24,587 | 409 | 409 | 116.5 | 124.7 | 219% | 1.5s | 1.6s | 0 | 4 | 0 |
| 5 | 60.0 | 28,624 | 0 | 28,624 | 477 | 477 | 101.1 | 106.2 | 255% | 1.9s | 1.9s | 0 | 5 | 0 |
| 6 | 60.1 | 35,788 | 0 | 35,788 | 596 | 596 | 108.7 | 114.0 | 319% | 748ms | 785ms | 0 | 6 | 0 |
| 7 | 60.1 | 38,712 | 0 | 38,712 | 645 | 645 | 99.1 | 105.1 | 345% | 1.0s | 1.0s | 0 | 7 | 0 |
| 8 | 60.1 | 40,292 | 0 | 40,292 | 671 | 671 | 92.4 | 99.0 | 359% | 1.0s | 1.1s | 0 | 8 | 0 |
| 9 | 60.1 | 44,005 | 0 | 44,005 | 733 | 733 | 88.3 | 97.6 | 392% | 1.8s | 1.9s | 0 | 9 | 0 |
| 10 | 60.1 | 43,719 | 0 | 43,719 | 728 | 728 | 77.0 | 86.1 | 389% | 1.3s | 1.3s | 0 | 10 | 0 |
| 11 | 60.1 | 46,032 | 0 | 46,032 | 766 | 766 | 78.4 | 89.0 | 410% | 1.3s | 1.4s | 0 | 11 | 0 |
| 12 | 60.1 | 48,643 | 0 | 48,643 | 810 | 810 | 71.8 | 82.8 | 433% | 2.0s | 2.1s | 0 | 12 | 0 |
| 13 peak | 60.1 | 50,382 | 0 | 50,382 | 839 | 839 | 70.9 | 82.8 | 449% | 677ms | 704ms | 1 | 12 | 0 |
| 14 | 60.1 | 47,925 | 0 | 47,925 | 798 | 798 | 61.7 | 73.4 | 427% | 1.3s | 1.4s | 0 | 14 | 0 |
| 15 | 60.1 | 48,906 | 0 | 48,906 | 814 | 814 | 59.3 | 71.5 | 435% | 1.9s | 2.0s | 0 | 15 | 0 |
| 16 | 60.1 | 49,630 | 0 | 49,630 | 826 | 826 | 57.4 | 70.6 | 442% | 1.7s | 1.7s | 0 | 16 | 0 |
| 17 | 60.1 | 47,629 | 0 | 47,629 | 793 | 793 | 50.5 | 63.5 | 424% | 2.7s | 2.7s | 0 | 17 | 0 |
| 18 | 60.1 | 48,878 | 0 | 48,878 | 814 | 814 | 48.8 | 62.8 | 435% | 596ms | 624ms | 0 | 18 | 0 |
| 19 | 60.1 | 43,822 | 0 | 43,822 | 730 | 730 | 42.5 | 59.6 | 390% | 2.0s | 2.2s | 0 | 19 | 0 |
| 20 | 60.1 | 42,649 | 0 | 42,649 | 710 | 710 | 38.7 | 56.9 | 380% | 1.7s | 1.7s | 0 | 20 | 0 |
| 21 | 60.1 | 41,359 | 0 | 41,359 | 689 | 689 | 35.4 | 53.3 | 368% | 1.5s | 1.5s | 0 | 21 | 0 |
| 22 | 60.1 | 37,553 | 0 | 37,553 | 625 | 625 | 30.5 | 50.8 | 334% | 587ms | 634ms | 0 | 22 | 0 |
| 23 | 60.1 | 35,169 | 0 | 35,169 | 586 | 586 | 27.5 | 47.4 | 313% | 827ms | 860ms | 0 | 23 | 0 |
| 24 | 60.1 | 31,331 | 0 | 31,331 | 522 | 522 | 23.6 | 43.6 | 279% | 388ms | 555ms | 0 | 24 | 0 |
* burst artifact — agents queued ~11–16s then generated in a short burst; the per-agent timer inflates the number. Trust combined throughput and TTFT columns. Peak row = highest combined (all-token) throughput.
Charts








Raw data & downloads
- report.md5 KB
- sweep-summary.csv2 KB
- sweep-summary.json11 KB
- combined_throughput.png120 KB · original matplotlib export
- combined_vs_per_agent.png105 KB · original matplotlib export
- dashboard_1_24.png165 KB · original matplotlib export
- outcome_breakdown.png57 KB · original matplotlib export
- per_agent_throughput.png113 KB · original matplotlib export
- scaling_efficiency.png76 KB · original matplotlib export
- time_to_first_token.png57 KB · original matplotlib export
- total_tokens_generated.png65 KB · original matplotlib export