# falcon3-1b-instruct Concurrency Benchmark

Full sweep of concurrent agents against LM Studio (`LM Studio's local endpoint`, model `falcon3-1b-instruct`).

- **Harness**: `benchmark/benchmark.js` (agent loop, streaming, token-level timing), `benchmark/sweep.js` (sweep driver)
- **Config per run**: 60s hard timeout per agent, temperature 0.9, agents run a multi-turn loop
- **Context / output caps**: falcon3-1b-instruct has only an **8k context window** (vs 32k for the other models). Runs used `--output-cap 2000` (max output per single request) with history trimmed to a rolling window (system + task + latest output + continue prompt) so every agent stays busy the full 60s without overflowing context. Total budget `--max-tokens 12000` (safety net, not binding in these runs).
- **Method**: 24 sequential runs at concurrency 1..24; results in `results/falcon3-1b-instruct/<N>/`, aggregate in `sweep-summary.{csv,json}` in the same folder
- **Reasoning tokens**: 0 across every run — falcon3-1b-instruct is a clean non-thinking model.

## Results (1–24 agents)

| agents | wall(s) | content_tok | reason_tok | total_tok | comb_tok/s | comb_all_tok/s | per-agent_tok/s | per-agent_all_tok/s | scale_eff | TTFT_mean | TTFT_max | ok | timeout | err |
|-------:|--------:|------------:|-----------:|----------:|-----------:|---------------:|----------------:|---------------------:|----------:|----------:|---------:|----|--------:|----:|
| 1 | 60.0 | 11214 | 0 | 11214 | 187 | 187 | 202.5 | 206.8 | 100% | 1.0s | 1.0s | 0 | 1 | 0 |
| 2 | 60.0 | 16684 | 0 | 16684 | 278 | 278 | 146.8 | 159.2 | 149% | 72ms | 91ms | 0 | 2 | 0 |
| 3 | 60.0 | 20889 | 0 | 20889 | 348 | 348 | 126.6 | 135.8 | 186% | 1.2s | 1.2s | 0 | 3 | 0 |
| 4 | 60.0 | 24587 | 0 | 24587 | 409 | 409 | 116.5 | 124.7 | 219% | 1.5s | 1.6s | 0 | 4 | 0 |
| 5 | 60.0 | 28624 | 0 | 28624 | 477 | 477 | 101.1 | 106.2 | 255% | 1.9s | 1.9s | 0 | 5 | 0 |
| 6 | 60.1 | 35788 | 0 | 35788 | 596 | 596 | 108.7 | 114.0 | 319% | 748ms | 785ms | 0 | 6 | 0 |
| 7 | 60.1 | 38712 | 0 | 38712 | 645 | 645 | 99.1 | 105.1 | 345% | 1.0s | 1.0s | 0 | 7 | 0 |
| 8 | 60.1 | 40292 | 0 | 40292 | 671 | 671 | 92.4 | 99.0 | 359% | 1.0s | 1.1s | 0 | 8 | 0 |
| 9 | 60.1 | 44005 | 0 | 44005 | 733 | 733 | 88.3 | 97.6 | 392% | 1.8s | 1.9s | 0 | 9 | 0 |
| 10 | 60.1 | 43719 | 0 | 43719 | 728 | 728 | 77.0 | 86.1 | 389% | 1.3s | 1.3s | 0 | 10 | 0 |
| 11 | 60.1 | 46032 | 0 | 46032 | 766 | 766 | 78.4 | 89.0 | 410% | 1.3s | 1.4s | 0 | 11 | 0 |
| 12 | 60.1 | 48643 | 0 | 48643 | 810 | 810 | 71.8 | 82.8 | 433% | 2.0s | 2.1s | 0 | 12 | 0 |
| 13 | 60.1 | 50382 | 0 | 50382 | 839 | 839 | 70.9 | 82.8 | 449% | 677ms | 704ms | 1 | 12 | 0 |
| 14 | 60.1 | 47925 | 0 | 47925 | 798 | 798 | 61.7 | 73.4 | 427% | 1.3s | 1.4s | 0 | 14 | 0 |
| 15 | 60.1 | 48906 | 0 | 48906 | 814 | 814 | 59.3 | 71.5 | 435% | 1.9s | 2.0s | 0 | 15 | 0 |
| 16 | 60.1 | 49630 | 0 | 49630 | 826 | 826 | 57.4 | 70.6 | 442% | 1.7s | 1.7s | 0 | 16 | 0 |
| 17 | 60.1 | 47629 | 0 | 47629 | 793 | 793 | 50.5 | 63.5 | 424% | 2.7s | 2.7s | 0 | 17 | 0 |
| 18 | 60.1 | 48878 | 0 | 48878 | 814 | 814 | 48.8 | 62.8 | 435% | 596ms | 624ms | 0 | 18 | 0 |
| 19 | 60.1 | 43822 | 0 | 43822 | 730 | 730 | 42.5 | 59.6 | 390% | 2.0s | 2.2s | 0 | 19 | 0 |
| 20 | 60.1 | 42649 | 0 | 42649 | 710 | 710 | 38.7 | 56.9 | 380% | 1.7s | 1.7s | 0 | 20 | 0 |
| 21 | 60.1 | 41359 | 0 | 41359 | 689 | 689 | 35.4 | 53.3 | 368% | 1.5s | 1.5s | 0 | 21 | 0 |
| 22 | 60.1 | 37553 | 0 | 37553 | 625 | 625 | 30.5 | 50.8 | 334% | 587ms | 634ms | 0 | 22 | 0 |
| 23 | 60.1 | 35169 | 0 | 35169 | 586 | 586 | 27.5 | 47.4 | 313% | 827ms | 860ms | 0 | 23 | 0 |
| 24 | 60.1 | 31331 | 0 | 31331 | 522 | 522 | 23.6 | 43.6 | 279% | 388ms | 555ms | 0 | 24 | 0 |

*Every run lasted the full 60s (agents were kept busy via continuation turns); `timeout` in the table is the intended 60s cap, not an error.*

## Key findings

- **Clean, all-content output**: 0 reasoning tokens everywhere.
- **Strong scaling, with a peak and then a rolloff**: combined throughput climbs from 187 tok/s solo to **839 tok/s at 13 agents (4.5x)**, holds ~730–840 through 18, then *declines* to 522 at 24 — unlike the other models, which plateaued. Under heavy contention with a small context/output cap, throughput rolls off.
- **Latency stays sane**: TTFT 72ms–2.7s across the whole sweep (no 10s+ spikes like the thinking models).
- **Per-agent speed**: falls from ~200 tok/s solo to ~24 at 24 agents.
- **Errors**: 0 across all 24 runs.

## Charts

Generated with `benchmark/charts.py` (matplotlib), saved as PNG in `results/falcon3-1b-instruct/charts/`:

| Chart | File | What it highlights |
|---|---|---|
| Dashboard (2x2 overview) | `dashboard_1_24.png` | Whole sweep in one view |
| Combined throughput | `combined_throughput.png` | 187 → 839 tok/s peak, rolloff past 18 |
| Per-agent throughput | `per_agent_throughput.png` | Per-agent slowdown under contention |
| Scaling efficiency | `scaling_efficiency.png` | 100% → 449% peak, then decline |
| Time to first token | `time_to_first_token.png` | 72ms–2.7s latency profile |
| Total tokens generated | `total_tokens_generated.png` | ~11k → 50k tokens/60s |
| Outcome breakdown | `outcome_breakdown.png` | Zero errors |
| Combined vs per-agent | `combined_vs_per_agent.png` | The concurrency trade-off on one chart |
