# chronos-1.5b Concurrency Benchmark

Full sweep of concurrent agents against LM Studio (`LM Studio's local endpoint`, model `chronos-1.5b`).

- **Harness**: `benchmark/benchmark.js` (agent loop, streaming, token-level timing), `benchmark/sweep.js` (sweep driver)
- **Config per run**: standard settings — 60s hard timeout, 12000-token total budget, `--output-cap 12000`, 32768-token context, temperature 0.9, agents run a multi-turn loop
- **Method**: 24 sequential runs at concurrency 1..24; results in `results/chronos-1.5b/<N>/`, aggregate in `sweep-summary.{csv,json}` in the same folder
- **Reasoning tokens**: chronos streams hidden thinking via `reasoning_content`; the share explodes with concurrency (~11% of tokens solo up to ~94% at 24 agents).

## Results (1–24 agents)

| agents | wall(s) | content_tok | reason_tok | total_tok | comb_tok/s | comb_all_tok/s | per-agent_tok/s | per-agent_all_tok/s | scale_eff | TTFT_mean | TTFT_max | ok | timeout | err |
|-------:|--------:|------------:|-----------:|----------:|-----------:|---------------:|----------------:|---------------------:|----------:|----------:|---------:|----|--------:|----:|
| 1 | 60.0 | 4703 | 565 | 5268 | 78 | 88 | 78.7 | 88.1 | 100% | 226ms | 226ms | 0 | 1 | 0 |
| 2 | 60.0 | 1042 | 8442 | 9484 | 17 | 158 | 8.7 | 79.2 | 180% | 130ms | 144ms | 0 | 2 | 0 |
| 3 | 60.0 | 1716 | 10283 | 11999 | 29 | 200 | 10.4 | 72.5 | 227% | 4.8s | 4.9s | 0 | 3 | 0 |
| 4 | 60.0 | 7843 | 5639 | 13482 | 131 | 225 | 38.7 | 66.9 | 256% | 7.0s | 7.0s | 0 | 4 | 0 |
| 5 | 60.0 | 6593 | 10091 | 16684 | 110 | 278 | 25.3 | 63.7 | 316% | 4.5s | 4.5s | 0 | 5 | 0 |
| 6 | 60.0 | 4546 | 13742 | 18288 | 76 | 305 | 14.7 | 59.0 | 347% | 8.3s | 8.4s | 0 | 6 | 0 |
| 7 | 60.0 | 5181 | 14391 | 19572 | 86 | 326 | 14.9 | 56.4 | 370% | 10.4s | 10.4s | 0 | 7 | 0 |
| 8 | 60.0 | 1589 | 18583 | 20172 | 26 | 336 | 4.7 | 52.8 | 382% | 11.0s | 11.1s | 0 | 8 | 0 |
| 9 | 60.0 | 5453 | 16614 | 22067 | 91 | 368 | 12.1 | 48.8 | 418% | 9.8s | 9.8s | 0 | 9 | 0 |
| 10 | 60.0 | 4150 | 18443 | 22593 | 69 | 376 | 8.6 | 47.1 | 427% | 12.0s | 12.0s | 0 | 10 | 0 |
| 11 | 60.0 | 3262 | 20593 | 23855 | 54 | 397 | 6.2 | 45.1 | 451% | 11.9s | 11.9s | 0 | 11 | 0 |
| 12 | 60.0 | 3545 | 20368 | 23913 | 59 | 398 | 6.4 | 43.2 | 452% | 13.8s | 13.9s | 0 | 12 | 0 |
| 13 | 60.0 | 6796 | 17121 | 23917 | 113 | 398 | 11.9 | 41.3 | 452% | 15.1s | 15.1s | 0 | 13 | 0 |
| 14 | 60.0 | 5642 | 19629 | 25271 | 94 | 421 | 8.9 | 40.0 | 478% | 14.8s | 14.9s | 0 | 14 | 0 |
| 15 | 60.0 | 4402 | 20902 | 25304 | 73 | 421 | 6.6 | 38.2 | 478% | 15.8s | 15.8s | 0 | 15 | 0 |
| 16 | 60.0 | 4577 | 21150 | 25727 | 76 | 429 | 6.5 | 36.8 | 488% | 16.3s | 16.4s | 0 | 16 | 0 |
| 17 | 60.0 | 2880 | 21882 | 24762 | 48 | 412 | 3.8 | 33.0 | 468% | 15.8s | 15.9s | 0 | 17 | 0 |
| 18 | 60.0 | 2194 | 23667 | 25861 | 37 | 431 | 2.7 | 32.4 | 490% | 15.7s | 15.7s | 0 | 18 | 0 |
| 19 | 60.0 | 2947 | 23277 | 26224 | 49 | 437 | 3.5 | 31.7 | 497% | 16.5s | 16.6s | 0 | 19 | 0 |
| 20 | 60.0 | 3081 | 22891 | 25972 | 51 | 433 | 3.6 | 30.7 | 492% | 17.7s | 17.8s | 0 | 20 | 0 |
| 21 | 60.0 | 1391 | 24762 | 26153 | 23 | 436 | 1.6 | 29.7 | 495% | 18.1s | 18.2s | 0 | 21 | 0 |
| 22 | 60.0 | 2508 | 24743 | 27251 | 42 | 454 | 2.6 | 28.9 | 516% | 17.1s | 17.2s | 0 | 22 | 0 |
| 23 | 60.0 | 2971 | 24555 | 27526 | 49 | 458 | 3.0 | 28.1 | 520% | 17.3s | 17.5s | 0 | 23 | 0 |
| 24 | 60.0 | 1577 | 25917 | 27494 | 26 | 458 | 1.5 | 26.8 | 520% | 17.3s | 17.4s | 0 | 24 | 0 |

## Key findings

- **The slowest model tested**: only **88 tok/s solo** (content 78), and combined all-token throughput tops out at just **458 tok/s at 24 agents** — far below every other model (llama 652, falcon3 839, qwen-deepseek 829, minicpm 983).
- **Impressive relative scaling** (520% at 23–24) — but that's mostly because the single-agent baseline is so low.
- **Reasoning dominates under load**: content-only throughput is tiny at high concurrency (26 tok/s at 24) — ~94% of output is hidden `reasoning_content` (25k–26k reasoning tokens per run at 22–24 agents).
- **TTFT degrades steadily**: 226ms solo → 17–18s at 20–24 agents.
- **Errors**: 0 across all 24 runs.

## Charts

Generated with `benchmark/charts.py` (matplotlib), saved as PNG in `results/chronos-1.5b/charts/`:

| Chart | File | What it highlights |
|---|---|---|
| Dashboard (2x2 overview) | `dashboard_1_24.png` | Whole sweep in one view |
| Combined throughput | `combined_throughput.png` | Slow absolute speed, big relative scaling |
| Per-agent throughput | `per_agent_throughput.png` | Per-agent slowdown under contention |
| Scaling efficiency | `scaling_efficiency.png` | 100% → 520% |
| Time to first token | `time_to_first_token.png` | 226ms → 17–18s latency growth |
| Total tokens generated | `total_tokens_generated.png` | Content vs reasoning stack per concurrency |
| Outcome breakdown | `outcome_breakdown.png` | Zero errors |
| Combined vs per-agent | `combined_vs_per_agent.png` | The concurrency trade-off on one chart |
| Content vs reasoning | `reasoning_vs_content.png` | Reasoning dominates under load |
