# qwen-deepseek-1.5b-agentic-distill Concurrency Benchmark

Full sweep of concurrent agents against LM Studio (`LM Studio's local endpoint`, model `qwen-deepseek-1.5b-agentic-distill`).

- **Harness**: `benchmark/benchmark.js` (agent loop, streaming, token-level timing), `benchmark/sweep.js` (sweep driver)
- **Config per run**: standard settings used for the other models — 60s hard timeout, 12000-token total budget, `--output-cap 12000`, 32768-token context, temperature 0.9, agents run a multi-turn loop
- **Method**: 24 sequential runs at concurrency 1..24; results in `results/qwen-deepseek-1.5b-agentic-distill/<N>/`, aggregate in `sweep-summary.{csv,json}` in the same folder
- **Reasoning tokens**: this agentic distill streams hidden thinking via `reasoning_content` (~35–65% of all tokens, growing with load).

## Results (1–24 agents)

| agents | wall(s) | content_tok | reason_tok | total_tok | comb_tok/s | comb_all_tok/s | per-agent_tok/s | per-agent_all_tok/s | scale_eff | TTFT_mean | TTFT_max | ok | timeout | err |
|-------:|--------:|------------:|-----------:|----------:|-----------:|---------------:|----------------:|---------------------:|----------:|----------:|---------:|----|--------:|----:|
| 1 | 60.0 | 6387 | 3528 | 9915 | 106 | 165 | 654.1* | 459.7* | 100% | 4.2s | 4.2s | 0 | 1 | 0 |
| 2 | 60.0 | 12566 | 3544 | 16110 | 209 | 268 | 929.8* | 469.4* | 162% | 10.6s | 18.2s | 0 | 2 | 0 |
| 3 | 60.0 | 11059 | 5753 | 16812 | 184 | 280 | 106.4 | 153.3 | 170% | 9.2s | 11.0s | 0 | 3 | 0 |
| 4 | 60.0 | 5889 | 16088 | 21977 | 98 | 366 | 70.8 | 214.6* | 222% | 10.6s | 15.5s | 0 | 4 | 0 |
| 5 | 60.0 | 11680 | 10495 | 22175 | 195 | 369 | 124.8 | 145.4 | 224% | 20.7s | 47.6s | 0 | 5 | 0 |
| 6 | 60.1 | 11444 | 10260 | 21704 | 190 | 361 | 59.6 | 70.9 | 219% | 14.1s | 27.6s | 0 | 6 | 0 |
| 7 | 60.0 | 11965 | 10740 | 22705 | 199 | 378 | 58.0 | 83.7 | 229% | 18.5s | 28.5s | 0 | 7 | 0 |
| 8 | 60.0 | 14271 | 8965 | 23236 | 238 | 387 | 59.0 | 70.8 | 235% | 20.4s | 45.2s | 0 | 8 | 0 |
| 9 | 60.0 | 16308 | 10225 | 26533 | 272 | 442 | 63.7 | 77.7 | 268% | 17.0s | 23.3s | 0 | 9 | 0 |
| 10 | 60.0 | 12477 | 15464 | 27941 | 208 | 465 | 53.3 | 85.9 | 282% | 21.5s | 31.5s | 0 | 10 | 0 |
| 11 | 60.0 | 14417 | 14907 | 29324 | 240 | 488 | 47.1 | 67.7 | 296% | 24.6s | 37.1s | 0 | 11 | 0 |
| 12 | 60.0 | 12716 | 15387 | 28103 | 212 | 468 | 52.6 | 77.4 | 284% | 26.4s | 57.9s | 0 | 12 | 0 |
| 13 | 60.0 | 15027 | 15888 | 30915 | 250 | 515 | 45.6 | 66.1 | 312% | 30.4s | 48.5s | 0 | 13 | 0 |
| 14 | 60.0 | 14590 | 14377 | 28967 | 243 | 482 | 41.6 | 54.2 | 292% | 27.4s | 49.6s | 0 | 14 | 0 |
| 15 | 60.0 | 10288 | 20860 | 31148 | 171 | 519 | 32.2 | 55.9 | 315% | 32.1s | 49.9s | 0 | 15 | 0 |
| 16 | 60.0 | 17958 | 14484 | 32442 | 299 | 540 | 39.6 | 49.8 | 327% | 31.2s | 49.8s | 0 | 16 | 0 |
| 17 | 60.0 | 12847 | 18512 | 31359 | 214 | 522 | 34.2 | 49.2 | 316% | 35.9s | 48.0s | 0 | 17 | 0 |
| 18 | 60.0 | 14861 | 17338 | 32199 | 248 | 536 | 32.7 | 46.6 | 325% | 30.2s | 37.0s | 0 | 18 | 0 |
| 19 | 60.0 | 15570 | 17254 | 32824 | 259 | 547 | 37.0 | 54.4 | 332% | 35.2s | 52.9s | 0 | 19 | 0 |
| 20 | 60.0 | 13697 | 18184 | 31881 | 228 | 531 | 29.7 | 45.6 | 322% | 33.1s | 47.8s | 0 | 20 | 0 |
| 21 | 60.0 | 16053 | 33731 | 49784 | 267 | 829 | 38.0 | 65.1 | 502% | 23.4s | 50.2s | 0 | 21 | 0 |
| 22 | 60.0 | 17576 | 26033 | 43609 | 293 | 726 | 71.6 | 59.2 | 440% | 26.8s | 41.0s | 0 | 22 | 0 |
| 23 | 60.0 | 17311 | 29157 | 46468 | 288 | 774 | 50.2 | 63.8 | 469% | 27.9s | 51.6s | 0 | 23 | 0 |
| 24 | 60.0 | 14588 | 31243 | 45831 | 243 | 763 | 39.6 | 63.7 | 462% | 27.8s | 54.5s | 0 | 24 | 0 |

*\* per-agent tok/s values marked * are burst-timing artifacts when agents wait in the queue then generate in bursts.*

## Key findings

- **Strong late-game scaling**: combined throughput climbs steadily to **829 tok/s at 21 agents (5.0x)** and holds 726–774 at 22–24 — the highest scaling of any model tested so far. A jump at 21+ likely reflects LM Studio batching kicking in more aggressively.
- **Visible content is modest**: content-only throughput peaks at only ~299 tok/s (16 agents) — over half the budget is hidden `reasoning_content` (31k–34k reasoning tokens per run at high concurrency).
- **Very slow TTFT — the worst profile**: 4.2s solo, and **20–36s mean under load**, with max values up to 58s (near the full 60s window). Agents frequently wait half the run before their first token. Bad for latency-critical use.
- **Per-agent speed**: ~30–50 tok/s (all-token) at high concurrency.
- **Errors**: 0 across all 24 runs.

## Charts

Generated with `benchmark/charts.py` (matplotlib), saved as PNG in `results/qwen-deepseek-1.5b-agentic-distill/charts/`:

| Chart | File | What it highlights |
|---|---|---|
| Dashboard (2x2 overview) | `dashboard_1_24.png` | Whole sweep in one view |
| Combined throughput | `combined_throughput.png` | Content vs content+reasoning curves |
| Per-agent throughput | `per_agent_throughput.png` | Per-agent slowdown under contention |
| Scaling efficiency | `scaling_efficiency.png` | 100% → 502% peak |
| Time to first token | `time_to_first_token.png` | 4s–58s reasoning-prefill latency |
| Total tokens generated | `total_tokens_generated.png` | Content vs reasoning stack per concurrency |
| Outcome breakdown | `outcome_breakdown.png` | Zero errors |
| Combined vs per-agent | `combined_vs_per_agent.png` | The concurrency trade-off on one chart |
| Content vs reasoning | `reasoning_vs_content.png` | Hidden reasoning budget share per run |
