Benchmarking local LLMs under agentic load
I build agentic systems and I test the models they run on — at home, on consumer hardware, under real daily-driver load. This blog documents those experiments with full raw data, charts, and honest methodology.
Posts
01
→
+
→
Testing LLM Concurrency on Consumer Hardware
2026-08-02 · 24 sequential runs of 1 to 24 concurrent agents per model, 11 models, one RTX 5060. Who scales, who serializes, and how far a single consumer box can push agentic workloads.
More posts incoming
Next up: the strongest models from this sweep tested against real-world agentic use datasets — tool calling, long-context drift, quality vs token budget.