ai.2it.onl

personal blog · experiments in agentic programming

Benchmarking local LLMs under agentic load

I build agentic systems and I test the models they run on — at home, on consumer hardware, under real daily-driver load. This blog documents those experiments with full raw data, charts, and honest methodology.

Posts

01

Testing LLM Concurrency on Consumer Hardware

2026-08-02 · 24 sequential runs of 1 to 24 concurrent agents per model, 11 models, one RTX 5060. Who scales, who serializes, and how far a single consumer box can push agentic workloads.

+

More posts incoming

Next up: the strongest models from this sweep tested against real-world agentic use datasets — tool calling, long-context drift, quality vs token budget.