There is no reliable GPU recommendation based only on the number of employees or a model’s parameter count. Capacity depends on the model, numerical precision, context length, output length, runtime and concurrent demand. Benchmark the actual workload before buying or reserving hardware.
Separate memory from throughput
The model must fit in available memory together with runtime overhead and the working state needed for requests. Longer contexts and more concurrent sequences can increase memory demand. A model that loads successfully may still fail under the intended traffic pattern.
As a rough starting estimate, weight storage is parameter count multiplied by bytes per parameter. This is not a complete sizing formula: runtime buffers, caches, quantisation metadata and other overhead still matter. Use it to reject obviously unsuitable configurations, then measure a working implementation.
Define the user experience
For an interactive assistant, measure time to the first useful response and total completion time. A background classification job may tolerate a queue. A voice interaction is sensitive to delays across speech recognition, model response and speech generation, not just GPU inference.
Create a workload sample with realistic prompt and output lengths. Include peak concurrency and sustained use. If retrieval adds several document passages, benchmark those contexts instead of a short “hello” request that flatters performance.
Include the rest of the platform
Storage throughput, model-loading time, network capacity and CPU work can become bottlenecks. Initial document ingestion may compete with user-facing inference. Decide which jobs can queue and which need reserved capacity.
Plan maintenance and failure headroom. If one accelerator fails, does the service stop, degrade or move to another endpoint? The answer changes the infrastructure budget. A theoretical maximum utilisation target can leave no room for recovery or unexpected demand.
Compare cost per useful task
- Measure accepted outputs, not only tokens per second.
- Include retries, evaluation and human correction time.
- Compare dedicated hardware with shared or managed options at realistic usage.
- Record power, support, licences and replacement assumptions for owned equipment.
- Re-test after model, runtime or context changes.
A smaller model that handles a bounded task well may be a better choice than a larger general model. Conversely, saving on hardware is unhelpful if users cannot trust the answers. Keep quality and performance results together in the capacity decision.
What to bring to a sizing discussion
Bring sample tasks, data boundaries, expected concurrent users, response targets, availability requirements and growth assumptions. If those are unknown, a benchmark engagement is a better first purchase than a large GPU cluster chosen from a generic specification sheet.
Sources & further reading
Primary references for the technical background and regional statements in this guide. Planning examples and checklists are Novacom’s practical guidance; examples are illustrative unless explicitly identified as project experience.