Private and on-premises AI · 2 min read

Choosing a server for private AI from measured workload

Size an AI server from model quality, context length, concurrency, latency and availability requirements—not from the model file size or employee count.

Define the workload

Pin model and precision, task, average and maximum context, concurrent requests, time to first token and generation rate, including any retrieval or speech services sharing the node.

Capture real context distributions and concurrency rather than using a maximum model context as if every request would consume it.

Budget the whole node

Account for weights, working memory and context cache on accelerators, plus system RAM, fast storage, CPU preprocessing, interconnect, cooling and recovery capacity.

Accelerator memory also holds runtime state and context cache, while ingestion, search and storage consume CPU, RAM and disk resources.

Benchmark the actual case

Run genuine prompt distributions on comparable hardware and measure quality, latency, throughput and memory under concurrency, long context, cold start and dependency failure.

Benchmark a pinned runtime with long prompts, simultaneous users, cold starts and dependency failures before approving the hardware order.

Use measured headroom

Derive reserve from peaks, updates and availability objectives, and define a scaling trigger such as queue depth or latency rather than buying blindly.

Set scaling thresholds from queue depth, latency and recovery requirements so future capacity decisions follow measured service behaviour.

pommeDeTerre

Need an estimate for your project?

Tell us about the project. We will break it into stages and explain the budget drivers.

View pricing