Running open models on-premise, honestly
What it costs, what it saves, and the three cases where it is genuinely the right call.
On-premise deployment of open-weight models is asked about far more often than it is appropriate. It is worth being specific about when it earns its overhead.
The three cases where it does: a contractual or regulatory bar on third-party inference that will not be waived, a workload with steady high volume where hosted per-token pricing exceeds hardware amortisation, and latency requirements tight enough that a network round trip is the bottleneck. Institutional data policy is the most common of the three by a wide margin.
The costs are not only the GPUs. Somebody has to own patching, monitoring, capacity and the upgrade path when a better open model ships in six months. For a company without an existing platform team, that ownership is the real expense, and it is the one most likely to be left out of the business case.
Quality is no longer the objection it was two years ago. On narrow, well-specified tasks with good tuning data, current open models perform close enough to hosted frontier models that the difference is not what should decide the question. On open-ended reasoning across long documents, the gap is still real.
Our default recommendation is hosted with zero-retention endpoints, and on-premise when one of the three conditions above genuinely applies. When it does, we build it properly, and we make sure the client knows who owns the box on the day we leave.