EMPVResearch NotesIT
EMPV / RESEARCH NOTESNOTE / 001

TECH NOTE / LOCAL AI

How to evaluate a local LLM before using it in production

A model name and a benchmark score are not enough. Before a local model receives production tasks, workload, runtime, constraints and failure modes need to be tested together.

Published2026-09-29 · 6 min
Local LLMEvaluationOn-premisePrivate AI
01 / The problem

A benchmark measures an abstract capability. A company needs an operational outcome.

A model can perform well on public benchmarks and still fail at the work we want it to perform. Production workloads include formats, tools, messy data, latency constraints, structured outputs and operational consequences that a benchmark does not represent.

The useful question is therefore not “which model is best?”, but “does this model + runtime + workload combination satisfy the contract we need?”

02 / The contract

Before testing, success has to be made explicit.

A useful test starts from an observable contract: allowed inputs, expected output, available tools, time limits, acceptable errors and conditions that must stop the flow.

Without those boundaries, an apparently correct result can become a false authorization.

  • define the workload before choosing the model
  • separate output quality from tool-use capability
  • record failure modes alongside successful cases
  • keep reproducible evidence of every test
03 / The test

Model and runtime need to be evaluated together.

Quantization, context window, prompt format, tool calling, available memory and runtime libraries can change model behaviour. The same model can therefore produce different outcomes in different environments.

A useful test records configuration, version, input, output, timing and conditions so the result can be repeated and compared.

04 / Authorization

Passing a test is not permanent permission.

A successful test demonstrates a capability under specific conditions. It does not prove that the model can operate without limits on different inputs, new data or additional tools.

For sensitive workloads it is useful to separate qualification from authorization: demonstrate the capability first, then authorize a narrow and observable operating boundary.

05 / What changes

Model selection becomes an engineering decision, not a leaderboard.

This approach can make smaller models useful when they satisfy a specific contract, while preventing a large model from receiving critical work simply because it performs well on average.

The goal is not to find the best model in general. It is to know what a model can do, in which environment, with what evidence and within which limits.

EMPV / TAKEAWAY

Before deployment, qualify the workload. The model comes after.

Research Notes