A benchmark measures an abstract capability. A company needs an operational outcome.
A model can perform well on public benchmarks and still fail at the work we want it to perform. Production workloads include formats, tools, messy data, latency constraints, structured outputs and operational consequences that a benchmark does not represent.
The useful question is therefore not “which model is best?”, but “does this model + runtime + workload combination satisfy the contract we need?”