EMPVAI LabIT
EMPV / AI LabB

Applied R&D / capability evaluation

Local AI

A lab for understanding what a locally executed model can actually do on a specific workload, using reproducible evidence and explicit operating boundaries.

Repository statePublic Core 0.3.1 RC1 · hitch-local-ai==0.3.1 · schema 1.0.1.
Local modelsEvaluationEvidence
Problem

The model name is not a guarantee

Two models with similar benchmark scores can behave very differently when tools, constraints, structured outputs and real operational sequences are involved.

Direction

Test the workload before deployment

The path separates discovery, observed candidate, test, preflight, bounded run, evidence and comparison without turning one successful test into operational authorization.

01 / What exists

A versioned and verifiable public core.

The project exposes a Python package, CLI, schemas, job/test packs and output contracts for describing and verifying local-AI work reproducibly.

02 / What we are looking for

A useful decision measure, not a model leaderboard.

The question is not which model is best in general, but which model + runtime + workload combination satisfies a concrete contract with enough evidence.

03 / Potential

A qualification layer for private and on-prem AI.

If the thesis holds across domains, the value could be an independent layer between local models and enterprise workloads: test first, qualify separately, operate only within demonstrated boundaries.

EMPV / NOTE

The public surface is intentionally separated from private R&D state, host inventories and operational recipes.

AI Lab