LLM evaluation and benchmarking
LLM evaluation and benchmarking that says which model is strong where
A single leaderboard number tells you almost nothing about whether a model can do your work. We run the same operational task across a model set under matched conditions and score every run against the agreed rubric, so the result is a per dimension comparison with the reasoning attached rather than one figure.
A service we provide to AI labs, data partnership programmes and companies training or fine tuning their own models. Active in the micro1 Company Data Partnerships Program.
How we evaluate
Comparability is the whole game. If the prompt, the context and the harness are not held constant, the differences you measure are differences in your setup.
Matched conditions, every model
The same task, the same context, the same harness, the same acceptance criteria. When something separates two models, that separation is attributable to the models. This is the part most internal evaluations get wrong, because the prompt drifts as the team learns what works.
Judged by people who do that work
Every run is scored by a practitioner who does that work for a living, dimension by dimension, with written notes on what the model handled well and where it went wrong. A number alone tells you a model lost. The notes tell you what it misread, which is the part that changes what you do next.
Ranked per dimension, not averaged
The dimensions are agreed with you before work starts and applied identically across the set. Averaging them away hides the finding: models rarely fail uniformly, they fail in a shape. We report the shape, and the notes that explain it.
What never leaves our side of the table
This work contributes operational expertise, not data about the people we work for. These are operating rules, applied before a workflow is written rather than checked afterwards.
Scrubbed or fictionalised from the start
Where a real workflow touches anything sensitive, it is rebuilt as a scrubbed or fictionalised version before any work begins. Not redacted afterwards.
Never client or customer material
No client, customer or employer confidential information. No proprietary material. Nothing covered by an NDA. This is the same commitment that governs our software delivery work.
No personal data, no credentials
No private or personal information about anyone, and no passwords, API keys or credentials in any prompt, file or environment.
Only tools we can properly grant
We work only in tools where the operator personally holds access and can safely extend that access to an AI system. Where a real process would run through a company system, we substitute a personal or fictionalised equivalent.
Never on unauthorised devices
Agent work does not run on any device we are not authorised to run it on.
The functions we can cover
Breadth from running these functions as a working software business, not from staffing up for a programme. We are also among the top one percent of official FlutterFlow partners globally, which is what betting early on a shift looks like when it pays off.
How an engagement runs
Scope
We agree the functions, the difficulty bar and the rubric dimensions with you before anything is built.
Build
Workflows are authored by the people who do the work, then scrubbed or fictionalised and checked against the data rules above.
Run
Each workflow is executed across the model set under matched conditions, so any difference that surfaces is attributable to the model rather than the setup.
Score and compare
Every run is graded on the agreed dimensions, then models are ranked against each other so the per dimension picture is legible.
Deliver
You receive the workflows, the runs, the scores and the comparison, in the structure your pipeline expects.
Questions partners ask
How is this different from a public benchmark?
Public benchmarks are saturated, contaminated by training data, and built from tasks that look nothing like operational work. Ours are authored from functions we run as a working business, they have never been published, and they are hard enough that current frontier models struggle with them.
Can you evaluate against our own rubric?
Yes, and we would rather. If you already have dimensions your pipeline expects, we work to those instead of ours. Where you do not, we agree the rubric with you before anything is built so the results are usable on arrival.
Which models do you cover?
The set is agreed per engagement, and we run whatever is in it under identical conditions. We do not publish which models we have evaluated or the results, and would extend the same discretion to your programme.
Who does the work?
The same practitioners who do the job for clients. The person who writes a QA workflow is a QA practitioner, and the person scoring a software engineering run is an engineer. That is the point: whoever judges the run has to know what good looks like in that function, or the score and the feedback are not worth much.
Do you ever use real client data?
No. Client confidentiality and intellectual property are non negotiable, and they do not bend for this programme. Where a workflow is drawn from real operational experience, it is rebuilt as a scrubbed or fictionalised version before the work starts, not cleaned up afterwards.
Can I join as an AI trainer, or can my company supply this to you?
Neither, and it is worth being direct about it. This page describes a service we sell to companies, not a programme you can join and not work we subcontract out. We are not recruiting AI trainers, evaluators or annotators, and we do not buy this capacity from other agencies, because the whole point is that the people writing and scoring the workflows are our own practitioners doing that job for clients. If you are looking for a role at Empiric, our open positions are at /career.
How do we start?
A short call to scope one batch. We would rather prove the work on a small, well defined set than talk about volume before you have seen the quality.
Related
Bring us the rubric, we will bring the tasks
One well defined batch is a better conversation than a capability deck. Tell us the functions you care about and the dimensions you score on.
Start a conversationSCOPE A BATCH
Tell us the functions you care about and the rubric you score on, and we will come back with what a first batch would look like.




