Adversarial capability evaluation
AI red teaming: the work frontier models still get wrong
Capability red teaming: operational tasks built specifically to defeat a frontier model, then scored dimension by dimension by a practitioner from that function, with written feedback on exactly where the model came apart. That account of how it failed is the finding, and it is the part that does not show up on a public benchmark.
A service we provide to AI labs, data partnership programmes and companies training or fine tuning their own models. Active in the micro1 Company Data Partnerships Program.
How we break a model
Failure is easy to produce and hard to produce usefully. A task that fails because it is badly specified teaches nobody anything. These fail because the work is genuinely difficult.
Long horizons and branching decisions
Real operational work does not resolve in one turn. It runs across many steps, changes direction on what it finds, and requires holding earlier context that has since gone stale. Models that look strong on single turn tasks come apart here, and that is where we aim.
Ambiguity that a practitioner resolves silently
Experts constantly fill gaps the brief never mentions, using judgement they would struggle to write down. We build tasks around exactly those gaps, because they are where a confident, plausible, wrong answer is most likely and most expensive.
Scored and explained, not just failed
A failure is only useful if someone can say why it happened. Each run is scored on every dimension by a practitioner from that function, who writes up what the model got right, where it went wrong, and what it appeared to misunderstand. That write up is what turns a low score into something you can act on.
What never leaves our side of the table
This work contributes operational expertise, not data about the people we work for. These are operating rules, applied before a workflow is written rather than checked afterwards.
Scrubbed or fictionalised from the start
Where a real workflow touches anything sensitive, it is rebuilt as a scrubbed or fictionalised version before any work begins. Not redacted afterwards.
Never client or customer material
No client, customer or employer confidential information. No proprietary material. Nothing covered by an NDA. This is the same commitment that governs our software delivery work.
No personal data, no credentials
No private or personal information about anyone, and no passwords, API keys or credentials in any prompt, file or environment.
Only tools we can properly grant
We work only in tools where the operator personally holds access and can safely extend that access to an AI system. Where a real process would run through a company system, we substitute a personal or fictionalised equivalent.
Never on unauthorised devices
Agent work does not run on any device we are not authorised to run it on.
The functions we can cover
Breadth from running these functions as a working software business, not from staffing up for a programme. We are also among the top one percent of official FlutterFlow partners globally, which is what betting early on a shift looks like when it pays off.
How an engagement runs
Scope
We agree the functions, the difficulty bar and the rubric dimensions with you before anything is built.
Build
Workflows are authored by the people who do the work, then scrubbed or fictionalised and checked against the data rules above.
Run
Each workflow is executed across the model set under matched conditions, so any difference that surfaces is attributable to the model rather than the setup.
Score and compare
Every run is graded on the agreed dimensions, then models are ranked against each other so the per dimension picture is legible.
Deliver
You receive the workflows, the runs, the scores and the comparison, in the structure your pipeline expects.
Questions partners ask
Is this safety red teaming or jailbreak testing?
Neither. This is capability red teaming: finding operational work that frontier models cannot yet do reliably. We do not test for harmful outputs, alignment failures, guardrail bypasses or prompt injection resistance, and we would point you elsewhere for that.
What counts as a successful adversarial task?
One that current frontier models get wrong in a way a practitioner from that function can explain. If everything passes it, the task is too easy to be informative. If it fails because the brief was ambiguous or the tooling was broken, that is a badly specified task rather than a hard one, and it does not ship.
Do you report the failures, or just the scores?
Both. A score tells you a model lost; the write up tells you where and how, which is the part that changes what you train on next. Every task ships with its runs, the per dimension scoring, and the practitioner notes behind each score.
Who does the work?
The same practitioners who do the job for clients. The person who writes a QA workflow is a QA practitioner, and the person scoring a software engineering run is an engineer. That is the point: whoever judges the run has to know what good looks like in that function, or the score and the feedback are not worth much.
Do you ever use real client data?
No. Client confidentiality and intellectual property are non negotiable, and they do not bend for this programme. Where a workflow is drawn from real operational experience, it is rebuilt as a scrubbed or fictionalised version before the work starts, not cleaned up afterwards.
Can I join as an AI trainer, or can my company supply this to you?
Neither, and it is worth being direct about it. This page describes a service we sell to companies, not a programme you can join and not work we subcontract out. We are not recruiting AI trainers, evaluators or annotators, and we do not buy this capacity from other agencies, because the whole point is that the people writing and scoring the workflows are our own practitioners doing that job for clients. If you are looking for a role at Empiric, our open positions are at /career.
How do we start?
A short call to scope one batch. We would rather prove the work on a small, well defined set than talk about volume before you have seen the quality.
Related
Send us a function you think is already solved
The most useful first batch is usually the work a team assumes models have covered. Tell us the function and we will build the tasks that test it.
Start a conversationSCOPE A BATCH
Tell us the functions you care about and the rubric you score on, and we will come back with what a first batch would look like.




