Home / CEST accuracy study

Ongoing study, 10 cases encoded

We ran HMRC's CEST tool against decided IR35 cases

CEST returns one word and no reasoning. So we took decided IR35 tribunal cases, encoded the judges' findings of fact as answers to HMRC's own questions, and ran the official decision tables over them. Every encoding is published below with a paragraph reference, so you can disagree with a specific answer rather than with the conclusion.

What we found

6

agreed with the tribunal

2

got the answer wrong

2

refused to decide

Across ten cases with a single clean holding, CEST reached the wrong answer in two, and produced no answer at all in two of them. Those refusals are cases that were contested hard enough to reach a tribunal, which is exactly the situation where a contractor most needs a determination.

Where it got the answer wrong

Note the directions. CEST is not simply leaning one way in this sample: it pushed a contractor the tribunal found outside into inside, and a contractor the tribunal found inside into outside. A wrong answer is expensive either way, and the second kind is the one a contractor would have relied on.

An undetermined result is not a neutral outcome. HMRC will not stand behind a result its own tool declined to reach, so the contractor is left carrying the status risk with nothing to point at.

What this sample does not show

2 refusals in 10 cases is a rate of 20%, but the 95% confidence interval on that figure runs from 5.7% to 51%. HMRC's own published figure, across 1,257,571 uses between September 2021 and June 2023, is 22%. That interval contains HMRC's figure, so this sample cannot show that CEST refuses more often on litigated cases than it does generally. Anyone quoting 20% as a finding, including us, would be over-reading ten cases. The named refusals stand on their own. The rate does not, yet.

Every case, and what each engine said

CaseTribunal heldCEST saidOur engine
Kaye Adams[2024] UKFTT 00037 (TC)Outside IR35Unable to determine46/100, inside (wrong)
Alan Parry[2022] UKFTT 00194 (TC)Inside IR35Unable to determine40/100, inside
Gary Hughes, design engineer at JCB[2011] UKFTT 411 (TC)Outside IR35Outside IR3557/100, outside
George Mantides, locum urologist, Royal Berkshire engagement[2025] UKUT 00124 (TCC)Inside IR35Inside IR3546/100, inside
Philip Winfield, database software developer at GlaxoSmithKline[2011] UKFTT 454 (TC)Outside IR35Outside IR3569/100, outside
Elaine Richardson, IT consultant at Vertex Data Services[2011] UKFTT 313 (TC)Outside IR35Outside IR3569/100, outside
Mark Fitzpatrick, design engineer at Airbus[2011] UKFTT 35 (TC)Outside IR35Outside IR3551/100, outside
Novak Brajkovic, IT contractor at Avecia[2010] UKFTT 150 (TC)Outside IR35Inside IR3547/100, inside (wrong)
Phil Thompson, football pundit at Sky[2024] UKFTT 38 (TC)Inside IR35Outside IR3543/100, inside
Stuart Barnes, rugby commentator at Sky[2024] UKUT 262 (TCC)Inside IR35Inside IR3532/100, inside

Our engine scores 0 to 100 as a probability of being outside IR35, treated as predicting "outside" at 50 or above.

Our own tool got two wrong

Two of them went against our own scoring engine, not just against CEST. In Kaye Adams we scored 46 against a threshold of 50, predicting inside, and the tribunal held Outside IR35. In Novak Brajkovic, IT contractor at Avecia we scored 47 against a threshold of 50, predicting inside, and the tribunal held Outside IR35.

Every one of those misses sits within 4 points of the 50 threshold, which is the part worth taking seriously: the engine is not reading these engagements wildly differently from the tribunal, it is landing just the wrong side of a line we drew. We publish that because a study that flatters the tool published alongside it is the version nobody should believe. It is also a fair warning about our own contract checker: a score close to the threshold is close to the threshold, and should be read as uncertainty rather than as a verdict.

Method, and its limits

The CEST engine is not a reimplementation. It is a verbatim port of the official CEST decision tables, version 1.6.0, published under the Open Government Licence. We look answers up in HMRC's own tables rather than modelling what we think the tool would do.

Encodings are committed before the engines run. Each case is encoded from the judgment's findings of fact, with a note per answer citing a paragraph or a contract clause. Answers are never revised after seeing what CEST returned. The interesting result would be CEST being wrong, which is exactly why it would be easy to drift towards it one field at a time.

These are different populations. HMRC's 22% comes from self-reported answers about mostly routine engagements. Ours comes from engagements contested hard enough to reach a tribunal, encoded from judicial findings. No sample size makes those the same population, and the comparison is context rather than a like-for-like test.

The two engines are not independent. Our scorer derives its substitution dimension from the CEST personal service result, and substitution carries 20% of our weighting. So the "our engine" column partly inherits the CEST column. It is published for interest and is not a second opinion.

The encoding is the contestable part. Several answers are inferences where a judgment makes no express finding, and each one says so. If you think an answer is wrong, the case page tells you which field, what we chose and why.

Cases we excluded, and why

Richard Alcock, IT contractor [2024] UKUT 00099 (TCC)

No status determination survives to compare CEST against. The Upper Tribunal set aside the First-tier Tribunal decision and remitted the appeal rather than remaking it, saying it was not "sufficiently equipped with appropriate findings of fact to remake the Decision". Scoring CEST against a decision that has been set aside, or against a first-instance finding an appellate tribunal has held to be legally flawed, would be scoring it against an answer that is no longer the law of the case. The tribunalOutcome field above is a placeholder required by the type and carries no meaning while excluded is set.

Excluded cases stay in the dataset rather than disappearing from it, because which cases were left out is part of what makes a small study readable.

Common questions

How accurate is HMRC's CEST tool?

On the ten decided IR35 cases we have encoded so far, CEST agreed with the tribunal on six, disagreed on two, and returned "unable to determine" on two. That is a small sample and it does not support a general accuracy percentage, but the refusals are notable because they are cases that were contested all the way to a tribunal.

Does CEST get IR35 cases wrong?

In this study CEST reached the wrong answer in two of ten cases and refused to decide two. An undetermined result gives a contractor no protection, because HMRC does not stand behind a result the tool would not reach.

How did you run CEST against old tribunal cases?

We use a verbatim port of the official CEST decision tables, version 1.6.0, published under the Open Government Licence. Each case is encoded from the judgment's findings of fact, with a note and a paragraph or clause reference for every answer, and the encoding is committed before either engine is run against it.

Is this study peer reviewed or definitive?

No. It is a small, openly published sample with every encoding on the page so that anyone can disagree with a specific answer rather than the conclusion. At ten cases the refusal rate carries a 95% interval of 5.7% to 51%, which contains the 22% HMRC publishes across all uses, so the rate is not yet a finding.

If CEST would not decide your engagement

An undetermined result leaves you where you started. Our contract checker scores the same engagement across six dimensions that tribunals weigh and shows the reasoning behind each one, so you at least know which factor is doing the damage.

Check my contract

Background reading: why CEST cannot decide so many cases and run one set of answers through both tools.