Six criteria, 30 points, the same six for every tool. Written and published before the first review, because a rubric written afterwards is a justification rather than a standard.
Think a score is wrong? Say so on the Ask page with the criterion and what you have seen. If you are right, the score changes and the page says it changed.
The exposure report splits every task in a job into three bands: reachable with a chat window today, needs purpose-built software first, and cannot be sped up at all. That split is the buying question. A tool charging for the first band is charging for something you already have. A tool promising the third is promising something nobody can deliver.
The middle band is where software earns its money, and it is the first and heaviest criterion below. Every category page says which tasks that means for that kind of tool, taken from the report rather than from the vendor.
Each is scored 0 to 5 on what we saw, not on what the vendor says. The total is a sorting key, not the review: a tool can score well on five of these and still be wrong for you because of the sixth.
Of the work the research puts in the band that needs software building first, how much does this tool actually do?
The exposure data already says which parts of a job a chat window handles on its own, so a tool charging for those is charging for something you have. The work worth paying for sits in the middle band, and that is what this criterion measures. It applies whether or not the product mentions AI at all: a workflow tool that genuinely moves the documentation scores well here, and a chat feature bolted onto a form does not.
From signing up, how long until it produces output you would actually use at work?
Not time to a demo and not time to a dashboard. Time to a thing you would put your name on. Implementation is where HR software goes to die, and the gap between the sales call and the first real output is the single best predictor of whether it gets used a year later.
Can you get everything out in a usable format, have it deleted, and does your data train their models?
Read from the contract and the data processing terms, not the marketing page. This is employee data: names, salaries, grievances, health information. It is the one thing on this list you cannot fix later by switching tools.
When it scores, ranks, filters or recommends a person, can you see the reason, and could you explain it to that person?
A recruiter who cannot say why a candidate was ranked eleventh has bought a liability, not a tool. This is also where the regulation is heading in both the UK and the EU, and vendors who cannot do it today tend to be quiet about it rather than clear.
Is the price published, and can you work out what year one actually costs?
Per seat, per employee and per hire are three different businesses wearing the same suit. The published number is rarely the number. What matters is whether you can forecast the bill before you sign, including implementation and the seats nobody mentioned.
What does leaving cost, in notice, in money, and in data?
Every HR system is easy to buy. The ones worth buying are the ones that are also easy to leave, and the auto-renew clause is where that is decided. Checked in the contract, because it is never on the pricing page.
The band definitions used by the first criterion come from the exposure ratings in Eloundou, Manning, Mishkin and Rock, GPTs are GPTs: Labor market impact potential of LLMs, Science 384, 1306–1308, 2024, applied to O*NET task statements. How to read those.