How this actually works
Written for the person who asked at the back of the room. About ten minutes, no maths, and it stops where the useful part stops.
If you have been in one of our sessions, you probably heard someone ask how these things work and heard the answer that it is not useful to you today. That is true, and it is also a bit unsatisfying. So here is the version that is worth knowing.
Everything below is about how to work with one of these things. None of it is about how one is built. That is a genuinely interesting subject and it will not change a single decision you make at work, which is why it is not here.
What it is doing
A language model reads what is in front of it and works out what should come next, one piece at a time. That is the whole mechanism. It read an enormous amount of text and it got extremely good at continuing a passage in a way that fits.
This explains almost every strange thing you have seen it do. It explains why it writes beautifully and confidently about something it has got wrong: a confident wrong answer fits the passage just as well as a confident right one. It explains why giving it three examples of what you want works better than describing what you want. And it explains why it will invent a policy clause rather than say it does not know, unless you have told it that "I do not know" is an allowed answer.
The one thing worth taking from this section: it has no idea what it does not know. Everything else follows from that.
The three things it needs, and you already know them
The most useful frame for anyone in a People team is one you use every time somebody joins. You decide what a new starter needs to read, which systems they get access to, and what they are allowed to settle on their own. Then you check their work for a while.
That is the same decision, three times over:
- Context. What you put in front of it. The policy, the handbook, the record, the three examples of good work.
- Tools. What you let it reach. The inbox, the folder, the system.
- Skills. What it does with those. There are about eleven, and they are at the bottom of this page.
Every question about whether something is allowed turns out to be one of those three, asked in legal language. People who have safely onboarded someone into a role with access to salary data and grievance files already know how to make this decision. They just have not been told it is the same one.
Prompts, skills, tools and agents
These four words get used as though they describe different worlds. They are the same thing at four different heights, and knowing which one you are talking about resolves most confused conversations.
| What | What it is | Who runs it |
|---|---|---|
| A prompt | You type, it answers, you use the answer. Nothing is kept. | You, in the moment |
| A skill | The same job written down properly: instructions, your rules, a couple of real examples, and what it must never do. It runs the same way every time. | Anyone on the team, still by hand |
| A tool | A skill with a way in and a way out that is not a chat window. A screen, a form, a button in something you already use. | Anyone, without knowing what is inside it |
| An agent | One or more skills with permission to reach a system and take the step. | Nobody. It runs. |
Almost everyone can already do the first one. The gap that matters is the second to the third, and nearly everybody believes it needs an engineer. It does not, and it has not for a while. The fourth is where it gets genuinely harder, and the difficulty there is permission rather than technology.
What makes the difference between a good one and a bad one
A bad skill and a good skill are the same technology. The difference is entirely how much of your own expertise you were willing to write down.
Take writing a job advert from a hiring manager's brief.
The bad version is "write me a job advert for a Finance Business Partner". You get something generic, in a house style it invented, with benefits you do not offer, and you rewrite it. Then you conclude that this stuff is overrated.
The good version gives it the brief, your template, your inclusive language rules, and three adverts you were genuinely happy with, pasted in whole. Then it tells it what it must never do: the salary comes from the brief and nowhere else, never use a word from the banned list, and if the brief does not say something, leave a gap and mark it rather than filling it in.
You get a draft with three gaps marked. You fill them, change what you want, and post it. Eight minutes instead of forty.
Notice how little of that was about the machine. It was your template, your rules, your examples, and your judgement about what it must not guess at.
The one instruction most people leave out
Tell it what to do when it does not know. The answer to a question it has no source for is "I do not know", and if you have not said so, it will produce something plausible instead, in your name, and you will not be able to tell.
This is the single highest-value sentence you can add to anything you write, and it costs nothing.
The eleven skills
We took the 79 tasks that make up People work across twelve functions, and classified each one by what a machine would actually have to do. They fall into eleven kinds. Twenty-two of the seventy-nine turned out to be the same one, which is a finding in itself: most of what people mean when they say they have tried ChatGPT is the first row.
| Skill | What it does | Tasks |
|---|---|---|
| Draft and review | Writes a first version a person then fixes | 22 |
| Summarise | Turns something long into something short | 13 |
| Check against rules | Reads a thing and says which rules it breaks | 9 |
| Compare and rank | Puts a set of things in an order | 6 |
| Workflow automation | Moves something from one place to another | 6 |
| Answer from source | Answers a question from documents you gave it | 5 |
| Monitor and flag | Watches something and tells you when it changes | 5 |
| Extract and structure | Pulls fields out of unstructured text | 4 |
| Transcribe and note | Turns speech into structured notes | 3 |
| Generate variants | Produces several versions of one thing | 3 |
| Classify and route | Decides what a thing is and sends it on | 3 |
The fourth one is the one to be careful with. Putting things in an order is easy for a machine and mostly harmless, right up until the things are people. Nearly every rule you have read about AI applies to that row and to almost nothing else on the list.
And most real tools are two of these in a row. Answering a candidate who asks where they are in the process is answer from source, to find out what is true, then draft and review, to write the reply. Two ordinary things in an order. That is how nearly everything gets built.
The eleventh row is not the end of the list of things that matter. Judgement is not on here. Deciding whether the work should happen at all, what it means, and what to do about the person it concerns, is not on this table and is not coming.
What we deliberately have not covered
Which model to use. There are over three million models published on the main open repository, and even a curated list of the ones that matter runs to about 157 from 23 different companies. You are not going to keep up with that and neither are we. For the eleven things above they are broadly interchangeable, and what actually decides it is what your employer already pays for, where your work already lives, and what the contract says about your data.
How they are trained, what a token is, what a context window is. Interesting, and none of it will change what you do on Monday.
Where to go next
- The report, which is which parts of which jobs this can already reach, across 923 occupations and 19,265 rated tasks.
- The library, where the builds are, with the source code and what each one cost to run.
- Ask us something, if there is a question this page should have answered and did not.