Foundations·10 min

    Choosing AI models for HR work

    No single best AI model for HR work exists. Match the model to the task, run two or three, and cross-check what matters. Here is the working stack.

    Matthew Bradburn··

    The question lands in almost every operating review the moment AI comes up: "Which AI model should we standardise the People team on?" The honest answer is that there is no single best AI model for HR work, and standardising on one is the wrong instinct. The right pattern is to run two or three models and reach for whichever one is strongest for the task in front of you: Claude for policy and sensitive comms, ChatGPT for structured drafting and workflows, Perplexity or Gemini for research you can source. The model is rarely the bottleneck. Knowing which one to pick, when, is.

    There is no best AI model for HR work

    "Best" is a vendor's word, not an operator's. It assumes one tool wins every task, which is not how the work behaves. Writing a redundancy letter and scaffolding a Zapier flow are different jobs that reward different strengths, and no frontier model tops the table on both at once. So the useful question is not "which one is best", it is "which one is best for this".

    Most teams never actually choose. They signed up for one model in year one, wired their habits around it, and never revisited the decision. That is understandable and it is expensive. The writing never quite lands and nobody can say why, because the cause is invisible: the wrong tool, applied to a task it was never the strongest at. The fix costs about £15 to £20 a month for a second subscription and roughly zero friction. Four tabs, four logins, and switching between Claude, ChatGPT, Gemini and Perplexity now takes seconds.

    The right model per task

    Two or three subscriptions, matched to task type

    Drafting, policy, research and structure each go where they are strongest

    A second model critiques what the first one wrote

    The model map lives in a one-page doc the team follows

    Reviewed every six months, because the curve moves quarterly

    One model, chosen once

    Signed up in year one, never revisited since

    Every task runs through the same model, good fit or not

    The writing never quite lands and nobody knows why

    One critique pass, from the model that wrote the draft

    Re-evaluation happens when something breaks, not on a cadence

    The cost of a second subscription is trivial. The cost of the wrong model on the wrong task compounds every week.

    What each model is actually good at

    Strip the marketing and four models carry the bulk of People work. Each has a real edge and a real weakness. Learn both and the choice mostly makes itself.

    ChatGPT (GPT-5 class). Strongest on structured output, multi-document workflows and automation scaffolding: anything where you need the model to follow instructions literally and hand back a clean table, a JSON block, or a spec a downstream tool can consume. Custom GPTs remain the easiest way to package a reusable play for a non-technical team. It also iterates without losing the thread: you can refine a draft through ten passes and it holds the brief. Where it is weaker: tone in long-form writing drifts toward generic, and it produces confident first drafts that need a critique pass before anyone reads them. Reach for it when you are drafting job descriptions, building interview kits, structuring a 30-60-90 plan, scaffolding an n8n or Zapier flow, or shipping a Custom GPT the team will reuse.

    Claude. Strongest on long-form writing, sensitive communications, policy drafting and performance-review prose: anything where tone and judgement decide whether the words land. It is more likely than the others to push back when something feels off, which is a feature, not a nuisance, when the reader is a real employee reading a real decision. It reasons well across multiple steps when you give it room to think. Where it is weaker: it is less aggressive about structured output unless you ask explicitly, and some tiers carry a smaller default context window. Reach for it when you are writing a difficult comms email, drafting a policy, framing a hard performance conversation, or redesigning a team ritual, anything where a careless sentence does damage.

    Gemini. Strongest on multimodal input and large research synthesis: feed it a screenshot, a PDF, a chart, or a pile of documents and ask it to read widely and pull the thread together. It is native to the Google stack, which matters if your team lives in Workspace. Where it is weaker: personality and tone read flatter than Claude, and it is less consistent than ChatGPT on tightly structured output. Reach for it when you are running a research scan, comparing vendors, reading survey-result screenshots, or summarising a long PDF where the input is varied and large.

    Perplexity. The answer engine for "what does the public web actually say about this", with citations built in and time filters that work. It is the first place to go for questions like "what is the current state of pay benchmarking for engineering at Series B", where you need sources you can open and check. Where it is weaker: it is not a drafting tool and not a workflow tool. Treat it as research, not as an assistant, and it earns its keep. Reach for it for market research, benchmarking, regulatory updates, and "is this actually true" checks on something another model has told you.

    The stack that works, task by task

    Here is a working default for a People function. It is not gospel; it is a starting map you adjust as the models move. The pattern that matters is the second column: a primary model does the work, a different model critiques it. Two perspectives on the same draft catches most of what a single model misses.

    TaskPrimaryCritique pass
    Writing HR policiesClaudeChatGPT
    Career pathing and coaching framesClaudeChatGPT
    Employee comms and tone rewritesClaudeChatGPT
    Recruiting: JDs, interview kitsChatGPTClaude
    Market and trend researchPerplexity or GeminiCross-check a second source
    Performance review supportClaudeChatGPT
    Meeting summaries from transcriptsGemini or ClaudeNone
    Internal knowledge base, SOPsClaudeChatGPT
    Workflow automation scaffoldingChatGPTNone
    Vendor and tool researchPerplexityGemini

    Read the table and the underlying logic is simple. Where the output is words a person will read and react to, Claude leads and ChatGPT checks. Where the output is structure a system will consume, ChatGPT leads. Where the output is a fact you are staking a decision on, an answer engine leads and you verify against a second source. The model follows the shape of the work.

    Two habits that beat any model choice

    Which model you pick matters less than two disciplines that sit underneath the whole stack. Get these wrong and the best model in the world still ships you bad work.

    The first: never ship a first draft. A first draft is a model's opening guess, not a finished thing. It exists to be argued with. The second: cross-check anything consequential with a second model. The cost is one extra prompt and one extra pass. The pay-off is work that does not come back to be redone, and mistakes caught before they reach a real employee rather than after.

    The cross-check habit earns its place most on numbers. A model will hand you a benchmark, a statistic, a "typical" figure with total confidence and no source. That is precisely where a second pass through an answer engine pays for itself.

    How to pick a model for a task you have not mapped

    The table covers the common jobs. New tasks turn up weekly, and you will not always have a row for them. Rather than guess, run the task through a short filter. It takes about thirty seconds and it beats defaulting to the open tab.

    The filter is not academic. It is the thing a champion writes down once and the team runs forever. Most of the value in a People function's AI stack is not the models; it is the fifteen minutes someone spent deciding which model meets which kind of work, and writing it where the rest of the team can find it.

    Where AI is never the decider

    One line runs under every model choice and never bends. AI is an assistant for HR decisions, never the decision-maker, and that is true of Claude, ChatGPT, Gemini and every model that follows them. All of them produce good interview question banks, calibration prompts and review draft frames. All of them will also produce a confidently wrong recommendation if you let them near a decision about a person without a human in the loop.

    The failure mode is subtle. The output reads well, so it feels authoritative, so the human review quietly thins from "judge this" to "rubber-stamp this". That is the moment the model became the decider by default, without anyone choosing it, exactly like the redundancy letter picked by the open tab. This is why AI in People work needs a written line on what a model must never decide, not a vibe. Governance for People teams is the mechanism that keeps the human at every consequential step, so speed on the drafting never turns into drift on the judgement.

    Make it a one-page doc, not a habit

    The single easiest win here is not a better model. It is writing the model map down. By task type, which model the team reaches for, which model critiques it, and where the human review is non-negotiable. Embed it in your AI workspace so it is next to the work, not buried in a wiki nobody opens. It is a one-page document. Most teams never write it, and pay for that omission every time someone drafts something sensitive in the wrong tool. This is one small piece of a wider AI workspace for People Ops: the models are the easy part, the decisions around them are the work.

    Then re-evaluate it every six months. The cost-quality curve moves quarterly; a model that trailed on tone last spring may lead on it now, and the reverse. Bake the review into a cadence rather than waiting for something to break. If you want a structured way to see where your function actually stands on tools, data and the habits around them, the Readiness Assessment scores it across four capability layers in about ten minutes.

    The model is not the interesting problem. Once you know which one to pick and when, the real question opens up: how to wire any of them into the workflows that carry real weight, which is where the move from prompts to systems begins.

    Common questions

    Which is the best AI model for HR work?
    There isn't one. The right pattern is two or three models matched to task type: Claude for policy, sensitive comms and anything where tone has to land; ChatGPT for structured drafting and workflow scaffolding; Perplexity or Gemini for research you can source. Most teams run one model out of habit and pay for it every time the writing does not quite land.
    Which AI model should we use for writing HR policies?
    Claude, for policy drafting and sensitive comms. It handles tone, edge cases and the implicit 'what would a reader take from this' check better than the alternatives. ChatGPT is fine for a structured first draft. Whichever model writes it, run the draft through a second model for critique. Two passes, two perspectives, far fewer mistakes reaching a real employee.
    Do we need more than one AI subscription for a People team?
    Usually yes. Two or three subscriptions run about £15 to £20 a month each, less than the time it takes to redo one first draft written in the wrong model. Four tabs, four logins, and the friction of switching between them is close to zero. The saving is real work that does not need redoing.
    Can AI make hiring or performance decisions for us?
    No, and no model changes that. Use AI to draft interview kits, calibration prompts and review frames, never to make the call. Every model will produce a confidently wrong recommendation if you let it touch a decision about a named person without review. Keep a human at every consequential step.
    10 min

    Not sure where your function stands yet?Take the Readiness Assessment

    When reading turns into doing

    The Grain Audit maps one People Ops process end to end, ranks the highest-return automations, and hands you a 90-day plan you keep whether or not we work together.

    Two weeks. £2,000, credited in full against a programme. Three slots a month.

    Book a Grain Audit

    If this resonated, there's more.

    Subscribe to receive new Intelligence pieces as they're published. No noise, just the work.

    By subscribing you agree to our Privacy Policy. Unsubscribe any time.

    Diagnostic

    Where does your People function stand?

    Score it yourself, free, in about ten minutes.

    Take the Readiness Assessment →