AI roadmap case study: FinEdge's first 90 days

    An AI roadmap case study: how a 280-person fintech People team went from scattered ChatGPT use to nine production workflows in 90 days, and what nearly broke.

    Matthew BradburnΒ·Β·

    Seven people. Two using AI heavily, two refusing to touch it, three somewhere in between. That was the FinEdge People team on day zero, with a six-figure AI budget already signed off and nobody assigned to spend it. This AI roadmap case study is what happened next: how that team went from scattered individual ChatGPT use to nine workflows in production in 90 days, sequenced as diagnose, prove, harden and scale. The budget was never the constraint. The absence of a method was. Almost every People function I meet has the same gap, money approved before a plan exists, so this is written in enough detail that another team could run the same play. Names, numbers and one sequencing detail are changed. The shape of the work is exact.

    Day zero: a budget with no plan

    FinEdge is a 280-person fintech, Series B, payments infrastructure for mid-market merchants. The People team is seven: a CPO, two HRBPs, two TA partners, a People Ops lead, and a learning lead. The CEO had approved a six-figure People AI budget at the previous board meeting. The CPO had that budget and no idea where to point it.

    The state of play before we started: individual ChatGPT use scattered across the team, no shared workspace, no policy, no measurement. Two of the seven were heavy users, three occasional, two avoiding it. Uneven adoption, which is the normal starting position, not a failure. What was missing was any read of where the work actually snagged. The instinct in the room was to spend month one buying tools. That instinct is almost always wrong, and the diagnostic is how you prove it.

    The four-phase arc we ran against that starting position looked like this.

    1. 01
      Days 1-14
      Diagnose

      Read the function at click level. Find the hot cells and the weakest readiness scores before naming a single tool.

    2. 02
      Days 15-60
      Prove

      Ship three workflows, one champion on each. Small audience, real data, a measured effect you can name.

    3. 03
      Days 61-90
      Harden

      Fail-close the error states, publish the policy, stand up the shared workspace and the evaluation review.

    4. 04
      Day 90+
      Scale

      Hand the pattern to the rest of the team so they replicate it without you in the room.

    The diagnostic that redirected month one

    We ran the People Ops diagnostic toolkit across three sessions over two weeks. The output was a one-page read of the function: a workflow heatmap, a readiness score across tooling, data and governance, and a shortlist of build candidates ranked by frequency, pain and AI fit.

    Two findings reshaped everything that followed. The first was that the team had guessed wrong about its own hot cells. Everyone expected the hottest cells to sit in talent acquisition, because that is where the noise is. The heatmap put them in onboarding (high frequency, high pain, strong AI fit on the document and Q&A steps) and in HRBP request triage (medium frequency, very high pain, very high fit on the routing and first-response steps). This is the value of reading before building: the loudest process is rarely the one with the most trapped time.

    The second finding was the readiness gap, and it is the one that saved the budget from being wasted.

    Readiness dimensionScore /10Month one as plannedMonth one after the diagnostic
    Tooling6Buy more toolsOne tooling decision, then stop
    Data3Assumed fineMake the handbook and Slack history retrievable
    Governance2DeferredDraft the policy and the evaluation habit

    The team had been about to spend its first month on tooling, its strongest area. The diagnostic redirected that month to data and governance, its weakest, with a single tooling decision made and then parked. You do not fix a 3-out-of-10 data foundation by adding a ninth tool. This is the same reason so many pilots stall short of production: the model works, the organisation cannot absorb it, a pattern worth understanding before you build anything, covered in why AI pilots stall at production.

    Prove value before you harden anything

    Days 15 to 60 were about proof, not scale. Three workflows in six weeks, one named champion on each, deliberately small audiences.

    WorkflowWhat it doesMeasured effectLive by
    Onboarding Q&ARetrieval assistant grounded on the handbook and 18 months of People Ops Slack answersReplaced ~40% of week-one new-hire questions, ~6 hours a week savedDay 28
    HRBP request triageClassifies inbound requests, routes them, drafts a first response for the partner to approve or rewriteMedian response time cut from 26 hours to 4Day 42
    Interview-loop schedulingCoordinates panels, holds slots, drafts candidate commsTA partner scheduling time roughly halvedDay 56

    The onboarding assistant mattered most because it grounded on real, retrievable context, not a generic model prompted blind. That grounding was only possible because month one had made the handbook and the Slack history reachable in the first place. The triage workflow mattered most to the HRBPs, because a 26-hour median response time was the thing quietly eroding their credibility with the business, and dropping it to four hours changed how the rest of the company saw the function. Interview-loop scheduling was less novel and carried less payoff, but it freed a TA partner from the single most-hated chore on the team, which bought goodwill for everything that came after.

    One ritual did more than any of the three builds. A weekly Friday demo, starting in week three, fifteen minutes, one champion shows one thing. It never stopped. It beat every formal training session combined, because it made progress visible and made the shy two-thirds of the team want in. If you take one thing from this AI roadmap case study, take the Friday demo.

    Harden and scale

    Days 61 to 90 added six more workflows, smaller and more specialised, owned by individual team members rather than dedicated champions. By this point the pattern was legible enough that the rest of the team could replicate it without hand-holding. That is the definition of scale that matters: the capability moved into the team and stopped depending on the person who built the first three.

    The harder work in this phase sat outside the workflows, in three pieces of infrastructure that make the difference between nine toys and nine production systems. Together they are what an AI workspace for People Ops actually is.

    • A published one-page AI policy covering data, vendors, evaluation and what the team is forbidden to automate. One page, so people actually read it.
    • A shared workspace with reusable context and standardised prompts, following the AI workspace setup pattern, so a new build starts from the team's accumulated context rather than a blank box.
    • A metrics pack tied to revenue and risk, refreshed monthly and presented at the exec team's business review, so the People function reported its AI value in the same language as every other function.

    By day 90: nine workflows in production, an estimated 38 hours a week of sustained team time freed, and a CFO who could answer what is the People AI spend returning? without flinching. Getting that reporting right is its own discipline, and it is what let the function report its AI value in the same language as every other line on the board pack.

    The two near-misses evaluation caught

    Both happened in the second month. Both were caught by the standing evaluation review, not by luck. Neither reached a single person outside the People team. That is the number that matters most, and it is the one you can only earn by building the review before you need it.

    0

    Critical issues that reached anyone outside the team, two months on. Both near-misses were caught in review. Neither was caught by luck, and both would have shipped without it.

    The gate we now run every workflow through before it reaches production came directly out of those two weeks.

    What we would do differently

    Two things, and both are about timing.

    Build the evaluation habit on day one, not day thirty. We added it in week five. Both near-misses formed in weeks six and seven. The five-week lag was very nearly the cost of the whole programme. The review is cheap to stand up and expensive to skip, and there is no reason it cannot exist before the first workflow does.

    Resist demoing every new workflow company-wide too early. Once the wider company has seen something, retiring it becomes political, and you will keep a mediocre workflow alive because pulling it looks like failure. Keep the audience small until the workflow has survived a month of weekly evaluation. The Friday demo is the right size of audience for that. The all-hands is not.

    What this AI roadmap case study actually transfers

    The temptation with a case study like this is to copy the nine workflows. Do not. Your hot cells will sit somewhere else entirely, and the specific builds are the least transferable part. What transfers is the order of operations, and the difference between the version that sticks and the version that stalls is almost entirely about where you start.

    Roadmap that sticks

    Starts with a two-week read of where the work actually snags

    Month one spent on the weakest scores, usually data and governance

    Workflows chosen from the heatmap's hottest cells

    Every item tied to an operating outcome a CFO can name

    Evaluation review standing before the second workflow ships

    Roadmap that stalls

    Starts with the tool the loudest vendor demoed last

    Month one spent buying seats and licences

    Workflows chosen by whoever shouts first in the room

    Items justified by 'it is impressive', not by outcome

    Evaluation bolted on after something has already gone wrong

    The left column is not more expensive than the right. It is the same budget spent in a different order.

    The demo-led roadmap optimises the vendor's pipeline. The diagnostic-led roadmap optimises yours. FinEdge got nine production workflows and 38 reclaimed hours a week for the same six-figure budget the demo-led version would have spent on licences and left the team no better off. The money was never the variable. The order was.

    If you want to run this shape against your own function, start where FinEdge did, with a read of the operating reality before a single tool decision. That is exactly what the Grain Audit delivers: one process mapped end to end, a ranked automation plan, and a 90-day plan you keep and run yourself.

    Common questions

    How do you build a 90-day AI roadmap for a People team?
    Sequence it by absorption capacity, not by tool appetite: diagnose for two weeks, prove value on two or three workflows for six, then harden and scale for the last four. FinEdge did exactly that and reached nine production workflows by day 90. Skip the diagnostic and you rebuild the same workflow three times.
    What was FinEdge's starting position?
    A 280-person Series B fintech, People team of seven, two heavy ChatGPT users, three occasional, two avoiding it entirely. No shared workspace, no policy, no measurement. The CPO had a six-figure budget approved and no plan. We see this exact split on most first calls: the money moves faster than the method.
    What did the 90-day roadmap deliver?
    Nine production workflows, a published one-page AI policy, a shared workspace with reusable context, a weekly demo, and a metrics pack tied to revenue and risk. Around 38 hours a week of team time freed by day 90, sustained. For the first time the CFO could say what the People AI spend was actually returning.
    What were the near-misses, and how were they caught?
    Two, both caught by a standing evaluation review rather than by luck. A comp-letter automation that would have leaked salary-band logic in an error state, caught in a critique pass. A candidate-screening agent that started encoding a hiring manager's preference with no link to performance, caught in the weekly review and removed. Neither reached anyone outside the team.
    Is this roadmap repeatable for another team?
    The sequence is repeatable. The specific workflows are not. Every team's hot cells on the workflow heatmap sit in a different place, so your first three builds will not be onboarding, triage and scheduling. Run the same shape, diagnose then prove then harden then scale, and let your own diagnostic pick the targets.
    11 min

    Not sure where your function stands yet?Take the Readiness Assessment→

    When reading turns into doing

    The Grain Audit maps one People Ops process end to end, ranks the highest-return automations, and hands you a 90-day plan you keep whether or not we work together.

    Two weeks. Β£2,000, credited in full against a programme. Three slots a month.

    Book a Grain Audit

    If this resonated, there's more.

    Subscribe to receive new Intelligence pieces as they're published. No noise, just the work.

    By subscribing you agree to our Privacy Policy. Unsubscribe any time.

    Diagnostic

    Where does your People function stand?

    Score it yourself, free, in about ten minutes.

    Take the Readiness Assessment β†’