What You Should Actually Use GPT-5 For, and Where Other Models Still Win
GPT-5 is excellent at planning and at inferring what you left out, but it is not the answer to everything. Here is how to choose the right model for writing, research, development, and reliable number work.

A small accounting firm was quietly losing hours to two problems that had nothing to do with accounting. Juniors spent afternoons hand-checking revenue across large client tables, verifying numbers a model kept getting wrong. And every time the partners tried to plan a new service, the AI they consulted handed back a grand twelve-month roadmap built for a firm three times their size, so the plan never got executed. This is the story of how the firm fixed both by matching the right AI model to each job, and what changed once they stopped asking one tool to do everything.
I am Madhuranjan Kumar, and I use this walkthrough because it cuts through the noise around GPT-5. A lot of the lukewarm opinions online come from people on a free plan with the deeper thinking turned off, or from people who do not do much planning or number work. Those are exactly the cases where the picture is more interesting than any headline. The firm's journey shows where GPT-5 genuinely leads, where other models still win, and why the real skill is routing each task to the tool that owns that category.
The starting point: one tool for everything, and it showed
Before the change, the firm used a single general model for all of it: drafting client emails, checking figures, and sketching out plans. The results were uneven in a way that is easy to recognize. The writing was fine. The planning was over-scoped, always imagining the firm as bigger than it was. And the number-checking was the worst of it, because a model that miscalculates on a large table is not a tool you can trust, it is a toy that occasionally lies to you with total confidence. A junior would run a total, get a plausible but wrong figure, and only catch it because the reconciliation did not tie out. That single weakness poisoned trust in the whole approach.
The insight that started the turnaround was simple: no single model wins every category yet, so stop pretending one does. Once the firm accepted that, the work became a routing problem. Which model leads writing, which leads planning, which leads reliable math, and how do you send each task to the right one.

Chapter one: solving the numbers with a private open model
The number-checking problem had a specific fix. Reliable math on large tables had been the missing piece in open models for a long time; most of them kept getting the arithmetic wrong on big tables, which is exactly why they never made it into an accounting workflow. That changed when a new small open model finally handled it, calculating revenue correctly from large tables in a way earlier open models could not.
For the firm, this mattered for two reasons at once. First, accuracy: a model that computes correctly turns a junior's afternoon of manual verification into a few minutes of checking. Second, privacy: because it is a small open model the firm can run itself, client data never leaves the building. That single choice, a reliable open model run privately, solved the trust problem that had kept open models out of accounting entirely. The firm set it up, tested it against a handful of known tables where they already knew the right answers, and only then trusted it on live work. The reconciliation errors that used to slip through started getting caught before they ever reached a client.

Chapter two: fixing the plans with GPT-5's restraint
The planning problem needed a different tool. GPT-5's biggest strength, the thing that genuinely surprised me when I first pushed on it, is restraint. When you hand it a large, real body of work and ask for a plan, it tends to meet the project where it actually is instead of reimagining it as a year-long startup effort. It reads what you already have and proposes a near-term path you can actually execute, rather than a sweeping vision that stays on the whiteboard forever.
The firm put that to work on tax-season workflow and a new advisory service. The critical move was telling the model the firm's real size up front. A plan built for a ten-partner firm is useless to a small shop, so the firm stated its actual constraints, three partners, a handful of juniors, a fixed capacity, and GPT-5 returned a plan scaled to that reality. Where the firm's old general model had quietly re-architected the advisory idea into a six-to-twelve-month program, GPT-5 gave them something they could ship that quarter. The model is also strong at inferring the parts of a request you left out while still respecting every part you spelled in, which meant the plans filled in sensible detail without ignoring the firm's stated limits.
Chapter three: the habit of comparing two plans
The most useful working habit the firm adopted was comparison. When a decision actually mattered, they would ask one model to compare its plan against another model's plan, and the differences exposed which one fit their situation. In one case the firm deliberately did not say which of two approaches it preferred and asked the model to weigh them. It correctly read where the firm actually was and recommended the tighter, near-term path, while the other model had quietly re-architected everything into a long-horizon vision. Seeing the two side by side made the choice obvious before anyone committed time to the wrong one.
This is a small ritual with outsized value. Plans are cheap to generate and expensive to execute, so spending a few minutes surfacing the trade-offs before you commit is the highest-leverage step in the whole process. The firm made it standard: no significant plan went forward until it had been compared against an alternative and scaled to the firm's real size.
What the results looked like
Put illustrative numbers on the change. Before the routing habit, the firm was catching maybe two reporting errors a month, and usually late. After matching the reliable open model to the number work and running planning through GPT-5, the caught-error count climbed steadily as the juniors verified faster and more often, into the high teens per month within a few months. Those figures are illustrative, but the mechanism is concrete: a model that calculates correctly means a junior spends minutes verifying instead of an afternoon, so more gets checked and more gets caught. A planning model that respects the firm's real scale means the new advisory service shipped in a quarter instead of dying on a twelve-month roadmap.
None of this asked the team to become developers. The wins came from writing down the four jobs the firm does most, drafting, planning, research, and numbers, and assigning the leading tool to each. For drafting clear client emails and summaries, GPT-5 stayed concise and human. For business and service planning, it produced realistic, executable plans. For deep research, another model still tended to lead. And for private, reliable calculations, the small open model became the firm's trusted engine. The cleaner client records and faster turnaround also fed the firm's CRM and website stack, so follow-up and onboarding stopped depending on anyone remembering to do it.
Chapter four: what changed in the day-to-day
The most telling sign that the routing habit had taken hold was not a metric, it was a change in how the office felt during a busy week. Before, the juniors dreaded the number-verification work because it was slow and the tool they used could not be trusted, so every figure had to be checked twice. After, they ran the reliable open model on a table, spot-checked a few known values, and moved on, because the model had earned the trust that the old general tool never had. The dread came out of the work.
The partners felt a parallel shift on the planning side. Before, asking an AI for a plan produced a document so ambitious it was demoralizing, a reminder of everything the firm did not have the capacity to do. After, telling the model the firm's real size up front meant the plans came back sized to the team, so a plan was something to start on Monday rather than something to file away. That small change, plans that respect your actual constraints, is the difference between AI as a source of anxiety and AI as a source of momentum.
The habit is portable, the tools are not
One reason I stress the routing habit over any specific model is that the models will change and the habit will not. The particular open model that finally handled large-table math will be replaced by a better one; the specific strengths of GPT-5 will be matched or exceeded by something else. None of that breaks the approach, because the approach is not attached to a product. It is a way of thinking: name the jobs you do most, learn which tool currently leads each one, tell every tool your real constraints, and compare options before you commit.
A firm that internalizes that stays ahead regardless of which model is on top this quarter, because it is practiced at evaluating and routing rather than loyal to a brand. The firms that struggle are the ones that pick one tool, use it for everything, and then judge all of AI by that single tool's weakest category. The accounting firm in this story did the opposite, and the payoff was not just fewer errors and faster plans, it was a team that had learned to treat AI as a set of specialized instruments rather than a single magic box. That distinction matters more every month, because the pace of new releases means brand loyalty ages badly. The firm that checks, quarterly, whether a better tool has taken the lead in one of its four core jobs spends an hour and stays current. The firm that locked in a year ago is quietly running the worst version of every task it does.
What this means for any small firm
The lesson generalizes well beyond accounting. Any business that writes, plans, researches, or works with numbers can benefit, which is nearly all of them. The point is not to crown one winner. It is to keep a short mental map of which tool leads which category and route each task accordingly, telling every model your real constraints so the answer is something you can act on rather than a vision you will never finish. For a firm handling sensitive numbers, the arrival of a reliable open model you can run privately is the piece that changes the calculus, because it removes the last reason not to automate the number work.
The clarity this creates has a way of rippling outward. A firm that trusts its numbers and ships its plans has real proof points to talk about, which strengthens the content behind SEO and organic search and sharpens the targeting on any Facebook and Instagram ad campaigns the firm runs to attract new clients, because the message is grounded in results the firm can actually stand behind.
To set this up yourself, write down the four jobs you do most and assign the leading model to each so routing becomes automatic. On a paid plan, switch on the deeper thinking for planning and analysis, since the quick default is where most of the disappointing opinions come from. Always tell the model your real scale before asking for a plan, and when a decision matters, ask it to compare two plans so the trade-offs are visible. For anything involving client numbers you cannot send to the cloud, run a small open model locally and test it on known tables first. You can build this routing habit yourself over a few sessions. If you would rather have the whole setup designed around your firm, with the private number model and the planning workflows wired in, that is the kind of thing you can hand to an expert and skip the experimentation.
That is exactly what we do at AI DOERS. Book a private 30-minute call with Madhuranjan Kumar and we will map the fastest path to it for your specific business.
Book your call →
