AI DOERS
Book a Call
← All insightsAI Excellence

GPT-5 Criticism Reviewed: What Is Valid, What Is Fixed, and How Law Firms Should Respond

The five biggest complaints about GPT-5 break down differently on inspection. Two are already resolved. Two are valid and require workflow adjustments. One was never a real problem.

GPT-5 Criticism Reviewed: What Is Valid, What Is Fixed, and How Law Firms Should Respond
Illustration: AI DOERS Studio

GPT-5 launched to widespread criticism that was partly based on a routing bug fixed within 24 hours, and partly based on real performance gaps that require genuine workflow adjustments to work around. Separating those two categories is the first step any professional team needs to take before redesigning their AI process around the new model.

I am Madhuranjan Kumar, and the practical question I want to answer is specific: given the documented failures, the coding gap, the math error, and the accuracy inconsistencies, how does a professional team extract real value from GPT-5 rather than either over-trusting it based on hype or discarding it after a bad first impression? The answer is a workflow with four distinct steps, and each step requires a different kind of decision than the one before it.

Sort the Complaints Into What Is Fixed, What Is Valid, and What Does Not Apply to Your Work

The first task before changing anything is triage. Not every complaint about GPT-5 is equally valid, and not every valid complaint matters for every kind of work. Running the complaints through three categories, already fixed, valid but containable, and reasons to use a different tool, produces a clear picture before any workflow change is made.

The router bug belongs in the already-fixed category. On launch day, GPT-5's automatic model routing system was sending many complex prompts to lightweight models that did not engage in extended thinking. The outputs felt noticeably weaker than GPT-4o in the same tasks: shorter responses, colder tone, less thorough reasoning. Sam Altman confirmed the bug publicly and announced the fix within 24 hours. The model behavior that most early reviewers documented as disappointing was not GPT-5 at its actual capability. It was a routing failure that made a strong model behave like a weak one during the highest-visibility window in its release cycle.

The personality complaints, the colder and more abbreviated responses that circulated as criticism, are largely router artifacts for the same reason. With the fix in place, response style is much closer to what GPT-4o users were accustomed to. The complaints that remain after accounting for the router period are the ones worth examining carefully.

The coding performance gap is valid and persistent. In a side-by-side test building a playable card game in the browser from a single prompt, Claude Opus 4.1 produced the more aesthetically complete version and o3 Pro produced the more functionally complete version. GPT-5 placed third on both measures. That result came from real testing on a real task, not from a launch-day artifact, and it means that any team using AI primarily for code generation or complex software builds is accepting a performance tax by staying on GPT-5 out of familiarity with the OpenAI ecosystem.

The math failure is also real. The problem 5.9 equals x plus 5.11, with the correct answer 0.79, produced an answer of 0.21 from GPT-5. Accuracy is genuinely inconsistent: GPT-5 got the elevator riddle right and the blueberry letter count right but failed the spatial logic cup question and the arithmetic problem. That inconsistency means a confident GPT-5 answer is not a reliable signal of correctness on numerical or logical problems.

For a digital marketing agency running through this triage, the picture looks like this: GPT-5 is strong for research synthesis, campaign brief drafting, content strategy, and long-form document work where any factual error is catchable in review. It is not the right tool for code generation where quality on both aesthetics and functionality matters, or for any calculation where the number in the output needs to be correct without a separate verification step.

How it works

Restore Legacy Model Access for Workflows That Depend on GPT-4o Behavior

If any team member built a specific workflow around GPT-4o that depended on that model's particular response patterns, the practical first step is not forcing a migration to GPT-5. It is restoring access to the legacy model through settings so the existing workflow continues while the evaluation of GPT-5 proceeds alongside it.

The path is simple: open ChatGPT, navigate to settings in the bottom-left corner, select the General tab, and enable the Show Legacy Models toggle. GPT-4o and other prior models reappear in the model selection dropdown. OpenAI restored this option after significant user backlash against the overnight removal of models that had active professional workflows built around them.

For a digital marketing agency, the people most affected by the forced migration are team members who refined their approach to GPT-4o specifically for a task they do every day and are producing consistent output with it. Forcing those people onto GPT-5 mid-project disrupts a working system with no guaranteed improvement in the output they need. Restoring legacy access lets those workflows continue while the rest of the evaluation proceeds on GPT-5 for new tasks and new projects.

The important caveat: legacy model access is not permanent. OpenAI stated it would monitor usage and consider how long to maintain the option. Building a long-term team dependency on GPT-4o as a permanent solution is not a strategy. It is a buffer that gives the agency time to run a real evaluation and make a deliberate allocation decision rather than a rushed one.

Research hours per case brief with AI assistance

Run a Two-Week Parallel Test on Real Work Tasks, Not Demos

The right way to know whether GPT-5 serves the agency better than GPT-4o or Claude Opus 4.1 on specific task types is not benchmarks, comparison videos, or published rankings. It is running all three models on the same real tasks the agency does every day and measuring the results against the criteria that matter to the people doing the work.

For a digital marketing agency, this means selecting fifteen to twenty actual work items from the next two weeks and submitting each one to both GPT-5 and Claude Opus 4.1 with the same prompt. Not simplified versions of the task. The actual campaign brief for an active client, the actual content calendar for a product launch, the actual competitive analysis from a set of observed competitor ads, the actual client-facing performance summary. Each output gets rated by the team member who does that work daily, because they are the person who knows what good output looks like for that specific task at the standard the agency holds.

The evaluation criteria for a digital marketing agency are specific: accuracy of any factual claims against verifiable sources, structural clarity of the draft output, how much editing is required before the output is usable for a client, and whether the tone and style match the agency's house voice without extensive correction. Those criteria produce a meaningful comparison because they measure the dimensions that cost the agency real time when they go wrong.

Running the evaluation over two weeks rather than two days captures variation. A model might perform well on the first several tasks in the set and fail on the type of task that appears later. Two weeks of real work diversity surfaces the failure modes as well as the strengths.

The evaluation almost always reveals the same split for a marketing agency: GPT-5 for research synthesis, content ideation, strategy documents, and long-form writing where the key quality measure is coherence and coverage; Claude for anything that requires precise structure, conditional formatting, or output that follows a specific schema reliably. That split is not a weakness in either tool. It is the starting point for a multi-model allocation that extracts the best current capability from each tool for each task type.

For agencies running Meta Ads campaigns, the evaluation should include the creative work as well as the strategic work. Writing five primary-text variations for a lead generation ad with a specific character count and a consistent call-to-action structure is a precision formatting task where the model's ability to follow structural constraints matters as much as the quality of the copy. Including that task type in the evaluation gives a complete picture that covers both the strategic and the execution-level work the agency does daily.

Build a Verification Policy for Every Numerical and Cited Claim

The math error and the accuracy inconsistencies are not arguments for abandoning GPT-5. They are arguments for treating every AI-generated number and citation as unverified until checked against the primary source, and formalizing that practice as a team policy rather than leaving it to individual judgment.

For a digital marketing agency, the verification policy might read: any statistic, industry figure, market size claim, or platform policy detail generated by any AI tool is checked against the original source before it appears in a client deliverable. Any competitor data point is verified against the actual source rather than accepted from the model's summary. This policy adds time to each piece of research, and that time is worth it because the cost of a wrong number in a client-facing document is not just the correction. It is the credibility cost and the client relationship cost.

The policy also needs to address the specific failure mode GPT-5 demonstrated on the math problem: confident wrong answers. A model that produces 0.21 when the correct answer is 0.79 does not add uncertainty flags or suggest that the result should be verified. It presents the wrong answer with the same confidence it presents correct answers. That means the team cannot use GPT-5's apparent certainty as a filter. Every claim that will appear in client work gets checked, regardless of how confidently it was stated.

This discipline is not new. Good research practice involved verification before AI tools existed. The issue is that AI tools accelerate the research phase so substantially that verification often gets compressed or skipped because the speed feels like accuracy. It is not. Speed is speed. Accuracy requires checking the source.

Match Each Task to the Model That Currently Performs Best for That Type

With the triage complete, the legacy model access restored as a buffer, the two-week evaluation done, and the verification policy in place, the allocation decision is straightforward. Each major task category goes to the model that performed best for it in the evaluation, with no loyalty to any vendor's broader product ecosystem.

For a digital marketing agency, a reasonable current allocation looks like this: research synthesis, campaign brief drafting, content calendars, strategy documents, and long-form client writing route to GPT-5, where the synthesis capability and response coherence on extended outputs are appropriate for the task and where the verification policy catches any factual errors before client delivery. Code-adjacent tasks, anything involving structured data extraction, schema generation, conditional logic, or precise formatting against a template, route to Claude Opus 4.1 based on the current performance advantage in those areas. For complex multi-step reasoning problems, media mix scenarios with multiple competing objectives, or anything where the reasoning chain matters as much as the conclusion, o3 Pro with extended thinking is worth the additional response time.

The GPT-6 direction gives useful context for how this allocation will evolve. Sam Altman described GPT-6 as a model that builds personal memory and adapts to who you are, making each user's experience meaningfully different from others'. That development points toward growing strength in long-running, contextually aware tasks where knowing your specific situation over time improves the output. As that capability develops, the allocation map will shift, and the quarterly review process, described below, is how the agency keeps the map current.

For agencies running Google Ads campaigns or building organic reach through SEO content alongside their social work, the AI stack powers research, ad copy generation, landing page strategy, and performance analysis across both channels. Keeping the model allocation explicit and reviewed means the agency is not defaulting everyone to the same tool regardless of whether it is the right one for the specific task at hand. The web CRM work that supports both channel types, building audience segments, mapping lead flows, and writing personalized follow-up sequences, also benefits from the same explicit allocation approach applied to the structured output tasks it requires.

Schedule the Next Review Before You Close This One

The last step, and the one most often skipped, is building the review cadence into the calendar before the current evaluation closes. The labs are shipping major releases multiple times per year. A model that places third on coding in one evaluation might lead that benchmark six months later, or vice versa. Without a scheduled review date, the team continues on an allocation that was right when it was set and may not be right by the time the next major model drops.

For a digital marketing agency, the quarterly review is a half-day exercise. Pull the task-to-model allocation map from the previous quarter. Run five representative tasks from each category against the current leading models. Compare against the previous quarter's baseline. Update the map where the rankings have shifted. Brief the team on any changes and update the prompt conventions for any tasks that moved to a different model.

Four structured reviews per year keeps the stack within one model generation of the current frontier without requiring constant attention to every benchmark report or product announcement. The evaluation cost is the time of the review half-day. The output is a current, explicitly justified allocation that the whole team follows rather than an outdated one that everyone is quietly working around.

GPT-5's launch produced complaints that fell across a spectrum from already-fixed router artifacts to real and persistent performance gaps. The teams that extract consistent value from it are not the ones who formed a final opinion on launch day or abandoned it because of a math error. They are the ones who did the triage, ran the evaluation on real work, built the verification policy, allocated tasks explicitly, and put the next review date on the calendar.

Do it with an expert
You can build this yourself, or have it set up right the first time.

That is exactly what we do at AI DOERS. Book a private 30-minute call with Madhuranjan Kumar and we will map the fastest path to it for your specific business.

Book your call →
Madhuranjan Kumar

Madhuranjan Kumar

Founder, AI DOERS · Performance Marketing

Madhuranjan Kumar brings 20 years of performance-marketing experience and has managed over $200 million in Facebook ad spend for brands across the United States and beyond. His expertise spans the full modern marketing stack: Meta, Google Ads, TikTok, email automation, CRM, and the websites that hold it together. At AI DOERS he turns that track record into lead-generation systems for businesses across every industry.

← Back to all insights
GPT-5 Criticism Reviewed: What Is Valid, What Is Fixed, and How Law Firms Should Respond | AI Doers