AI DOERS
Book a Call
← All insightsAI Excellence

Anthropic's Own Data Makes The Strongest Case Yet That AGI Is Already Here

Anthropic never says the words, but its report When AI Builds Itself shows Claude solving open-ended problems with no clear answer, jumping from 26% to 76% success in six months. Here is what that means for a normal business, with a worked example for an electrician.

Anthropic's Own Data Makes The Strongest Case Yet That AGI Is Already Here
Illustration: AI DOERS Studio

In November 2024, when Anthropic's researchers put their own AI against human researchers at 129 real decision points in a live research workflow, the AI chose better 51 percent of the time. By April 2025, that number had risen to 64 percent. The task duration the system could handle went from four minutes two years ago, to ninety minutes a year ago, to twelve hours today, with a single internal run logged at sixteen hours. These numbers come from Anthropic's own internal report on how Claude is used inside Anthropic. They are not from a benchmark designed for a press release.

The report also noted that 80 percent of Anthropic's current codebase is now written by Claude. Not assisted by Claude. Written by it.

What makes these figures relevant for a business with nothing to do with AI research is the specific category of task where the improvement happened: open-ended work. Tasks where there is no predefined correct answer, where even the person assigning the task cannot fully specify what a good result looks like before the work is done. That category is where most of the high-value time in any organization lives. It is also exactly the category that most businesses are still handling entirely without AI.

This playbook is about changing that, one task at a time.

Separate your open-ended work from your fixed-answer work

Before anything else is decided, the work on your plate needs to be sorted into two categories.

Fixed-answer work has a clearly correct output that can be verified against a known standard. Scheduling a meeting, calculating a quote from a price list, reformatting a spreadsheet, summarizing a document with a specific required structure, converting a file format. These tasks are valuable and AI handles them well. But they share a property that limits their impact: the person assigning them already knows what the result should look like. The task can be checked against a template.

Open-ended work is different. It has a goal rather than a prescribed output. Turning site-visit notes into a near-final estimate. Drafting a follow-up email after a client meeting. Researching three approaches to a problem and recommending one. Writing a client proposal from a conversation transcript. The person assigning this work knows roughly what they want but cannot fully specify it in advance. The quality of the result has to be judged by someone with context, not checked against a format.

Most businesses, when they first adopt AI tools, use them almost entirely on fixed-answer work. The reason is intuitive: if you know exactly what the output should be, you can tell immediately whether the AI got it right. The feedback loop is fast and the risk of an error reaching a client is low.

The problem is that fixed-answer work is rarely where the most significant hours are being lost. An hour of reformatting is a real hour, but it is usually delegatable or templatable already. The hours that genuinely constrain a business live in the open-ended category: the estimate that takes forty-five minutes to produce because it requires judgment about materials, labor, and scope; the client communication that requires synthesizing three different conversations into one coherent next step; the proposal that requires understanding a problem before recommending a solution.

Write two lists. The first: everything you regularly do that has a clearly correct output. The second: everything you regularly do where the result requires judgment and is evaluated by someone rather than verified against a known answer. The second list is where this playbook applies.

How it works (short)

Write the goal, not the steps, when handing a task to AI

The single most common mistake when moving to open-ended AI tasks is writing prompts like instructions rather than briefs.

Instructions describe the steps. "First, read these notes. Then, identify the materials needed. Then, calculate the labor hours. Then, format the output as a table." This works for fixed-answer tasks where the process is known and the steps are fixed.

A brief describes the goal and the context. "Here are the notes from a site visit with a homeowner who needs a full panel replacement and three new circuits for a home office addition. Produce a near-final estimate suitable for sending to the client, a permit checklist for this jurisdiction, and a follow-up email that confirms the scope and asks for the go-ahead to schedule. The tone should be professional but direct."

The difference matters because open-ended tasks require the AI to make judgment calls. When you specify every step, you constrain those judgment calls to the steps you already thought of, which means the output is bounded by what you already knew. When you describe the goal and the context, the AI can fill gaps, flag ambiguities, and produce a result that you evaluate on its merits rather than check against a predetermined format.

Anthropic's internal data on the jump from 26 percent success on open-ended tasks to 76 percent in six months was not driven primarily by better hardware. It was driven by better task specification and models trained on more open-ended work. For a business owner, model quality is largely outside your control. The quality of the brief is entirely within it.

The brief should contain three things: what you want as output (concrete and specific), the context the AI needs to make good judgment calls (client background, constraints, prior decisions), and the standard by which you will evaluate the result (what makes this good rather than just technically complete).

Quotes sent per week (illustrative)

Build the review checkpoint before anything reaches a client

The jump from 26 to 76 percent success on open-ended tasks is genuinely impressive. It also means that in roughly one out of four open-ended tasks, the output requires meaningful correction before it is ready.

For fixed-answer tasks, an error rate of twenty-five percent would make a tool unusable. For open-ended tasks, that rate is actually quite good, but it means that a review step is not optional. It is structural.

Before any open-ended AI task output goes to a client, it passes through a review checkpoint. The review is not a full rewrite. It is a check against three questions: Is the output factually accurate given what I know about this project? Is the tone appropriate for this specific client? Is there anything missing that the client will notice?

Building this checkpoint into the process before the first open-ended task is delegated prevents the dynamic where a single poor output creates lasting skepticism about the tool. The review step is what makes the 76 percent success rate usable rather than risky: good outputs go out unchanged, outputs that need adjustment get adjusted, and the client never sees the difference.

In practice, the review checkpoint for most open-ended tasks takes two to five minutes. An estimate review, a communication review, a proposal check. The time savings from not producing the document from scratch are typically ten to thirty minutes per task. The review step is not a tax on the process. It is what makes the process safe enough to run at production speed.

Run the first open-ended task end to end and measure the before-and-after

Do not plan to delegate ten tasks at once. Delegate one, run it completely, and measure it.

Take an electrical contractor as an example. The open-ended task that consumes the most time per occurrence is producing a complete quote package from a site-visit writeup. Before AI, the process works like this: the contractor reads through handwritten or voice-memo notes from the site visit, translates them into an itemized estimate using a pricing guide and labor tables, writes a permit checklist specific to the jurisdiction, and drafts a follow-up email to the homeowner confirming the scope and requesting authorization to proceed. Total time per quote: approximately forty-five minutes for a standard residential panel replacement and circuit addition.

After delegating this task to AI with a well-written brief that includes the site-visit notes, the relevant pricing guide, and a description of the standard permit requirements for the jurisdiction, the process changes. The contractor reads the notes once to confirm they are complete, writes the brief (eight to twelve minutes), reviews the AI output (three to five minutes), and makes any needed adjustments (two to four minutes). Total time per quote: approximately twelve to sixteen minutes.

That is a reduction from forty-five minutes to fourteen minutes per quote on average. The contractor was quoting six jobs per week before. At fourteen minutes per quote, the same working hours support nineteen quotes per week, assuming the other constraints including site visits and scheduling are also addressed.

By week twelve of running this workflow consistently, with the brief refined over several iterations based on what the review step keeps catching, the contractor's quoting volume has moved from six to nineteen per week. At a close rate of forty percent and an average job value of $3,200, the weekly revenue potential moves from approximately $7,700 to approximately $24,300. Not because the contractor found more hours. Because the open-ended work that consumed those hours was delegated.

Measure the before state before you start. Time the task, count the steps, note the output quality. Then run the first AI-delegated version and measure the same things. The gap between these two measurements is what you are actually working with, and it tells you whether this task is worth systematizing or whether the next item on the list is a better starting point.

Identify the next task once the first one runs cleanly

Once the first open-ended task is running consistently at the new efficiency level, with the brief refined and the review checkpoint functioning, the question is not whether to expand. It is which task to expand to next.

The selection criterion is the same as before: which open-ended task consumes the most time per occurrence, relative to how clearly the goal can be specified and how straightforwardly the result can be evaluated?

For the electrical contractor, the next task after quoting might be permit application preparation, which requires assembling the same information from the estimate in a different format for the jurisdiction's submission form. Or it might be the client onboarding communication sequence: three specific emails over the first week of a job that vary by project scope. Or it might be the end-of-job summary for the homeowner, which documents what was done and why, and serves as the primary reference document for future warranty or inspection questions.

Each of these is open-ended work that requires judgment and context. Each is currently consuming time that could be recovered. Each is a task where a well-written brief and a two-to-five-minute review checkpoint would convert forty minutes of production into eight minutes of oversight.

The Anthropic data points toward a specific trajectory: the tasks that AI can handle are growing longer, more complex, and more judgment-intensive at a rate that doubles the effective task length every four months. Two years ago, the ceiling was four minutes. Today it is twelve hours. That trajectory does not mean businesses should hand over twelve-hour tasks immediately. It means the category of work that is genuinely delegatable is expanding faster than most businesses are moving to capture it.

The playbook here is not about a wholesale reorganization of how work gets done. It is about moving one task at a time from "I do this myself" to "I write the brief and review the output," running each transition until it is stable, and then identifying the next one.

The businesses that are moving systematically through that list today will have recovered material hours at material volume before their competitors recognize what happened. The advantage is not in the model. It is in the discipline of actually identifying which open-ended tasks to delegate, writing briefs that allow good judgment calls, and building review checkpoints that make the whole system safe to run fast. None of those steps require a technical background. They require the same judgment you already use to manage the work itself.

The Anthropic data is not a projection. It is a measurement of what is already happening inside one of the organizations that builds these systems. The improvement rate is fast enough that tasks handed over today at 76 percent reliability will run at higher reliability in six months without the business having to do anything differently. That is not a reason to wait. It is a reason to start now, while the effort of establishing the habit is still small relative to the payoff it will produce as the underlying capability continues to compound.

Do it with an expert
You can build this yourself, or have it set up right the first time.

That is exactly what we do at AI DOERS. Book a private 30-minute call with Madhuranjan Kumar and we will map the fastest path to it for your specific business.

Book your call →
Madhuranjan Kumar

Madhuranjan Kumar

Founder, AI DOERS · Performance Marketing

Madhuranjan Kumar brings 20 years of performance-marketing experience and has managed over $200 million in Facebook ad spend for brands across the United States and beyond. His expertise spans the full modern marketing stack: Meta, Google Ads, TikTok, email automation, CRM, and the websites that hold it together. At AI DOERS he turns that track record into lead-generation systems for businesses across every industry.

← Back to all insights
Anthropic's Own Data Makes The Strongest Case Yet That AGI Is Already Here | AI Doers