AI DOERS
Book a Call
← All insightsAI Excellence

The One AI Skill That Actually Matters: Giving an Agent the Right Context

AI is collapsing fifty tools into one. The only skill worth mastering is knowing what context to feed your agent and how to check its work.

The One AI Skill That Actually Matters: Giving an Agent the Right Context
Illustration: AI DOERS Studio

Before the first context-loading experiment, the firm's paralegal was spending approximately twelve hours per week on first-draft work that none of the three partners read closely enough to learn from. The drafts were competent. They were produced from templates that had not changed in four years. They were reviewed quickly, edited occasionally, and filed. The paralegal had suggested twice that the templates needed updating. The suggestions had been acknowledged and deferred. There was always a more pressing matter.

This is a recognizable pattern in small legal practices: work that is systematic enough to be templated but consequential enough to require attorney review, occupying an uncomfortable position in the firm's workflows. Too important to automate entirely, too routine to deserve the partners' best attention. It accumulates, produces friction, and represents a real bottleneck when client volume increases. The firm, three lawyers and one paralegal serving primarily small and mid-sized businesses, started an experiment with a context-fed AI agent in early spring. What they learned over twelve weeks changed the practice's staffing conversation in ways that surprised everyone involved.

The question the experiment was designed to answer was narrow and practical: how much does output quality improve when the agent is given the firm's actual standards rather than a generic prompt? The answer was enough to change the firm's staffing calculus for the following year.

The Before State: Drafting Work That Consumed a Third of the Week

The paralegal's workload before the experiment broke into three main categories. Engagement letters, the firm's most frequent first-touch document, were produced at an average rate of about eight per week. Each required approximately forty-five minutes to draft, review against the client intake notes, and prepare for partner sign-off. Deposition summaries ran longer: an average of two per week at approximately three hours each. Contract deviation flagging, reviewing incoming client contracts for non-standard terms, ran at roughly four contracts per week at about ninety minutes each.

The weekly totals: eight engagement letters at forty-five minutes equals six hours; two deposition summaries at three hours equals six hours; four contract reviews at ninety minutes equals six hours. Eighteen hours per week of drafting-category work, representing close to half the paralegal's total billable hours and a substantial share of the firm's administrative overhead.

The partners spent an average of about twenty-five minutes per engagement letter on review, ten to twelve minutes per deposition summary, and thirty to forty minutes per contract flagging report. This added another five to six hours of partner time per week to tasks that were, in the partners' own assessment, not the work they had trained for. The firm was not unhappy in a dramatic sense. The cost structure was simply the reality of operating at their scale, and none of them had a clear picture of what it looked like until they added it up.

The motivation for the experiment was pragmatic. One of the partners had attended a continuing legal education session where an AI agent demonstration was included almost as an afterthought. The demonstration was not sophisticated. But it prompted a question: what would happen if the agent knew the firm's actual document standards before being asked to draft?

How it works

Week One: The Context-Loading Experiment

The experiment started with the engagement letter, because it was the most frequent document and the most templated. The paralegal spent one afternoon loading context: the firm's standard engagement letter in three versions (business litigation, contract advisory, and employment matters), five examples of letters the partners had edited heavily with the edits preserved so the agent could see what had changed and why, the firm's standard fee schedule, and a plain-language description of the firm's approach to scope limitation and client communication.

The first request was simple: draft an engagement letter for a new business litigation client using the intake notes from that morning. The output was not ready to file. But it was closer to correct than the template-starting-point drafts the paralegal had been producing, because it had already incorporated the specific facts from the intake notes rather than leaving placeholder language to fill in manually. The paralegal's editing pass took twelve minutes instead of forty-five.

Over the following three days, the paralegal loaded additional context for deposition summaries: a guide to the firm's preferred summary structure, three examples of summaries the partners had approved without revision, and notes about the two most common errors in previous drafts. The first error was missing the key timeline dispute. The second was underweighting the witness's statements about internal communications. The agent, given the deposition transcript and the context folder, produced a summary that the reviewing partner described as ready with minor adjustments. The minor adjustments took eighteen minutes. The prior baseline was three hours.

The partners noted two things about the first week. The output quality was higher than they expected. And the quality came directly from the quality of the context provided, not from the sophistication of the prompting. The firm had excellent standards. They had simply never externalized those standards in a form that any tool or new team member could use directly. The context-loading exercise was valuable independent of the AI it was built for.

Tools needed for daily work

The Two-Bucket Discovery That Changed How the Firm Assigned Work

By week four, a pattern had emerged that the partners had not anticipated when they started. The tasks the agent handled well were not simply the routine ones. They were the tasks where the agent could do the work itself: receive information, apply standards, produce a document. The tasks where the agent underperformed were ones where the work was fundamentally relational or judgment-dependent in ways that required a human to be present in the room or on the call.

The firm began calling these the two buckets informally. Bucket one: tasks the agent could execute fully, with human review at the end. Bucket two: tasks where the agent could prepare and coach, but a human had to execute the final step.

Contract deviation flagging was a clear bucket-one task. The agent could read a contract, compare it against the firm's standard term set, identify every deviation, and produce a flagging report the reviewing attorney could use to conduct the client conversation. The attorney's job became reviewing the report and deciding how to address the deviations, not finding the deviations in the first place. The time saving was meaningful: the attorney was doing strategic work rather than extraction work.

Client intake calls were a clear bucket-two task. The agent could not conduct a call. But it could prepare the attorney by summarizing the prospective client's email correspondence, identifying key risk factors in the described situation, listing questions the intake form had not answered, and suggesting analogous matters the firm had handled previously. The attorney went into the call better prepared. The call itself remained human. The preparation was not.

This two-bucket framework changed how the firm thought about work allocation across the board, not just for AI-assisted tasks. It surfaced a third category that had previously been invisible: tasks sitting in bucket two that were being done as if they were bucket one, where a human was doing extraction work the agent could handle, freeing the human for the judgment step that actually required their presence.

How the Paralegal's Role Changed in Eight Weeks

By week eight, the paralegal was spending approximately four hours per week on drafting tasks that had previously consumed eighteen hours. The drafting itself had not disappeared. It had changed character. The paralegal was now primarily a context curator and review editor rather than a first-drafter.

The context curation role had not existed before the experiment. Someone needed to maintain the context folders: update them when the partners adopted new fee structures, add exemplary documents as they were produced, annotate examples with notes about why they worked, and periodically remove outdated material. This took about two hours per week and required genuine judgment about what was worth preserving. The paralegal took on this work naturally, without being asked, because the connection between context quality and output quality was direct and visible.

The review editing role also changed character. Editing an agent-produced first draft differs from editing a human-produced first draft in one specific way: the errors are more consistent. A human drafter makes idiosyncratic errors based on personal habits and assumptions. An agent makes systematic errors based on gaps or ambiguities in its context. When you identify one of the agent's systematic errors and correct it in the context folder, the error stops appearing across all subsequent outputs. The paralegal was fixing categories of error rather than individual instances. This is more efficient and also more substantively interesting than catching the same idiosyncratic human errors repeatedly.

The partners also noticed a change in the questions the paralegal brought to them. Instead of clarifying questions about document format and filing procedures, the paralegal was asking substantive questions: why certain terms were standard in the engagement letters, which case outcomes had led to changes in how the firm handled deposition summaries, how the partners wanted to address a type of deviation that kept appearing in contracts from one industry sector. These were questions the context-curation role required them to resolve. The paralegal was engaging with the firm's institutional knowledge rather than executing around it.

Week Twelve: The Numbers That Changed the Staffing Conversation

At the end of week twelve, the firm ran a formal accounting of the time changes. Paralegal drafting time had fallen from eighteen hours per week to four hours. Partner review time had fallen from five to six hours per week to approximately two hours, because the first drafts arriving for review were already closer to final versions. Total staff time devoted to drafting-category work had fallen from approximately twenty-three hours per week to six hours.

The seventeen hours recovered per week translated to approximately 680 hours over a full year at the current staffing model. At the paralegal's billing rate, the recovered time represented a capacity increase of roughly 17,000 dollars in billable hours annually, assuming that time could be directed to billable work. The reality was more complicated: some of the recovered paralegal time had gone into context curation (non-billable) and some into client intake preparation (billable). The net billable gain in the first three months was closer to 12,000 dollars, which the firm considered a conservative estimate for year one.

The staffing conversation that changed was about the next hire. The firm had been planning to bring on a second paralegal to handle the growth in client volume over the past two years. That hire was deferred for at least six months, not because the current paralegal was now doing twice the work, but because the capacity that would have required a second paralegal was now being handled by the agent. The deferral represented a salary and benefits cost of approximately 55,000 dollars for the year.

The partners' assessment at the end of week twelve was more nuanced than the numbers suggested. One partner noted that the reduced drafting time had surfaced a set of judgment calls that had previously been buried in routine work, and that these judgment calls were harder to make, not easier. The cases where the agent's output required significant rethinking were precisely the cases with the most legal complexity, and those cases now arrived at the partners' desks with less buffering than before. The agent had made the routine faster. It had made the complex more visible. That was not a problem exactly. But it required a different kind of management attention than the firm had previously needed.

The firm's conclusion, discussed at the end of the twelve weeks, was that the experiment had succeeded at what they had intended to test. The context-fed agent had handled bucket-one tasks reliably. It had accelerated bucket-two preparation. It had surfaced the distinction between the two buckets in a way that had permanent value for how the firm thought about work allocation going forward. The one skill that had mattered throughout, the only one the partners had needed to develop to make the system work, was knowing what context to provide and how to review what came back. Madhuranjan Kumar notes this as the consistent finding across every context-fed AI experiment he has mapped: the technical setup is simple, and the skill that separates the operators who get results from those who do not is almost always the quality of the context they prepared before the first prompt.

The question the experiment was designed to answer was narrow and practical: how much does output quality improve when the agent is given the firm's actual standards rather than a generic prompt? The answer was enough to change the firm's staffing calculus for the following year.

Do it with an expert
You can build this yourself, or have it set up right the first time.

That is exactly what we do at AI DOERS. Book a private 30-minute call with Madhuranjan Kumar and we will map the fastest path to it for your specific business.

Book your call →
Madhuranjan Kumar

Madhuranjan Kumar

Founder, AI DOERS · Performance Marketing

Madhuranjan Kumar brings 20 years of performance-marketing experience and has managed over $200 million in Facebook ad spend for brands across the United States and beyond. His expertise spans the full modern marketing stack: Meta, Google Ads, TikTok, email automation, CRM, and the websites that hold it together. At AI DOERS he turns that track record into lead-generation systems for businesses across every industry.

← Back to all insights
The One AI Skill That Actually Matters: Giving an Agent the Right Context | AI Doers