How to Build a Team of AI Researchers That Work Together Automatically
A single research agent gives uneven results. A director, manager, and researcher working together deliver consistent quality and fill a spreadsheet on their own. Here is how it works for a consulting firm.

The most expensive mistake in AI-assisted research is not the wrong answer. It is the wrong answer delivered with total confidence, accepted without verification, and embedded in a client report before anyone catches it. The problem is not that AI research tools fail. The problem is that they fail in ways that look like success: fluent, well-structured responses that are wrong in subtle ways a reader cannot identify without independently knowing the subject.
A single research agent, given a hard question, will often return the first plausible-sounding answer it finds rather than the most accurate one. It will accept a secondary source over a primary one if the secondary source is easier to retrieve. It will fill in details it cannot verify with details that fit the pattern of the question. And it will present all of it in the same confident tone it uses for verified facts. This is the quality problem that the three-agent architecture solves, not by making the researcher more capable, but by placing a demanding reviewer between the research and the final output.
What a single research agent actually does when the answer is hard to find
A single research agent has no external standard to meet. It calibrates its own quality threshold based on the instructions it received, and when those instructions are vague or when the task is genuinely difficult, the threshold drifts lower. The agent is optimizing for completion. When an answer is hard to find, the fastest path to completion is accepting the most plausible available answer and moving on. That is precisely what it does.
The failure mode looks like this in practice. A consulting firm needs to know the pricing model for a software company that does not advertise prices publicly. The single agent searches the company's website, finds no pricing page, searches third-party review sites, finds a price range mentioned in a year-old post by someone whose claims are unverified, and returns that range as the answer. It cites the review site. The answer is plausible enough to pass a quick read. It enters the competitive intelligence document. Two weeks later, when the consulting firm references that pricing in a conversation with the software company, they learn the number was wrong by a significant margin.
The single agent did not invent the number. It found a source. But it applied no judgment about source quality, recency, or reliability. It had no mechanism for recognizing that the answer was inadequate and the search should continue with more specific sources. Without that mechanism, it defaulted to the nearest available approximation. This is not a failure of intelligence. It is the predictable behavior of a system with no external quality check and no instruction about what to do when a definitive source is not found.
The pattern repeats consistently across research tasks of meaningful difficulty. The agent produces results that look complete at a surface scan. The errors are in the details that matter most: the specificity of a pricing tier, the recency of a funding event, the accuracy of a client category description. These are exactly the details that make research useful in a client context, and they are exactly the details where a single agent without a reviewer will cut corners invisibly.

The manager agent is not a feature, it is the quality mechanism
The function of the manager agent in a three-agent research architecture is to be the external quality check the researcher cannot apply to itself. The manager reads the researcher's output, evaluates it against the research goal, and decides whether it meets the standard or whether the researcher needs to try again with more specific guidance.
What "pushing back" actually means in this context is not a general instruction to try harder. It is a specific description of what is missing and where better information is likely to exist. A manager that responds "the answer is incomplete, please try again" produces a marginally different second attempt. A manager that responds "the pricing page is not publicly listed, but the company publishes a partner pricing guide through their reseller program, check their partner portal and their recent press releases for any deal size mentions" produces a meaningfully better second attempt because it redirects the researcher to more productive sources with specific guidance about what to look for.
The specificity of the pushback is what makes the manager valuable. It requires the manager to be designed with enough domain awareness to know what constitutes a complete answer for this type of research task and where such an answer is likely to exist. A manager prompt that instructs the reviewer to demand primary sources over secondary ones, to reject any answer that cannot be traced to a specific dated document or direct page content, and to provide a concrete alternative search direction when rejecting an answer will consistently produce better outputs than a researcher working alone against its own judgment of what is good enough.
This is the mechanism that makes research trustworthy. Not a smarter researcher, but a demanding reviewer that refuses to let weak answers through and tells the researcher specifically where to look next. The combination produces results that a single agent cannot reliably achieve because it has no mechanism for recognizing its own inadequacy on a particular question and no external force requiring improvement.

Why the director-reads-list-writes-results pattern keeps memory from collapsing at scale
A research team processing hundreds of items faces a memory problem that does not appear in small test batches. Language models have a context window, a limit on how much information they can hold in working memory at once. A single agent processing a list of a hundred items by loading the whole list, generating all research, and writing results at the end will eventually run into that window limit. The later items in the list receive less thorough attention than the early ones because the model is maintaining context for everything processed so far.
The director-reads-list-writes-results pattern solves this by keeping the active memory lean throughout the run. The director reads one item at a time from the list, delegates it to the research team, receives the completed result, writes it back to the output spreadsheet, and then releases that item from working memory before moving to the next one. At any point in the run, the active context contains only the current research task, the manager-researcher exchange for that task, and the output being written. Completed items are stored in the spreadsheet, not in the model's memory.
This pattern allows the team to process a list of five hundred items with the same quality on item four hundred and ninety as on item three. The quality does not degrade with list length because the memory footprint stays constant. Each item is processed fresh, in its own context, against the same quality standard enforced by the same manager prompt. Without this architecture, a long list produces a well-researched first section and an increasingly shallow second section as the model's attention narrows under context pressure.
The pattern also makes the system resilient to interruption. If the run is stopped partway through, the completed results are already written to the spreadsheet. Resuming from the last incomplete row requires no reconstruction of what was already processed. This resilience matters for research runs that take several hours, where the probability of some interruption is not trivial.
Calibrating the manager: the tradeoff between strict and lenient review
The quality of the entire research team depends on where the manager's standard is set. This calibration is the most consequential design decision in the three-agent architecture and the one most commonly neglected in favor of a default prompt that sounds appropriately demanding without specifying what demanding actually means for the specific task.
A manager set too strictly generates retry loops on answers that are already adequate. If the manager demands three independent primary sources and a direct quote for every finding, the researcher will produce unnecessary retry loops even for questions where one clear, sourced answer from the company's own website is genuinely the best available. Each unnecessary loop consumes model usage, increases the time per item, and raises the cost of the run without proportionate quality gains. A strict manager miscalibrated to the actual task will produce a research run that is expensive, slow, and produces results that are no more accurate than a reasonably calibrated manager would have produced in a fraction of the time.
A manager set too leniently passes weak answers through and the overall quality stays low across the full run. If the manager approves any answer that includes a citation regardless of citation quality or recency, the researcher learns that citing a low-quality source is sufficient. The output list will be consistently filled but will contain a predictable rate of weak or inaccurate findings that require manual review to identify, which defeats much of the purpose of running the team in the first place.
The right calibration depends on the specific research task and the cost of a wrong answer in that context. For a prospect list where a wrong industry classification means an irrelevant sales call, a moderately strict manager that demands answers be sourced to the company's own website or a verified business database is appropriate. For competitive intelligence informing a client's pricing strategy, a strict manager that demands primary sources and rejects secondary commentary is appropriate. For general background research where approximate information is useful, a more lenient manager accepting reputable secondary sources is appropriate. Determining which level fits your task requires running the team on a small set you can verify manually before committing to the full list.
A consulting firm profiles 40 competitors overnight
Madhuranjan Kumar would set this up for a consulting firm client in the following way. The engagement requires competitive intelligence across forty companies, each profiled for funding stage, pricing model, primary product positioning, and main client categories. A junior analyst doing this manually would spend two to three days searching websites, reading press coverage, checking funding databases, and formatting findings. The research team processes the same list in a few hours, running overnight.
The director reads the list of forty companies from a spreadsheet, one row at a time. For each company, it passes the research task to the manager and researcher with a clear brief specifying exactly what is needed: the funding stage with a date and source, the pricing model described with any available tier information, the top three client categories stated in the company's own language from their website or case studies, and a direct link to the primary source for each finding.
The researcher searches the company website, the company's LinkedIn page, recent press releases, and relevant funding databases. For a company whose pricing is not publicly listed, the researcher checks the partner portal, looks for deal size mentions in press coverage, and notes explicitly that direct pricing was not publicly available rather than inferring a number from secondary estimates. The manager reads the findings and evaluates them against the brief. If the client category information came from a generic about-page description rather than specific case studies or client lists, the manager pushes back and directs the researcher to the company's case study library or client testimonial page.
When the manager approves the findings for a company, the director writes the completed profile to that row in the spreadsheet and moves immediately to the next company. By morning, the spreadsheet has forty completed rows. The consulting team reviews a sample, identifies two companies where pricing was noted as unavailable and requires manual follow-up, validates the rest against their own industry knowledge, and has the competitive intelligence foundation ready for analysis the same day.
The research phase that previously consumed two to three days of analyst time completes overnight. The analysts begin interpretation and recommendation work the next morning rather than the day after. For a consulting firm billing by the hour, that shift in how time is allocated from gathering to thinking is a direct improvement in both the quality and the profitability of the engagement.
The calibration test you must run before scaling to hundreds of rows
Every research team needs a calibration test on a small, manually verifiable set before it is trusted at scale. The test is simple in design and essential in practice.
Take ten to twenty items from your list where you already know the correct answers or can verify them quickly through your own manual research. Run the team on those items. For each one, compare the team's output to what you know is correct and evaluate three things: accuracy, source quality, and the number of retry loops the manager generated for each item.
Accuracy tells you whether the researcher is finding real information or producing plausible approximations that pass casual review. Source quality tells you whether the manager's standard is demanding enough to require primary sources or lenient enough to accept secondary commentary. Retry loop count tells you whether the manager is calibrated appropriately for the task or generating friction that does not improve the final answer.
If accuracy is low, the researcher's prompt needs more specific guidance about where to look and what constitutes an acceptable answer for this type of question. If source quality is low despite reasonable accuracy, the manager's standard needs to be raised with more explicit requirements about source type and recency. If retry loops are excessive without proportionate quality gains, the manager's standard needs to be loosened slightly or made more precise about what it actually requires rather than being broadly demanding.
This calibration test costs perhaps an hour of time and a small amount of model usage. What it prevents is a full run of hundreds of rows that produces output at the wrong quality level, either too low to use in a client context or too aggressively reviewed to complete in a reasonable time and cost. Running the test before the full list is the investment that makes the full run reliable. Skipping it to save an hour is a decision that costs more than an hour to correct when the full run produces the wrong output at scale.
That is exactly what we do at AI DOERS. Book a private 30-minute call with Madhuranjan Kumar and we will map the fastest path to it for your specific business.
Book your call →
