Harness Engineering: Why I Let the AI Prompt Me
Harness engineering means building a thin, self-modifying layer that lets a coding agent control real tools, then flipping the workflow so the agent proposes work and you only approve it. One founder now runs an entire startup this way from a chat app.

The most counterintuitive shift I have seen in AI-assisted work is not automating the task. It is reversing who does the prompting. I am Madhuranjan Kumar, and the example I want to trace here is a founder who built a startup on this inversion, and what the mechanics behind it mean for any business willing to think the same way.
The setup is specific enough to be convincing. He connects a 24/7 cloud machine to his email, his team's messaging, his notes, and his code repository. He logs it into a coding agent and gives it one instruction: make this startup successful. From there, every thirty minutes, the agent surfaces a list of proposed actions. He reviews each one on his phone and swipes yes or no, the same gesture as approving or declining a message. What changed is not the technology. What changed is the direction of initiation. He stopped writing a stream of prompts. The agent writes the proposals. He reviews them. That inversion is what harness engineering is trying to make reliable.
The founder who runs a startup from a chat app and why the mechanics behind it matter
Harness engineering is not a new concept in the generic sense. Engineers have always built the scaffold around a capability to make it usable in production. What is new is that the underlying model is capable enough for that scaffold to be thin. The browser harness at the center of this example is approximately 600 lines long. In the old world of AI orchestration, a production-ready harness for a task of comparable scope would have been thousands of lines of rigid, brittle logic that broke every time a web page changed its structure. At 600 lines, the harness is small because the model handles the reasoning that previously required explicit code. The harness defines the permissions and the connections. The model decides what to do within them.
The self-modification property is the part that changes the development dynamic most. Because the agent has access to its own source code, when it hits friction, a page that renders differently, a file-upload flow that was not anticipated, a browser dialog that was not accounted for, it writes the fix into the harness itself. The next session starts with the improved version. The old architecture for this kind of task required three agents: one to act, one to judge, and one to fix. One self-correcting agent with write access to its own harness collapses all three into a single loop that runs roughly ten times faster.
The trust transfer effect that followed this setup is worth paying attention to. Once people observed the agent handling code reviews reliably, the same harness started receiving parking tickets, online shopping tasks, and eventually routine credit-card purchases. That escalation is not evidence of recklessness. It is evidence that the error rate dropped low enough that delegation to adjacent tasks felt reasonable. The harness kept the error rate low. The model's capability made the harness worth building at all.

Reversing the prompting loop: why the agent should be asking you, not the other way around
The most expensive part of working with AI tools right now is not the API cost. It is the cognitive load of writing prompts. Every time a person writes a prompt, they are doing a form of specification work: deciding what to ask, how to frame it, what context to include, and what output to request. For recurring tasks, this specification is rebuilt from scratch every session. Harness engineering eliminates that overhead by setting a high-level goal once and letting the agent figure out the actions.
The instruction to the agent to get him to accept as many ideas as possible is doing more work than it appears. It forces the agent to make the case for each proposal before it is approved. Every item in the queue comes with an explanation of the expected impact. The owner is not rubber-stamping a list. He is reviewing a briefed recommendation and deciding whether it serves the goal. That distinction is the difference between a principal making a judgment and a reviewer waving things through. The former keeps the human genuinely in the loop. The latter produces the worst outcome: the appearance of oversight without the substance of it.
The communication architecture supports the inversion. Email, team chat, notes, and the repository all connect to a single interface. Different topics in that interface become separate agent sessions, each holding its own context and goal. The agent that read a complaint in team chat, composed an email to an external party about a naming conflict, and closed the thread before any human noticed it had opened was operating within exactly this structure: connected to the real systems, given the real context, trusted to initiate the appropriate action, and held to the approval constraint before anything external was sent.
For a business running Facebook and Instagram ad campaigns and managing leads through a CRM, the same architecture produces a concrete daily change. Instead of an analyst manually checking the ad dashboard and drafting a summary, the agent checks the dashboard on its own schedule, identifies budget anomalies, drafts the recommended response, and adds it to the approval queue. The analyst reviews a pre-drafted recommendation rather than starting from a blank monitoring task. The work that used to take forty-five minutes arrives as a two-minute review.
The hardest design decision is the approval step, and the right answer is to never eliminate it for anything client-facing or financially consequential. A drafted follow-up email that a broker reviews and approves before sending is safer than a sent email that the broker never read. An ad budget change the analyst reviewed and approved is defensible. A change that the agent made autonomously because the approval step was removed is not. The harness is what defines the boundary between what the agent can do autonomously and what requires a human check. That boundary should be set at the point where an error has a real cost, not at the point where the approval feels like overhead.
The practical application for a CRM-driven sales or service operation maps directly onto the founder's setup. Connect the inbox, the scheduling system, and the contact records. Set one goal: keep every active lead engaged and every scheduled interaction prepared. Let the agent propose the daily queue. Review the queue each morning. The agent never forgets a follow-up because a meeting ran late. It never misses an account that went quiet because there were more urgent things to handle. It never skips the weekly pipeline report because it was a busy Friday. What changes is not the quality of the work. What changes is that the work happens consistently, which in most businesses is the actual problem. Consistency is the thing that erodes when humans are under pressure. It is the thing the agent does not sacrifice.
The three-to-six-months-ahead framing is the most useful operating principle for anyone building these systems. The founder's current problem is not whether agents can do the work. It is the interface: the right way for a business owner to hand a high-level goal to an agent they trust with real consequences. That interface is not yet fully solved, and the people who are building toward it now, noticing the friction in their own setup and writing the harness piece that addresses it, are the ones who will have a working answer when the rest of the market is still figuring out what question to ask.

That is exactly what we do at AI DOERS. Book a private 30-minute call with Madhuranjan Kumar and we will map the fastest path to it for your specific business.
Book your call →
