How to Run an Always-On AI Agent for Your Business
An always-on AI agent lives on a cheap server or mini PC and works around the clock, answering customers, drafting follow-ups, and remembering your business. Here is how it works and how I would set one up.

Ninety-three percent of prospects who message a service business after hours have moved on to a competitor by morning if no one responds. That figure is easy to absorb in the abstract and painful to live with once you have spent Sunday evenings calculating how many warm leads your Monday morning replies were too late to convert. This is the story of a medical aesthetics clinic that set up an always-on AI agent to solve exactly that problem, what the first two weeks looked like, where they ran into trouble, and where the system sits six months in.
Before the agent: what three hours of after-hours admin looked like each week
The clinic offers botox, filler, microneedling, and a suite of laser treatments. It closes at six in the evening on weekdays and does not operate on weekends. The owner is the sole practitioner and runs the business with one part-time front-desk coordinator.
Before the agent, three hours of the owner's Sunday evenings went to clearing the backlog of messages that had arrived since Friday afternoon. On a typical weekend, seven to twelve messages came in through the clinic's website contact form and through the Facebook message channel tied to her Facebook and Instagram ad campaigns. The messages sorted into predictable categories. A large portion were treatment questions: someone wanted to know which service addressed their specific concern, whether dark spots, lip asymmetry, or early jowling, along with the price and recovery time. A smaller portion were follow-up messages from people who had attended consultations but not yet booked, circling back to ask whether a particular slot was still available. A handful were general inquiries about the clinic's process for first-time visitors.
None of those messages required clinical judgment. They required three things: knowledge of the treatment menu and prices, a clear explanation of what each treatment addresses, and the ability to reply in a way that captured the person's contact details for a morning call. The owner had all of that memorized. What she did not have was a way to send it at 9:30 on a Friday night without sitting down at her desk.
The hidden cost was not the three hours. It was the leads that slipped. She estimated her conversion rate on Friday and Saturday messages hovered around thirty percent, because a meaningful share of those contacts had already moved on before she reached them on Sunday. She knew this not from modeling but from asking new clients during consultations how they had found the clinic. A number named a competitor they had considered first, which confirmed the loss was real rather than theoretical.
There was a second category of after-hours work running in parallel: review requests. After each completed treatment, the owner sent a short personalized note asking the client to share their experience. She wrote these by hand, one at a time, and sent them when she had time, which meant some clients received a request within two days of their appointment and others waited three weeks. The inconsistency produced inconsistent review velocity. She was performing a high volume of treatments and seeing lower review accumulation than the volume warranted.
Three hours per week. One part-time coordinator handling daytime traffic. An after-hours window from six in the evening to ten in the morning where the only thing responding to a warm prospect was silence.

The decision to run something that never clocks out
The owner had read about AI tools in a general way for several months without acting on any of them. The change came from a specific conversation. A regular client who worked at a software company asked, at the end of a consultation, whether the clinic used any AI tools for operations. The owner said she had been exploring options but had not found anything that seemed practical without a significant technical setup.
The client mentioned she ran a small agent on a rented server that handled the first reply in every inquiry she received for her freelance work. The agent did not do everything, and she reviewed its output regularly, but it meant no inquiry sat unanswered overnight. The server cost was around eighteen to twenty dollars a month. The agent learned from plain text files she had written herself, describing her work and how she liked to communicate.
The owner asked one follow-up question: could it learn a price list and treatment menu. The client's answer was yes. You write the information into a text file and point the agent to it as part of its memory. There was no specialized technical background required. One command installed it, and one afternoon of writing the memory files made it sound like a competent clinic assistant rather than a generic chatbot.
That conversation happened in the second week of March. The agent was installed by mid-April.

The first two weeks: a small server, one command, and a very specific first task
The owner chose a rented virtual private server rather than a local mini PC because she did not want to manage hardware. The server runs continuously on hosted infrastructure maintained by the provider, and she pays a flat monthly rate of eighteen dollars. Installation was a single command pasted into a terminal window, followed by a brief setup flow where she selected a model provider, set an API key with a twenty-dollar monthly spending cap, and connected the agent to a dedicated Telegram account she had created for the clinic.
The first task she assigned was deliberately narrow: answer factual questions about services, prices, and wait times from the website contact form and the Facebook message channel. No booking changes. No payments. No clinical recommendations. Just information and lead capture.
She spent roughly two hours on the three memory files that determine how the agent behaves in practice. The personality-and-rules file told the agent it was the clinic's front-desk assistant, warm and direct with visitors, not a diagnostic tool, and that its purpose in any conversation was to answer factual questions clearly and capture the visitor's name and phone number so a staff member could reach them the next morning. The business context file held the complete service menu, current prices, brief descriptions of what each treatment addresses, and a FAQ drawn from the thirty most common questions she had received over the past year. The running to-do file started blank.
In the first two weeks she reviewed every reply the agent drafted before it went out. Fourteen messages came in during week one and eleven during week two. Of the twenty-five, twelve were handled cleanly. Eight needed small corrections in tone or a factual detail, and each correction was written into the rules file so the same situation would be handled better the next time. Five fell outside the agent's defined scope, typically specific questions about skin conditions or medical contraindications, and for those the agent sent a brief acknowledgment saying a staff member would be in touch.
She made no changes to the agent's permitted scope during those two weeks. She focused entirely on refining the memory files based on what she observed. By day fourteen, the most recent twelve messages had required no correction.
Month one results: 26 after-hours inquiries handled, zero dropped leads
At the end of the first full month she reviewed the numbers. The agent had handled twenty-six after-hours inquiries independently. The clinic's three-month historical average for total after-hours inquiries was thirty-one per month, and the remaining five had arrived during staffed hours and been handled by the coordinator normally.
Of the twenty-six, twenty-two received a reply within four minutes of the original message. Four arrived during a brief server maintenance window and waited approximately fifteen minutes. Zero went overnight without a response.
The agent captured a name and phone number in nineteen of the twenty-six conversations, either because the visitor offered their details unprompted or because the agent asked at the natural close of the exchange. The remaining seven did not respond after the initial reply or said they would call.
Of the nineteen captured contacts, the front desk converted eight to booked consultations at the morning follow-up call. A forty-two percent conversion rate on overnight captured contacts, which the owner notes is higher than her typical walk-in conversion rate. Her interpretation is that the overnight exchange had already resolved the prospect's main questions, so by the time the front desk called the next morning, the person was past the information-gathering stage and closer to a decision.
She also ran review requests through the agent that month. After each completed treatment, the front desk marked the client as complete in the scheduling log. Once a week, the agent read that log, identified clients marked complete but not yet sent a review request, drafted a short personalized message for each one, and placed the messages in a staff review queue for approval before sending. Eleven review requests went out through this process. Four produced new reviews on the clinic's Google profile.
All contact data gathered by the agent synced into the clinic's CRM and website stack, giving the coordinator a complete view of each overnight conversation before making the morning call.
What the owner learned about security the hard way, and the rule they now follow
At the end of month one, with the first task running cleanly, the owner gave the agent access to the booking software. She connected it to the scheduling API and granted read-and-write permission, intending for it to check appointment availability and confirm times during after-hours conversations. The logic seemed straightforward: if the agent could answer questions and confirm a slot in one exchange, conversion would be higher.
Within the first week of expanded access, an appointment appeared in the scheduling software that no one on staff had changed. A regular client's Tuesday afternoon slot had moved to Thursday morning. The agent had received a message that looked like a reschedule request, interpreted it as genuine, and used the write permission it held to process the change. The message had almost certainly come from someone probing or testing the system rather than an actual client, but the agent had no mechanism to make that distinction. It had a job, a permission, and it did the job.
The owner found the error during her daily log review the following morning. The client was reached before the original appointment date, the Tuesday slot was restored, and the client received a small service credit for the confusion. No serious harm occurred. But the sequence was clarifying: a speculative permission granted because it seemed useful in theory had been used in a way that required a recovery call.
She removed the write permission from the booking software the same morning. The agent can now query availability but cannot modify the schedule. Any modification request gets flagged and placed in a staff review queue. She then audited every other permission the agent held and tightened each one. The API spending cap dropped from twenty dollars to eight, because six months of real usage had never exceeded six. She removed access to the staff contact list, which had been added speculatively and used once in a way she had not intended.
The rule she follows now is stated simply: the agent gets the lowest permission level that allows it to complete its defined job and no higher. Read access where read is sufficient. No access where a task has not been tested in a controlled review queue for at least two weeks. Permissions expand only after evidence of safe behavior, not before.
Where the agent lives six months in
Six months after the first installation, three agents are running on the same server. Each has a single focused job.
The original front-desk inquiry agent handles overnight messages across the contact form and the Facebook channel. It is the most active of the three and is still refined monthly through reviews of the rules file. A second agent runs once each week and produces a structured summary of every inquiry, treatment booked, and review request from the past seven days. The owner reads it Monday morning in about five minutes rather than pulling separate reports from the scheduling tool, the contact form inbox, and the message channel individually. A third agent handles treatment-question drafts: when a client messages about combining two procedures or asks what to expect in recovery, the agent drafts a reply using the clinic's aftercare documentation and places it in a staff review queue for a final check before it goes out.
Monthly operating cost across all three agents: thirty-one dollars. Eighteen for the server and thirteen in model API calls based on actual monthly usage tracked over the past five months. The owner estimates the three agents recover approximately nine staff hours per month at a blended staff rate, putting direct labor savings at just under two hundred dollars a month against a thirty-one-dollar operating cost.
Her after-hours lead recovery rate, measured as the proportion of overnight inquiries that result in a booked consultation or a live phone conversation within twenty-four hours, has moved from roughly thirty percent before the agent to sixty-eight percent now. The agent did not create new demand. It captured demand that was already arriving and had been slipping past the clinic while it was closed.
The most common question she receives from other service business owners is about the technical side: which hosting provider, which model, which connector. Her consistent answer is that the technical setup takes under an hour and is the least important part of the project. The memory files are what matter. Writing them carefully, reading every reply the agent sends during the first two weeks, and correcting the rules file each time something is off is what produces an agent that sounds like a competent front desk rather than a confused system. Start with the narrowest possible first task. Keep all permissions as low as they can be while still allowing that task to be done. Expand only after two weeks of clean operation.
That is the full pattern she would repeat if she were starting over, and the one she recommends to every other service business owner who asks how it works.
That is exactly what we do at AI DOERS. Book a private 30-minute call with Madhuranjan Kumar and we will map the fastest path to it for your specific business.
Book your call →
