AI DOERS
Book a Call
← All insightsSearch & Video

The AI Video Breakthroughs That Could Transform How Hair Salons Attract New Clients

Seven research papers published in the past month reveal that virtual try-on, automatic video green screening, and AI-generated talking portraits are converging into tools that beauty businesses can use for marketing content today.

The AI Video Breakthroughs That Could Transform How Hair Salons Attract New Clients
Illustration: AI DOERS Studio

Seven research papers landed in a short window, each solving a different problem in AI video. Taken individually, each one is interesting. Taken together, they describe a near-complete toolkit for creating professional visual content without a camera crew, a studio, or a significant production budget, and that convergence is the actual news.

What virtual try-on research actually solves that professional product photography cannot

Professional product photography solves one problem well: it shows a garment on a model, in a studio, under controlled lighting, looking exactly as the brand wants it to look. It does not solve the problem the customer actually has, which is whether the product will look right on them, with their body, their coloring, their specific proportions. That gap between the studio image and the fitting room reality is where a significant percentage of e-commerce returns originate, and it is the problem that virtual try-on research is directly attacking.

CatVITON, one of the seven papers, approaches this with a lightweight model that takes two inputs: a photograph of a person and an image of a clothing item. The model analyzes body pose, proportions, and lighting conditions from the person photograph, then renders the garment onto that person while preserving their identity, expression, and the original lighting conditions of the source image. The result is a preview that shows the customer what the item would look like on their specific body rather than on a model who may share little with their actual proportions. The model is open-source, runs without prohibitive hardware requirements, and can be deployed as a web service at very low cost per query, making it practical for consumer-facing use at scale.

Any-to-Any VTON goes further by accepting multiple garment items simultaneously and allowing text-guided modifications. A customer can upload a photo, select a jacket and trousers, and specify that the jacket collar should be modified or the color should shift toward a different shade. In direct comparisons with CatVITON, Any-to-Any VTON produces higher-fidelity outputs and removes the complex masking and pose estimation steps that made earlier try-on systems brittle under real-world photo conditions. These two papers together represent the practical resolution of a customer experience problem that product photography was never designed to solve. The question they raise for any product business is not whether to use this technology but when, because the integration cost is low enough that delay is the more expensive choice once the model quality is sufficient for real consumer use.

How a Hair Salon Uses AI Video for Marketing

How automatic background removal at scale changes the economics of creative content production

The second set of papers addresses a different but equally significant problem: the cost and complexity of isolating subjects from backgrounds in video. This has historically required either a physical green screen setup during filming or painstaking manual rotoscoping in post-production. Both approaches are expensive in time, equipment, and specialist skill. They are the reason professional video production stays out of reach for most small businesses even as the cameras themselves have become cheap enough to sit in every pocket.

Diff Eraser and Matte Anyone together change this economics. Diff Eraser is a diffusion-based video inpainting model that masks out any person or object in a video and fills in what would have been visible behind them, resolving the ghosting artifacts that plagued earlier removal tools and making the filled result hold up cleanly across motion. Matte Anyone is a video matting system that takes any video, lets you define a subject, and exports a clean alpha-channel version automatically, treating the subject as if they had been filmed against a green screen regardless of the actual background. Both tools work on footage already captured, which means a business does not need to plan for green-screen during filming. The capability applies retroactively to existing video libraries.

For any business producing video content regularly, this collapses one of the largest fixed costs in professional video: the logistics of controlled-background filming. A founder can record a product walkthrough against a busy office background, and Matte Anyone produces a clean compositable version that can be placed against a branded backdrop in post. A tutorial filmed in a kitchen can be delivered with a professional-looking background that matches the brand. The equipment requirement for professional-looking video output drops to a reasonably capable phone and access to either of these models, which is a fundamentally different cost structure than the one that ruled video production two years ago.

Monthly Leads from AI Try-On Widget vs Standard Gallery

The convergence point where these seven tools meet inside one production workflow

The reason these seven papers matter together rather than individually is that they describe something approaching a complete end-to-end content production pipeline, with each component addressing a different gap that previously required a separate specialist, a separate tool, and a separate budget line.

The pipeline looks like this in practice. A product image goes through a virtual try-on model so customers can see it on themselves. Raw video footage of a spokesperson gets the background removed with Matte Anyone so it can be composited cleanly against any branded environment. Omnium One, the ByteDance paper, takes a single photograph of the spokesperson and a single audio clip and generates a realistic talking portrait video, eliminating the need for filming altogether in cases where fresh footage is impractical. VideoJam's training approach, applied across video generation platforms, ensures the motion in any generated clip looks physically correct rather than uncanny: gymnasts actually execute gymnastics, movements follow realistic physics, faces behave consistently through motion. Film Agent's multi-agent coordination shows how these components can eventually be orchestrated automatically by a system acting as director, writer, and cinematographer simultaneously.

No single business will need all seven capabilities in one workflow today. But a medium-sized e-commerce brand might realistically use three or four of them in a single month: virtual try-on for the product catalog, automatic background removal for spokesperson clips, talking portrait generation for founder social content, and physics-corrected video generation for promotional clips. Each of those previously required a separate vendor relationship, a separate learning curve, and a separate cost. The convergence is that they now share an underlying approach compatible enough to chain together in one production workflow without specialist overhead at each transition.

Which of the seven papers is furthest from real business use and why that matters for planning

Film Agent is the most impressive paper in academic terms and the furthest from practical business deployment. It coordinates a director agent, a scriptwriter agent, actor agents, and a cinematographer agent inside a Unity 3D environment, with the agents communicating through structured messages to produce a coherent short film. The human evaluation score for plot coherence reached 3.98 out of 5, which is meaningful evidence that multi-agent creative coordination produces coherent narrative output, not just technical artifacts that look like film.

But the gap between a 3D environment running coordinated agents and a tool a business owner opens to produce a promotional video is wide. Film Agent is a research demonstration of a principle, not a product. The Unity 3D dependency means content creation requires 3D scene-building skills that most businesses do not have and do not want to acquire. The multi-agent coordination overhead makes it computationally demanding and slow to iterate on. The output is a CGI short film, not a realistic-looking commercial video. Understanding this is not a criticism of the research but a calibration for planning. Film Agent represents where multi-agent creative coordination is heading, and probably what commercial production tools will look like in three to four years. It is not the paper to act on now. The papers to act on in the next six months are CatVITON, Matte Anyone, and Omnium One.

VideoJam is similarly positioned as a near-term infrastructure improvement rather than an immediate deployment decision. Its contribution is a training approach that corrects motion physics in AI-generated video, and that contribution will show up in the commercial video generation tools that adopt this method rather than as something a business deploys directly. Runway, Kling, and comparable platforms will incorporate this kind of training into their standard models, and the improvement in output quality will arrive for business users as a platform update rather than a new deployment decision. Understanding which papers are infrastructure-layer improvements and which are application-layer capabilities is what makes planning realistic rather than aspirational.

The e-commerce workflow that uses three of these capabilities with specific production numbers

Consider a direct-to-consumer apparel brand selling through its own website and through social channels, with a catalog of forty active products. The current content production process involves one studio shoot per quarter, producing roughly sixty to eighty images across active products at a cost of approximately three thousand dollars per shoot. Between shoots, product photography is static and cannot respond to trend cycles, seasonal styling contexts, or customer feedback about what they want to see on specific products.

The team integrates CatVITON for virtual try-on, deployed as a web tool on the product detail page. A website visitor uploads a selfie, selects a product, and sees a preview of that product on themselves within about twelve seconds. The integration cost is a one-time developer project of approximately eight hundred dollars and ongoing hosting under fifty dollars per month at typical catalog query volumes. In the three months after deployment, the brand tracks conversion rate on try-on-enabled pages against control pages without the tool. Try-on pages convert at a measurably higher rate, and the return rate on orders placed through the try-on preview is lower than the catalog average, because customers who have seen the product on themselves have fewer surprises when it arrives. The return rate reduction alone, across a catalog generating a few hundred orders per month, recovers the integration cost inside the first quarter.

For social content, the team begins using Matte Anyone to process their existing video library. Product walkthrough clips filmed in the studio get new branded backgrounds composited in without reshooting. A founder testimonial clip recorded at a conference gets a clean background so it can be used in paid social without the distracting conference environment behind the speaker. This retroactive processing of existing clips produces more than a dozen new distinct social assets from content already paid for, with editing time of approximately two to three hours per clip compared to four to six hours that a professional editor would charge for manual rotoscoping at previous rates.

For the founder's direct-to-camera social content, the team tests Omnium One during weeks when filming is impractical. A single professional photograph of the founder plus a recorded audio walkthrough produces a talking portrait video suitable for Instagram reels and short-form content. The team produces eight such videos in a single afternoon session, covering product features, styling context, and seasonal themes, that would have required eight separate filming sessions to produce by traditional means. Production time per video drops from approximately ninety minutes of filming and editing to approximately twenty minutes of recording audio and processing. Total impact over a four-month integration period: more social content per month than in any previous quarter, at roughly forty percent of the previous per-asset production cost, all without hiring additional creative staff or expanding the photography budget.

What to do with this information today

Understanding these seven papers as a converging toolkit rather than seven separate research items is the practical takeaway. The toolkit is not fully assembled yet. Some components are production-ready and open-source today. Others are platform improvements that will arrive as updates to the commercial tools already in use. And one, Film Agent, represents a future state that is meaningful to understand as a direction without being actionable today.

The production-ready items to evaluate now: CatVITON for any business that sells products where customer fit or appearance is part of the purchase decision, Matte Anyone for any business producing regular video content that wants more flexibility in post-production without adding green-screen logistics to the filming process, and Omnium One for any founder or spokesperson who wants to produce regular direct-to-camera content without scheduling regular filming sessions. Each of these has a clear integration path, open-source availability, and a cost structure that makes evaluation practical without significant commitment. The right first step is testing the capability on a real example from the actual business before deciding whether to build around it.

Do it with an expert
You can build this yourself, or have it set up right the first time.

That is exactly what we do at AI DOERS. Book a private 30-minute call with Madhuranjan Kumar and we will map the fastest path to it for your specific business.

Book your call →
Madhuranjan Kumar

Madhuranjan Kumar

Founder, AI DOERS · Performance Marketing

Madhuranjan Kumar brings 20 years of performance-marketing experience and has managed over $200 million in Facebook ad spend for brands across the United States and beyond. His expertise spans the full modern marketing stack: Meta, Google Ads, TikTok, email automation, CRM, and the websites that hold it together. At AI DOERS he turns that track record into lead-generation systems for businesses across every industry.

← Back to all insights
The AI Video Breakthroughs That Could Transform How Hair Salons Attract New Clients | AI Doers