The Red Queen Effect: What Self-Improving AI Teaches Small Businesses
AI recently taught itself to beat decades of human strategy simply by competing against copies of itself. The lesson for owners is that the businesses that test, adapt, and keep the winner pull ahead of the ones that copy and stand still.

What self-play actually revealed about learning under pressure
I am Madhuranjan Kumar, and the AI experiment I keep referencing to clients is one most people will never try themselves: putting copies of an AI model inside an old programming game called Core War and having them compete against each other without any human examples to learn from. The setup sounds like a curiosity. The result is a case study in how genuine competitive pressure produces learning that no amount of studying the existing record can replicate.
Core War is a game where small autonomous programs fight to overwrite each other's memory until one survives. Researchers loaded an AI into this environment, gave it no human strategies, and let copies of it battle each other for hundreds of rounds. The model had one signal: whether it survived or died. Everything it learned, it learned from the outcome of its own experiments against an opponent that was getting better at the same rate.
The result that surprised everyone was not that the AI got good at Core War. It was what it got good at. After hundreds of rounds of self-play, the model independently rediscovered strategies that human players had taken decades to develop through community collaboration and accumulated wisdom. It converged on the same solutions without ever seeing them, not because someone told it the right answers but because the right answers were the ones that survived.
This is convergent evolution in learning systems. The same selection pressure, winning the game, produced the same solutions across two completely different developmental paths, human trial-and-error over decades and machine iteration over hours. The implication for how businesses should think about improvement is direct: the answers that work in your market exist and are discoverable through testing. You do not need to copy them from someone who found them first. You can rediscover them through your own iteration, often faster.

Why copying rivals is a losing move from the start
The conventional business wisdom says to study what successful competitors do and adapt it. There is some value in this, but the self-play experiment reveals its ceiling. When you copy a strategy from a competitor, the best possible outcome is that you execute it as well as they do. You start where they are. You are already behind by the time you implement because they have had time to iterate further. And you are operating with a strategy designed for their audience, their positioning, and their specific offer, not yours.
The Red Queen effect is named for a character in Lewis Carroll who must run as fast as she can just to stay in the same place. In a competitive market, the mechanism is the same. Your competitors are not standing still while you copy them. They are also iterating. Copying their current state means you are perpetually one or more iterations behind a target that keeps moving.
The alternative the self-play model demonstrates is more powerful: run your own contest. Rather than studying what others have learned, put your own ideas into competition and let the evidence determine which ones survive. The outcome is a strategy that is calibrated to your specific market, your specific customers, and your specific constraints, rather than to someone else's context. And because it emerged from iteration on real data, it carries higher confidence than any borrowed strategy.
There is a secondary finding from the self-play research that is equally important: the model developed an ability to look at a proposed strategy and estimate how strong it was without running it. This is the equivalent of intuition built from experience, the ability to read the board and know roughly what will happen before committing to a move. In a business context, this is what experienced marketers call having a sense for what will land. The self-play experiment suggests this sense is not innate but is built through rapid, feedback-rich iteration.

Translating the loop from a game to a real marketing channel
The mechanics that produced learning in Core War translate directly into a repeatable improvement loop for any marketing channel. The parallels are precise enough to use as a template.
In Core War: define what winning means, run several programs simultaneously, see which survive, promote the survivors and kill the rest, introduce new challengers in the next round, repeat.
In a marketing channel: define what winning means for this channel, a cost per lead below a threshold, a reply rate above a target, a show rate in a specific range. Generate several genuinely different approaches to this channel, not minor variations but real alternatives that embody different assumptions. Run them against real traffic or real prospects simultaneously. Measure which ones produce results against the defined metric. Kill the losers without sentiment or justification beyond the number. Double the budget or effort on the winner. Then immediately introduce new challengers to test against the champion.
The word simultaneously matters. Sequential testing, one approach tried, then another, introduces time variables that contaminate the results. What performed better: the approach itself, or the season, the news cycle, the day of week? Running variations simultaneously controls for all of these confounders and produces cleaner signal.
The commitment to killing losers also matters. Most businesses test variations, find a winner, and then continue running the losers at reduced scale to hedge. This is a mistake. Every dollar or hour spent on a known loser is a dollar or hour taken from a known winner. The discipline of eliminating bad options immediately is what drives the compounding improvement. Sentiment and sunk-cost thinking are what make most businesses stop improving after one or two iterations.
The compounding return structure: why rounds three through six beat rounds one and two
The data from self-play experiments shows an important pattern: improvement is not linear. The early rounds eliminate the obviously wrong approaches, which is valuable but not dramatic. The middle rounds begin refining among genuinely competitive options, which is where the most significant gains occur. The late rounds produce incremental improvements on what is already working well.
In business terms, this maps to a testing cycle that most teams run too short. A single round of A/B testing produces a winner over a loser. That winner might be 20 percent better than the loser. If the team stops there, they have captured the low-hanging fruit but not the compounding return. A second round that tests the winner against two new challengers might produce another 15 percent improvement. A third round produces 25 percent. By round five or six, the tactics being refined are nothing like the starting point and are performing at two to three times the original baseline.
The businesses that pull ahead in competitive markets are not the ones with the best first-round ideas. They are the ones that run the loop longest. An average idea tested across six rounds of iteration will outperform a brilliant idea tested once, because iteration is the mechanism that discovers what the market actually responds to rather than what sounds compelling in a planning meeting.
For a real estate agency illustrating this with numbers: a single round of testing on listing ad copy that costs 20 percent less per lead is worth capturing. Six rounds of the same discipline, each refining what the previous round produced, can realistically deliver 100 to 150 percent improvement in lead cost efficiency over three to four months. The compounding is real because each round starts from a higher baseline than the previous one.
The practical investment to run six testing rounds over three months is modest. Each round requires generating several variations, which takes two to three hours with AI assistance and one to two hours of review and setup. The cost of running the variants is zero beyond the budget already allocated to the channel. The return is the compounding improvement that accumulates across rounds. No single round looks impressive. The cumulative result is the business case.
What the field-reading capability means when applied to a real funnel
The second striking result from the self-play experiment was that the model developed predictive judgment: the ability to evaluate a strategy's likely performance without running it. This predictive capability appeared after sufficient rounds of iteration, when the model had seen enough outcomes to begin recognizing patterns.
In a business context, this is the value of a team or an advisor who has run enough iterations in a specific channel to know, before testing, which approaches are likely to work and which are likely to fail. This expertise is real and valuable, and it develops from exactly the process the experiment described: rapid, feedback-rich iteration with honest measurement.
The practical implication is that AI can begin to simulate this field-reading capability by scanning a funnel for patterns that predict performance. What is the typical cost per lead for this type of offer in this market? Which ad formats in this category historically produce the highest reply rate? Which follow-up timing patterns have the highest conversion rate for this buyer profile? The model is not predicting with certainty. It is narrowing the search space before the iteration begins, which means the early rounds of testing start from a better informed starting point than random experimentation would produce.
Combining this predictive narrowing with disciplined iteration is the full Red Queen strategy for a business. Use available data and expertise to identify the most promising starting hypotheses. Run them in parallel against real traffic. Kill the losers, keep the winner, immediately introduce new challengers. Repeat across enough rounds to capture the compounding return. The result is tactics calibrated to the actual market, arrived at faster than copying ever could have produced them, and more defensible than borrowed strategies that any competitor can copy the same way you did.
The lesson of the self-play experiment is ultimately a permission structure. You do not need to study your competitors to find what works. You can discover it faster through your own iteration, and what you discover will fit your specific situation in ways that borrowed strategies never fully can. The arms race is real, the treadmill keeps moving, and the only durable answer is to keep running.
The measurement problem that most businesses have not solved before they try to apply the loop
The self-play learning loop that AlphaGo demonstrated requires one prerequisite that most businesses skip: a clear, measurable definition of what winning looks like at each stage of the process. In the game context, winning is unambiguous. In a business context, winning is rarely defined with the same precision.
A marketing team that applies the self-improvement loop to its ad creative process without a precise definition of what constitutes a better creative will iterate in circles. Better by what measure? Higher click-through rate on the first interaction, or higher conversion rate on the third? Lower cost per click, or higher lifetime value per customer acquired? Each answer implies a different iteration direction, and without specifying the answer before beginning, the team will optimize for whichever metric happens to be most visible in the week they are reviewing.
Define the metric before you start. One metric per iteration loop, measured over a sufficient time horizon to be reliable, evaluated against a baseline that was established before the first iteration began. This sounds obvious but is violated in almost every business application of iterative improvement because the temptation to add a second or third metric when the first one moves unexpectedly is nearly universal. A multi-metric optimization without a clear priority ranking produces teams that disagree about whether performance improved.
The businesses that apply self-play principles successfully define the single outcome they are optimizing for, measure it consistently, and resist the temptation to add metrics until the first one has been reliably improved.
The loop does not stop at the point where the strategy is working. It continues into the next phase of the market, the next platform update, the next shift in what the audience responds to. The businesses that internalize the loop as a permanent operating discipline rather than a project to complete are the ones that compound advantage over time instead of defending a position that was won once.
That is exactly what we do at AI DOERS. Book a private 30-minute call with Madhuranjan Kumar and we will map the fastest path to it for your specific business.
Book your call →
