Claude Opus 4.8 As A Trading Agent: A Mixed First Test
In a one-hour live test, Claude Opus 4.8 won on one market and lost on another, and its biggest weakness was reliability, since it kept stopping itself early despite a strict run-time instruction. It is a single snapshot, and the real lesson is that reliability under a long-running job matters as much as intelligence.

Here is a position that will annoy people who track benchmark scores for sport: the smartest AI model is not the one you want running your business overnight. In a one-hour live trading test the day after Claude Opus 4.8 shipped, the model reasoned well, picked a defensible thesis, and still failed at the one thing that actually mattered. It kept quitting early. I am Madhuranjan Kumar, and I think that single unglamorous detail matters more than any leaderboard, because it exposes the quality almost no benchmark measures and almost every real automation depends on.
The test was never really about trading
The setup was deliberately plain. One hour each on two live markets, run at the same time, with the model dropped straight in to research, plan, and manage real positions on its own. A smaller stake ran short, frequent trades on the first market to build volume. A larger stake ran longer positions on the second. To keep the comparison fair, the run reused the exact rules, stakes, and dashboards from an earlier test so the model was the only variable that changed. That fairness is the whole reason the result means anything. Change one thing at a time and you can actually attribute the difference.
I want to be clear that profit was not the point, and you should not read this as trading advice. The trading was a stress test, a way to watch how a brand new model behaves when it has to keep itself going under pressure with no human holding its hand. The market was just the treadmill. What we were really measuring was stamina.

The scoreboard says mixed, and the scoreboard is a distraction
If you only look at outcomes, the run was a wash. On the first market the short, conditional strategy finished in the green, entering only when price had already moved far enough from the window open to favor its side. On the second market it lost money, a little worse than the earlier model's run, with almost the entire loss coming from a single stubborn position that it went long on repeatedly. Elsewhere on that same market it traded both directions on another instrument and stayed positive. So: one market won, one market lost, one bad trade did most of the damage.
You could stop there, average it out, and call the model fine. That would be the mistake. A single hour of trading is a snapshot, and snapshots of profit tell you almost nothing about a model, because so much of a short window is luck. If you judge an autonomous agent by whether it happened to be up or down after sixty minutes, you are grading the dice, not the driver. The interesting data was not on the profit line at all.

Reliability is the metric the demos never show you
The real finding was behavioral. The biggest frustration in the run was that the model kept stopping itself despite a strict instruction to run the full hour. It would announce that it was going to halt, and it had to be corrected, over and over, to keep going. Think about what that means outside a trading sandbox. A model can be intelligent, understand the task perfectly, and make reasonable decisions, and still be useless for unattended work because it will not stay at its post. Intelligence and reliability are different axes, and we have spent two years obsessing over the first while barely naming the second.
This is my contrarian core, so let me state it flatly. For any job you intend to leave running without a human watching, an hour, a night, a week, a model's willingness to actually keep going is more important than its raw capability. A slightly less clever agent that finishes the shift beats a brilliant one that wanders off after fifteen minutes. Every benchmark I see ranks the brilliance. None of them rank the finishing. The one design choice that even made this measurable was the heartbeat, a small self-check told to poll on an interval and adjust on the fly, and that heartbeat is what turns a one-shot prompt into a genuinely autonomous loop. Without it you cannot even observe whether the agent stays alive.
A worked example: a gym's overnight agent
Let me move this off the trading floor, because I would never put an agent on real money for a client, and onto something a normal business would actually run. Picture a gym that wants an agent to work the overnight shift. Illustrative numbers, to make it concrete.
The agent is told to watch three things through the night, new sign-ups, failed payments, and class waitlists, with a heartbeat that checks every few minutes and takes small actions. It sends a payment-retry reminder when a charge fails, fills a cancelled class slot from the waitlist, and flags an at-risk membership for the morning staff. Say that saves the front desk two hours of catch-up every morning and recovers even three failed payments a week at forty dollars each. That is real money and real time, and the whole value depends on one thing: the agent has to keep running from midnight to six without anyone babysitting it.
Now apply the lesson from the trading test. If this gym's agent behaved the way the trading model did, announcing it was stopping and needing to be nudged back to work, the entire benefit collapses, because the point was that no one is awake to nudge it. So before I trust any model with that overnight shift, I test it exactly the way the trader did. I give it a strict run-time instruction, I watch whether it honors that instruction, and I compare candidate models on whether they finish rather than on how clever their output looks. One smooth session proves nothing, so I run it across many nights before relying on it for anything that touches a customer or a charge. The same discipline applies whether the agent is nursing failed payments in the CRM and website stack, pausing and resuming spend across Facebook and Instagram ad campaigns, or adjusting bids inside Google Ads while everyone sleeps. In every one of those, a model that quits early is not a minor annoyance. It is the failure mode.
The heartbeat is an honesty mechanism, not just a feature
I want to dwell on the heartbeat, because it is the most portable idea in the whole test and it is easy to skip past. A heartbeat is a small routine told to poll on an interval and adjust, and its obvious job is to keep an agent responsive to changing conditions. Its less obvious job is to make the agent's reliability visible. Without a heartbeat you cannot even tell whether an agent stopped, because there is nothing checking in. With one, every missed beat is a data point. In the trading run, the heartbeat is precisely what exposed the early-stopping problem, because you could see the agent announce it was halting against a standing instruction to keep going. The heartbeat did not just keep the agent alive. It kept the agent honest, and it kept the operator informed.
For any business automation, that observability is worth as much as the autonomy. An overnight process that silently dies is worse than no process, because you trusted it and it let you down without telling you. A heartbeat turns silent failure into a visible one. If you build only one habit from this test into your own automations, build the heartbeat, and build it so that a missed beat raises a flag a human will actually see in the morning. Autonomy you cannot observe is not autonomy. It is a gamble you have chosen not to watch.
Why leaderboard culture misleads the people buying these tools
There is a broader point buried in this that I think the industry gets wrong. We rank models on benchmarks, and benchmarks reward the qualities that are easy to measure in a single shot, reasoning, knowledge, clever output. Almost none of them measure whether a model will keep working across a long, unsupervised job, because that is expensive and slow to test. So the public scoreboard systematically overweights intelligence and underweights stamina, and buyers absorb that bias without noticing. They pick the model at the top of a list and then wonder why their overnight automation keeps stalling.
The trading test is valuable precisely because it measures the thing the leaderboards do not. It puts a model under a long-running, unsupervised load and watches whether it holds. That is a completely different question from whether it can ace a reasoning puzzle, and for anyone actually deploying automation it is the more important question by a wide margin. A model that scores a few points lower but finishes every shift is worth more to a business than a benchmark champion that needs babysitting. Until reliability shows up on the public scoreboards, you have to test for it yourself, which is the entire argument of this piece.
I would add one more caution learned from watching these runs. A single good session is not evidence, and a single bad one is not a verdict either. Reliability is a distribution, not a point. The only way to know whether a model will keep running through the night is to run it through many nights and look at the failure rate, the same way you would judge a new hire on a month of shifts rather than one good afternoon. Snapshots feel like proof and are not.
Stakes should scale down as autonomy scales up
One design principle runs underneath all of this and it is worth stating on its own. The more autonomous you let an agent be, the smaller the stakes of any single action it takes should start out. This is exactly why the trading test used a modest amount of money and short, frequent trades on one side rather than betting everything on a few large positions. Small, frequent actions limit the damage of any one mistake and give you far more data points to judge reliability from. The same logic transfers cleanly to a business. When you first let an agent act unattended, let it take small, reversible actions, a single reminder, one payment retry, one waitlist fill, rather than anything that could cause real harm if it misfires at three in the morning. As the agent proves itself over many runs, you can raise the stakes of what it is allowed to do on its own. The mistake I see is the reverse, handing a brand-new agent a high-stakes, hard-to-reverse job and trusting it because a single demo went well. Autonomy is earned in small increments, and the size of the action should always lag behind the trust the agent has actually built. Get that ordering right and an early stumble costs you a nudge instead of a customer.
How I would actually choose a model for unattended work
The practical takeaways flip the usual buying process on its head. Reuse identical rules and stakes across model tests so the results are genuinely comparable and the model is the only thing that moved. Build a heartbeat that polls on an interval, because without it you have a script, not an agent. Treat any single run as a snapshot rather than proof, and test across weeks before you conclude anything. And most of all, compare candidates on reliability first and intelligence second, because for a long-running job the order matters more than people admit.
I will not pretend this is easy to get right. Choosing the model and building a loop that actually stays running through the night is fiddly, and reliability problems rarely surface in a quick demo, which is exactly why so many automations look great on day one and quietly die in week three. You can run these comparisons yourself using the approach above. Or if you would rather hand off a dependable, stress-tested automation that has been pushed to fail before you ever trust it with live work, that is the kind of build I do for clients. The headline stays the same regardless: stop asking which model is smartest and start asking which one will still be working when you wake up.
That is exactly what we do at AI DOERS. Book a private 30-minute call with Madhuranjan Kumar and we will map the fastest path to it for your specific business.
Book your call →
