Building an Evolving Strategy with Checkpoints and an LLM Coach
Letting an AI edit a running strategy sounds great until you ask one question: how do you know the new version is better?
In August we showed how an LLM Reaction can supervise a Code Strategy and patch it when something breaks. This post takes the same idea further. The AI may now replace the whole trading logic, and a small set of rules keeps that from turning into guesswork.
We call the setup Evo. It trades BTCUSDT on paper, and it is built only from parts you already have on Sigrex: one Code Strategy, two webhooks and one LLM Reaction. We built and tested it through the Sigrex MCP server.
The loop

The strategy reports to a Data Webhook, the webhook wakes the coach, and the coach sends back the next version of the trading logic. Signals go to a Bot Webhook with no bot attached, so they are only logged.
One file, two owners
The trick that makes rewriting safe is a split inside the strategy code. Comment markers divide it into a part the AI owns and a part nobody touches.
Block | Owner | What it holds |
|---|---|---|
| The coach | Settings and a |
| Nobody | Indicator helpers (EMA, RSI, ATR and friends) |
| Nobody | Price, stops, the paper trade journal, checkpoints |
The coach can swap a trend strategy for a mean-reversion one. It cannot change how trades are booked or how results are reported. So every version, good or bad, is measured by the same ruler.
Here is the whole seed version. It trades pullbacks in the direction of the EMA 50/200 trend on 15-minute candles.
//@@GENOME-BEGIN
const GENOME = {
v: 1,
name: "ema-trend-rsi-pullback",
tf: "15m",
bars: 300,
p: { fast: 50, slow: 200, rsiLen: 14, rsiLo: 40, rsiHi: 60, atrLen: 14, slAtr: 2, tpAtr: 3, maxBars: 48 }
};
function decide(m) {
const p = GENOME.p;
const n = m.c.length - 1;
const fast = m.ind.ema(m.c, p.fast);
const slow = m.ind.ema(m.c, p.slow);
const rsi = m.ind.rsi(m.c, p.rsiLen);
const atr = m.ind.atr(m.h, m.l, m.c, p.atrLen);
if (slow[n] === null || rsi[n - 1] === null || atr[n] === null) return { act: "HOLD", why: "warmup" };
const up = fast[n] > slow[n];
const dn = fast[n] < slow[n];
if (m.pos) {
if (m.pos.bars >= p.maxBars) return { act: "EXIT", why: "time" };
if (m.pos.side === "LONG" && dn) return { act: "EXIT", why: "trend-flip" };
if (m.pos.side === "SHORT" && up) return { act: "EXIT", why: "trend-flip" };
return { act: "HOLD" };
}
const px = m.c[n];
if (up && rsi[n - 1] < p.rsiLo && rsi[n] >= p.rsiLo) {
return { act: "LONG", sl: px - p.slAtr * atr[n], tp: px + p.tpAtr * atr[n], why: "pullback-up" };
}
if (dn && rsi[n - 1] > p.rsiHi && rsi[n] <= p.rsiHi) {
return { act: "SHORT", sl: px + p.slAtr * atr[n], tp: px - p.tpAtr * atr[n], why: "pullback-dn" };
}
return { act: "HOLD" };
}
//@@GENOME-ENDdecide gets finished candles (m.c, m.h, m.l), the open paper position (m.pos) and the helpers (m.ind). It answers with LONG, SHORT, EXIT or HOLD, plus a stop and a target.
The harness does the rest. It uses a time gate to tick every 3 seconds, checks stop and target on the live price, and calls decide once per finished candle. Every trade is booked on paper with a 0.12% round-trip cost and written to $.Storage.
Signals still go out for real, to a Bot Webhook with no bot attached. That is the forward-testing setup: live prices, real signals, nothing executed.
Checkpoints decide when the coach wakes up
An LLM that reviews every candle is expensive and jumpy. Evo reports only when there is something to review.
Reason | When it fires |
|---|---|
| 6 new closed trades |
| At least 2 new trades and -3% or worse since the last checkpoint |
| 20 errors in a row, or |
| 12 hours passed |
Checkpoints are at least 45 minutes apart. Trimmed down, the logic looks like this:
let reason = null;
if (minutesSince >= 45) {
if (newTrades >= 6) reason = "trades";
else if (newTrades >= 2 && resultSince <= -3) reason = "drawdown";
else if (errorsInARow >= 20 || decideFailsInARow >= 3) reason = "errors";
else if (minutesSince >= 12 * 60) reason = "time";
}
if (reason) {
await $.Http.post(HOOK_URL, report, { "X-Key": $.Env.EVO_KEY });
}The report is one JSON object: the live version, results since the last checkpoint and per version, the last 12 trades, the open position, a market summary and recent errors. It goes to a key-protected Data Webhook with a plain $.Http.post.
The coach never has to dig for numbers. Everything it needs to judge the strategy arrives in the payload.
The coach
The coach is an LLM Reaction on that Data Webhook. Signal sending is off, so it gets no trading tools. It can read and rewrite the strategy, run JavaScript, and keep its own storage. That is all it needs.
At each checkpoint it picks one of five answers.
Decision | Meaning |
|---|---|
KEEP | Change nothing |
TUNE | Small change to the current logic |
REPLACE | Different logic |
ROLLBACK | Bring back the best proven version |
FIX | The current version throws errors, repair it |
Patience is a rule, not a hope
Most of the prompt is about when not to act. One line from it:
A strategy that changes at every checkpoint teaches you nothing, because no version lives long enough to be measured, so patience is part of the job.
In practice that becomes three rules:
- Trial. A new version is on trial until it has 12 closed trades or 4 days live. On trial the answer is KEEP, unless it throws errors or loses more than 4%.
- Champion. A version that ends its trial positive, with a profit factor of at least 1.1, becomes the champion. Its code is saved in the coach's memory.
- One experiment at a time. Each change is a new trial with a written hypothesis. If it fails, the coach rolls back to the champion.
A lab for candidates
Before anything is deployed, the coach tests it. It calls run_javascript_code with the candidate GENOME, the LIB block and a small replay script from its prompt. The script loads about 2,000 recent candles and runs them through the same rules as the harness: same decide input, same cost, same stop cap.
This works here because the seed only looks at candles. Logic that reads a live API or asks an LLM can't be replayed like that, which is why live paper results always outrank the lab.
A candidate has to beat the live version on the newest 40% of candles, with at least 8 trades, and stay sane on the older 60%. The coach may test two candidates per checkpoint. More tries only raise the odds that the winner is luck.
Deploy, check, undo
A deploy is set_code_strategy with the full code: the new GENOME, then LIB and HARNESS copied unchanged. Right after, the coach reads the strategy back with get_code_strategy and compares the error count from get_error_logs before and after. If anything looks off, it restores the previous code.
The harness notices the new version number on its next tick. It closes any open paper position, tags the exit genome-swap, and starts fresh statistics for the new version.
Memory
The coach saves one JSON object with set_storage after every checkpoint, including the ones where it does nothing. It holds the champion, the current trial and its hypothesis, a history of every version, a short list of lessons, and a log of decisions.
That memory is what makes the loop consistent. Without it, every checkpoint is a first date.
All the tools are in the LLM tooling reference.
What the first review looked like
We sent the first checkpoint by hand, half an hour after the strategy started. Zero trades, nothing to judge yet. The coach did what the prompt asks on a first run: it tested the seed in the lab to get a baseline.
Lab segment (15m, 1,999 candles, about 21 days) | Trades | Result | Profit factor |
|---|---|---|---|
Older 60% | 23 | -3.32% | 0.57 |
Newest 40% | 10 | +2.94% | 4.46 |
All | 33 | -0.38% | 0.96 |
A mixed picture: bad on the older candles, good on the recent ones, flat overall. The coach decided KEEP and wrote this into its memory:
Do not replace the seed yet: its newest lab segment is positive, so live evidence must decide.
It also parked two ideas for later, a stricter trend filter and a faster timeframe. It did not try either. That restraint is the behaviour we wanted.
We then ran a deploy drill on a copy of the strategy. The model rewrote the full file with a new version number, and the harness picked it up on the next tick with no errors.
That first review ran on GPT-5.5. The drill passed on a much cheaper model, step-3.7-flash through Nous, so at least the deploy step does not need the most expensive model you have.
Things to keep in mind
- It is a paper setup. Results are booked at the live price with a flat cost. Real fills and slippage will differ.
- It learns slowly. A 15-minute strategy makes one or two trades a day, so a trial takes days. That is the price of not chasing noise.
- It can learn the wrong thing. The trial and lab rules slow that down. They do not remove it, and nothing here is a promise of profit.
- Each checkpoint is one LLM run on your key. With the 12-hour heartbeat that is at least two a day.
- Watch the first deploys. The coach rewrites the whole file, so read the first few versions it ships before you trust it.
- Going live is one step. Attach a Signal Bot to the paper webhook and the same signals become orders. Do that only when the journal has earned it.
Build your own
- Create a Bot Webhook and attach nothing to it.
- Create a key-protected Data Webhook.
- Write the Code Strategy. Keep the trading logic in one marked block. Let the rest journal to
$.Storageand post checkpoints. - Create an LLM Reaction on the Data Webhook with signal sending off. Put the rules in the prompt: what it may change, when it must wait, how to test, how to undo.
- Send the first checkpoint yourself. Read what the coach writes to its storage before you let it run alone.
Or skip the clicking. Connect the MCP server to your AI assistant and describe the loop you want.


