⌂ Research Index · tradebotsecrets research record · evidence-first, negatives included
Primitives Registry — BTC Trading System

BTC · Unsupervised discovery · Live registry

Primitives Registry

Every signal this project has discovered, in one place, so nothing gets lost between sessions. Each entry is a real, holdout-tested finding — worked, didn't work, or not tested yet — never a hypothesis dressed up as a result.

35
Primitive families
8
Validated signal
17
Negative result
10
Untested / marginal
0
Planned / in progress

Candle-Shape Entry Primitives

evolution/candle_shape_entry_primitive_v1.py · length=1

Each discovered shape cluster gets a fitted entry direction (long / short), computed directly from real historical forward returns — no GA, no exit logic. A regime switch to a different primitive is what closes or flips a position. Fee-aware, and now requiring SIGNIFICANT profit: a cluster's real directional lean is forced to "no entry" unless its edge clears at least 2× a real round-trip fee (2×5bps taker), not bare breakeven — the operator's own framing for why a choppy/sideways edge shouldn't trade just because it's technically positive. At horizon=1 this wipes out every sub-Daily timeframe except a single surviving cluster each at 15m and 12h, cuts Daily from 7 to 5 surviving clusters, while Weekly/Monthly (much larger real edges relative to fee cost) mostly still clear it. Leverage-scaled, but as a CEILING, not the entry leverage: a surviving cluster's max_conviction_leverage is fitted from the same real edge, saturating (tanh) toward the real exchange max — but the composed policy always starts a fresh entry at the floor and only ratchets up toward this ceiling on confirmed favorable movement (see the Wave-Attention Selector section below for how this actually performed on holdout — still a real loss).

Timeframe Clusters Long / short split Holdout samples Avg. holdout hit rate Leverage range Status

Order-Book Microstructure Signals

scripts/run_orderbook_signal_research_v1.py · read-only from evolive2

Magnitude-only signals — they predict that a big move is coming in the next 1–15 minutes, not which direction. That makes them position-sizing primitives (bet small to protect the downside, or bet big if direction is already known from elsewhere), not entry primitives. ~19M real ticks, chronologically split — train May 23–Jul 14 2026, holdout Jul 14–25 2026.

Order-Book Volatility Sizing Primitive Validated, modest

evolution/orderbook_volatility_sizing_primitive_v1.py · scripts/run_orderbook_volatility_sizing_check_v1.py

Turns the order-book magnitude signal above into the position-sizing primitive it was always framed as: never picks direction, only ever scales an entry primitive's size_fraction DOWN (0.5× in this pass) when tick_count or mid_relative_volatility flags a likely large move — "bet small when it fires," the operator's own framing from the original research round.

First attempt was a real dead end, reported plainly. Composed directly with the fitted 1m entry primitive (length=1, horizon=1) to get real directional trades to test against — found zero: all 17 clusters are suppressed by the entry-side fee-churn gate at that granularity (a 1-minute BTC move essentially never clears real round-trip taker fees). Refitting at horizon=15 finds exactly one surviving cluster, itself fit on a single training occurrence with zero real occurrences anywhere in the holdout pool — not a usable trade source either. Consistent with this project's standing conclusion that direction has no learnable structure at short candle scales.

Pivoted to test the primitive's actual claim directly: does scaling down size on flagged real minutes shrink real tail exposure? Reused the order-book research's own real per-minute joined table (same 15,678-minute holdout, same chronological split) and measured unsigned |forward_return| at 1 minute under baseline (always full size) vs. the sizing primitive applied.

ConditionMean |return|StdTop-decile mean
Baseline (full size always)0.02810%0.03061%0.09709%
Sizing primitive applied (40.0% flagged)0.02145%0.02403%0.07525%
Random control — same 40.0% fraction, same 0.5×, 200 real trials0.02248%0.02648%0.08307%

The random control is the load-bearing check. Most of the raw reduction (baseline → sizing primitive, ~22–24%) is mechanical — cutting size on any 40% of minutes by half shrinks the aggregate roughly that much regardless of which 40%. The real question is whether the order-book flag targets the ACTUAL large-move minutes better than a random 40% would. It does, on all three metrics, most clearly on the tail: sizing-primitive top-decile mean (0.07525%) vs. random-control top-decile mean (0.08307%) — a real ~9.4% tail reduction beyond what blind size-cutting achieves, 13.4 standard deviations below the random control's own trial-to-trial spread (200 real random draws, not one seed) — not explainable by chance. Modest in absolute size, genuinely real.

Volatility Straddle Order Primitive Negative, all 10 timeframes

evolution/volatility_straddle_order_primitive_v1.py · scripts/run_volatility_straddle_backtest_v1.py

Operator's design: use the magnitude signal (above) at every timeframe 1m–D to decide when a move is worth a real round-trip fee, then place a buy-stop AND a sell-stop around the current price — whichever the market actually takes becomes the position. Wrong-side/fakeout risk is handled the same way this project's own DCA-up entry primitive already does (start small, grow only on CONFIRMED continuation) rather than by waiting for confirmation before entering — this project already tested that directly (Trend Confirmation, above) and found it a null with a real lag cost. The halving-cycle theory (this project's one validated DIRECTIONAL signal) breaks the straddle's symmetry: whichever side agrees with the current halving-cycle phase ratchets toward full size on confirmation; the other side stays capped at its starting size for as long as it's held.

TimeframeTraining returnBuy & holdTradesFee dragHoldout
1m-30.44%-0.86%1414.45%0.00% (bh -3.38%, n=0)
5m-30.94%-4.24%843.22%
15m-27.54%-4.13%822.79%
30m-20.18%-4.10%632.18%
1h-21.12%-4.22%331.27%-0.18% (bh 0.18%, n=3)
2h-26.75%-4.19%551.68%
4h-22.29%-4.23%381.09%
6h-23.77%-4.18%371.03%
12h-28.42%-4.33%420.99%
Daily-33.62%20.54%120.43%+16.57% (bh -21.53%, n=2)

A clean, decisive negative at every timeframe on training data — fee drag (0.4%-4.4%) is small and NOT the driver; the loss comes from real, repeated stop-outs. The straddle keeps buying real breakouts and selling real breakdowns that then genuinely reverse — the same underlying phenomenon the Trend-Following Controller's Donchian-breakout test already found (also negative, above), confirmed again through a structurally different mechanism. A real bug was caught and fixed before trusting any of this (a fill candle's own close was immediately re-triggering its own stop-loss); the fix barely moved the numbers, confirming this is a real result, not an artifact.

The halving-cycle bias did not help, and the honest per-trade attribution suggests it hurt. Comparing trend-aligned vs. counter-trend trades' own raw returns at 1m/1h/D: trend-aligned trades were NOT less negative — at Daily they were dramatically worse (-12.03% mean, 0/3 hit rate, vs. counter-trend's -0.89% mean). A macro, multi-month signal doesn't predict whether any ONE local breakout holds; letting the "aligned" side ratchet toward full size just means losing bigger on the frequent misses, while the capped counter-trend side's losses stay bounded by construction. The one holdout number that looks positive (Daily, +16.57%) is on only 2 trades while the same mechanism lost -33.62% on Daily's own training pool in the same run — not trustworthy evidence, stated plainly rather than cherry-picked.

Wave-Attention Selector

scripts/run_trading_selector_train_v1.py · scripts/run_entry_primitive_selector_train_v1.py

Same router mechanism throughout (nonlinear phase-coherence attention + multi-timeframe context), only the primitives underneath change. Daily holdout, genuinely disjoint from training in every run.

Discovery trainer (100x leverage cap) Real discovery, liquidated

evolution/discovery_trading_trainer_v1.py · real fees/leverage, no hand-coded exit rule
Holdout return ratio0.621
Buy & hold (same window)0.785
Holdout trades1
Liquidatedyes

Discovery trainer (10x leverage cap) Safe, never trades

same seed/architecture, TRADING_MAX_LEVERAGE 100x→10x
Holdout return ratio1.000
Buy & hold (same window)0.785
Holdout trades0
Liquidatedno

Discovery trainer v2 (3x safety margin) Improved, still liquidated

evolution/discovery_trading_trainer_v2.py · leverage capped by real liquidation buffer + zero-trade penalty
Holdout return ratio0.624
Buy & hold (same window)0.785
Holdout trades3
Liquidatedyes

Discovery trainer v2 (8x safety margin) Mixed — demoted by forward validation

same architecture, wider liquidation buffer margin
Holdout return ratio1.096
Buy & hold (same window)0.785
Holdout trades5
Liquidated (holdout)no
Forward continuation (fresh 2026-08 candles)LIQUIDATED
With 9.4x tail cap (same genome)survives, holdout 1.160

Forward-validation demotion (2026-08-18), then the fix. Replaying this champion on real candles newer than every training decision exposed what the "validated" mark-to-market number was hiding: the holdout run ended carrying a live 73.4x short opened on the second-to-last holdout candle (the magnitude-based safety cap read a momentarily quiet signal), and the very first out-of-sample candle (+1.05%, 2026-08-03) liquidated it — the fourth independent hit of the same leverage-vs-liquidation failure mode. The fix: a hard tail cap derived from the real Daily training pool's own 99th-percentile adverse excursion (10.2% → ~9.4x ceiling), applied as a min() over the adaptive magnitude cap. Same genome, bounded hands: no liquidation on the same continuation, fresh-segment loss reduced (0.956 vs 0.923), and the holdout itself improved to 1.160. The v3 maker/taker champion is byte-identical under the cap (it never left 1x) — the ceiling binds only where risk was excessive.

Operator's own design: "let it discover it and train to come up with its own strategies... don't want to give it the answer [trailing stop]." Reused the exact NEAT/GA substrate that produced the 5-primitive result below, gave it richer real information (the volatility magnitude signal, unrealized PnL, halving-cycle phase) and a genuine close-to-flat action — no exit rule hand-coded anywhere. Traced the 100x-leverage champion's real per-candle holdout behavior before trusting anything: it held flat for 77 of the first 80 real candles (genuine, unprompted patience), entered long once, and one candle later — with distance-to-liquidation down to 2.16% — issued its own close order, unprompted. A real, discovered defensive reflex. It fired one candle too late: at ~73x leverage the liquidation buffer (~1.4%) is smaller than a typical real Daily candle's own range. Capping leverage to 10x removed the liquidation risk entirely but the search fell into the fitness function's own already-documented "never trade" trivial optimum instead — zero trades on holdout either way.

Operator: "close it." Built a real fix, not an artificial cap: leverage now squashed into [min_leverage, max_safe_leverage] where the ceiling is derived from the SAME expected-move magnitude signal already computed each step, so a wilder market genuinely earns a tighter leverage ceiling — the same logic a real exchange risk engine applies. Plus a fitness penalty specifically for never trading at all, breaking its tie with careful trading. First re-run (3x margin) found a real bug in the fix itself: before two real candles exist, the magnitude signal is exactly 0.0, and the fallback for "no signal" was "assume it's safe" — letting a genome open at 73x leverage on literally the second candle of the episode. Fixed (fallback to the safe floor instead), re-ran: leverage was now genuinely bounded by real recent volatility (~39x/15x vs. the uncapped run's ~73-100x) and drawdown dropped (0.376 vs. 0.592) — still liquidated, but for an honest reason this time: a real move exceeded what a 3x margin on the PRIOR candle's volatility anticipated, the same fat-tail risk any volatility-sizing scheme faces.

Widened the margin to 8x: a genuine validated win. Holdout return ratio 1.096 (real +9.6%), never liquidated, beats buy-and-hold by a wide margin on the same real bear-market window. Traced the real behavior: opened a small short at the safe floor, built it up as price fell ~81,700→63,700 (matching buy-and-hold's own -21.5% on this window) with equity climbing through the decline, closed the position on its own near the bottom, then re-shorted near the end — coherent entry, position management, and exit, none of it hand-coded. Honest caveat: n=1 seed on one holdout window where being short was simply correct — the two diagnosed mechanics are confirmed real and fixable, not that this exact genome is proven durable; a walk-forward check (matching the rigor already applied to the DCA ladder work) is the natural next step, not yet run.

5 NEAT trading primitives Validated

primitives/daily_D_growing_fib/ · zero-fee "brave" profile
Holdout return ratio1.0630
Best single primitive alone1.0555
Holdout trades5
Holdout churn rate5.6%

11 entry-only primitives Whipsaw

no fee-churn gate · zero-fee "brave" profile
Holdout return ratio0.8797
Training fitness (in-sample)1.0995
Holdout trades33
Holdout churn rate36.7%

7 fee-gated entry primitives Improved

candle_shape_entry_primitive_v1_compute_round_trip_fee_cost · REAL fees, fixed 1x leverage
Holdout return ratio1.0425
Training fitness (in-sample)1.0823
Holdout trades38
Holdout fee drag0.63%

+ conviction-scaled leverage Overfit

candle_shape_entry_primitive_v1_compute_leverage_for_edge · REAL fees, 1x–100x fitted leverage, sized ALL AT ENTRY
Holdout return ratio0.4771
Training fitness (in-sample)5.386
Holdout trades1
Holdout fee drag2.29%

+ DCA-up ratchet Still negative

low start, ratchets toward the ceiling only on confirmed favorable moves · 5 clusters, stricter "significant profit" gate
Holdout return ratio0.4924
Training fitness (in-sample)4.935
Holdout trades3
Holdout fee drag1.69%

+ take-profit / trailing-stop exit Improved, still negative

fitted take-profit while unconfirmed, 3-candle trailing stop once armed · see Holdout Timeline
Holdout return ratio0.9306
Training fitness (in-sample)3.346
Holdout trades17
Holdout fee drag0.30%

+ damping-anticipated take-profit Liquidated

target from the real fitted volatility-damping cycle, not a static number · see Holdout Timeline
Holdout return ratio0.4924
Training fitness (in-sample)5.085
Candles processed8 / 90
Liquidatedyes

4 of Daily's 11 clusters had a real historical directional lean too small to survive a real round-trip fee (2×5bps) and were suppressed to "no entry" instead of trading them anyway. That alone flipped the entry-only approach from a real holdout loss (0.8797) to a real holdout gain (1.0425), and improved the surviving clusters' own average holdout hit rate too (62.8% → 65.2%, see the entry-primitives table above).

Sizing leverage off that same edge (bigger historical edge past breakeven → more leverage, up to the real exchange max, committed all at once on entry) made it worse, not better. In-sample training fitness jumped to 5.386 (nearly 5× the fixed-leverage run) purely because leverage amplifies whichever trades the training segments happened to get right — the GA started selecting for lucky training-segment leverage, not genuine timing skill, and that didn't transfer to holdout. Fee drag alone (2.29%, one trade) exceeded the entire 38-trade fixed-leverage run's total, because fee cost scales with leverage exactly as directly as profit does.

The operator's own read on the failure: "we enter low leverage and dca in when in our favor... the only way to use leverage well is to not make mistakes on entering too early or exiting too late." Rebuilt accordingly — every fresh entry now starts at the exchange floor (1x) and only ratchets toward the cluster's own conviction ceiling while price keeps confirming itself favorable since the last add, never on the unconfirmed prediction alone. Entry itself also got stricter: clusters now need a SIGNIFICANT profit margin (edge ≥ 2× round-trip cost, not bare breakeven) to be tradeable at all, which is why Daily dropped from 7 survivors to 5 (3 long, 2 short) for this run. Result: still a real loss (0.4924), and training fitness is still ~4.5× the fixed-leverage baseline — the same overfitting signature persists. This makes sense once you notice the DCA-up mechanism only fixes HALF of what the operator named: entering too early is now cheap (you start small on an unconfirmed move), but nothing in this architecture addresses exiting too late — there is still no take-profit, no reversal detection, and no mechanism that closes a position except a different primitive's opposite-direction fill (a deliberate, still-standing "no independent safety net" call, see AGENTS.md). Once leverage has ratcheted up deep into a confirmed run, a reversal before the router happens to select the opposite primitive still lands at whatever leverage was reached, fully exposed. The operator's own diagnosis holds up on both real attempts: leverage genuinely requires solving the exit-timing half too, not just the entry-timing half — that piece isn't built, and isn't started without being asked, same discipline as every other deferred item in this file.

Built the exit half the operator specifically asked for: "with sideways, it makes sense to anticipate where take profit... in bull run or bear, its better to take profit with trailing stop... a 2 to 3 candle trail gives it room to retrace and continue the run." A position that hasn't proven itself extending yet (no confirmed DCA add) now targets a fitted take-profit at its cluster's own historical mean_forward_return; once it HAS proven itself extending (≥1 confirmed add), it switches to a 3-candle trailing stop instead, so a real trend isn't capped by a fixed target. Inspected evolive2's own v2/gates/tp_trail_v2.py (a sophisticated arm/trail/rate-of-change-adjusted engine) for design context — this is a deliberately simpler first pass, not that whole feature set.

Result: still a real loss (0.9306), but a big real improvement — closes most of the gap to breakeven that both leverage attempts opened (0.4771, 0.4924), and training fitness dropped back toward the fixed-leverage baseline (3.346, down from 4.935-5.386), meaning the overfitting signature eased too. Fee drag fell to 0.30%, the lowest of any leverage variant. An important honest nuance, visible on the Holdout Timeline dashboard: leverage never once exceeded 1x during this specific 90-candle holdout window — every one of the 6 exits fired via the fitted take-profit before any position accumulated a single confirmed DCA add. The trailing-stop branch is built and would activate for any position that DOES ratchet, but it was never exercised on this window, so today's improvement is entirely the take-profit/quick-exit side of the design proving itself, not the "let a real trend run" side — that half remains genuinely untested, not validated, pending a holdout period with an actual sustained trend in it.

Operator's own follow-up fixed exactly that gap: "we should be able to see the dampen cycle as it goes sideways and anticipate an optimal take profit... it doesn't always reach the last high of the last cycle, as dampening reduces" — plus "if we anticipate a breakout we should not take profit early and rely on trailing stop." The static mean_forward_return target is replaced with one anticipated from evolution/multiscale_oscillator_v1.py's own already-fitted Daily volatility damping rate (~0.0206/candle): the last comparable swing's real range, decayed by that rate for however long the trade has run, buffered to 85% so it can actually fill. The SAME projection also detects breakouts: if the realized move already exceeds what damping alone predicts, the position is treated as armed immediately (trailing stop takes over) rather than waiting for the slower DCA-confirmation throttle. Verified by direct computation before trusting the run: for the first real holdout trade, the new target sat at $82,514 vs. the old static target's $81,326 — genuinely wider and slower to hit, exactly as intended, so positions get a real chance to develop into confirmed trends instead of getting capped early.

Result: holdout return ratio 0.4924 — but the real story isn't take-profit timing at all. The Holdout Timeline now shows only 8 of 90 candles processed, because the episode ends the instant a position is liquidated for insufficient margin — leverage ratcheted deep enough (this cluster's ceiling: 91.7x) that one ordinary Daily candle wiped the position outright, well before the take-profit or trailing-stop logic ever mattered. The damping-based take-profit itself is real and verified working as designed; it just never got the chance to matter here, because a much bigger, more fundamental gap sits upstream of it — max_conviction_leverage ceilings in the 55x-92x range are simply unsafe for Daily BTC's normal volatility, independent of how well-timed the entry or exit is. This is the same root cause diagnosed all the way back at the very first leverage attempt ("a cluster can have a real, fee-clearing edge on average and still get liquidated... from an adverse swing"), now seen directly in a real trace rather than argued from theory. A hard cap on leverage tied to actual liquidation distance (not just edge size) was the load-bearing next piece — and as of 2026-08-18 it exists: the tail cap in discovery_action_decode_v2_compute_tail_cap_leverage (99th-percentile real adverse excursion → ~9.4x at Daily), validated on the exact candle that liquidated the v2 champion's forward continuation (see the Discovery trainer section above).

Cohort Selector Validated — 8/8 seeds + forward-tested

evolution/cohort_selector_v1.py · scripts/run_cohort_selector_train_v1.py · scripts/run_cohort_selector_multiseed_v1.py

Robustness update (2026-08-18): the n=1 caveat is retired. Forward validation first: the trained router, replayed on real candles newer than every training decision, survived the exact fresh candle that liquidated its own v2 member standalone (cold-start +0.1%, continuation −0.9%, never liquidated) — the composed architecture genuinely absorbs member risk. Then 8 independent GA seeds, retrained from scratch with the tail-capped v2 member (9.36x) and the same guard rail: every seed lands holdout 1.112–1.190 (mean 1.163 ± 0.023), zero liquidations across all 24 replayed windows — the halving-backbone + discovery-for-entries division of labor is a real attractor, not one seed's luck. Honest counterweight: every seed also lost ~4% on the fresh 16-candle forward segment (mean 0.960 ± 0.011) — uniformly bear-positioned in a window where BTC rose ~1.9%; that is the halving-cycle phase call being tested live (registered bottom date 2026-09-26), not a seed-dependent defect. Summary: logs/cohort_selector_multiseed_v1_summary.json.

Operator: "build the combine selector aha get a baseline and we can add more indicators or strategies to it later if its not sufficient." Same router mechanism as the Wave-Attention Selector above, generalized from a fixed catalog of same-shaped NEAT genomes to an open-ended list of heterogeneous primitives — each just a factory returning a fresh policy function, so a rule-based strategy and a NEAT-trained one sit side by side. Only the router genome is trained; every primitive underneath is already fixed and validated.

Baseline cohort, 2 members, both real (real fees, real leverage, genuine holdout, nothing zero-fee or toy): the halving-cycle calendar-only timing rule, and the discovery trainer's own 8x-safety- margin champion from the round above. A third candidate (the candle-road lookback+feedback controller) was considered and deliberately left out — its real-BTC conversion already tested negative standalone, but a router's whole point is deciding WHEN to use something, so it's flagged as the first "add more later" candidate rather than discarded.

ConfigurationHoldout return ratioTradesLiquidated
Halving-cycle alone1.11721 (mostly zero-fill top-ups)no
Discovery-trainer v2 (8x margin) alone1.0965no
Buy & hold0.785
Cohort selector1.5062 real fillsno

A real, decisive win, and genuinely better than either member alone — not just "picked the better one and stuck with it." Traced the real per-candle selections before trusting the number: the router used the discovery trainer for the smart, risk-aware ENTRY (candles 1–71, building a short as price fell — identical to that primitive's own standalone behavior), then at candle 72 — right around where the standalone champion would have gotten nervous and closed — handed off to halving-cycle to hold the already-open position steady at safe 1x leverage through the rest of the real decline, flicking back to the discovery trainer three times to reassess (each time it correctly chose to do nothing). Different primitives specializing in different PHASES of the same trade — an emergent division of labor neither primitive produces alone. Only 2 real fills the entire episode (fee drag 0.52%) — the environment's own "top up to target, don't stack" logic makes every repeated same-direction request after the target is met a free no-op.

Honest caveat: the position was still open at the holdout cutoff, so 1.506 is a real mark-to-market equity figure (standard practice here, same as every other backtest), not a fully realized number — and this is n=1 seed, one holdout window, the same discipline as every other single-run positive in this project. What's genuinely demonstrated is that the cohort MECHANISM itself adds real value beyond its individual members on real data, not just in principle.

Operator: "we already have a trailing stop algorithm developed in evolive folder... it should be the guard rail, acts independently... very conservative giving lots of room, trailing stop is there only to lock profit in if the selector makes a mistake." Confirmed nothing like this existed in this project yet, then ported the REAL algorithm math-for-math from evolive2/v2/gates/tp_trail_v2.py (read-only source, unmodified) into trading/trailing_stop_v1.py — arm threshold, trailing distance, a hard giveback cap, and two real velocity-adaptive levers (a fast favorable move widens the trail; an accelerating drawdown tightens it). Calibrated from real data, deliberately wide: every threshold is a real multiple of the Daily training pool's own mean |return| (172.75bps) — arm ~15%, trail ~8%, giveback cap ~10% — and the velocity reference was rescaled from evolive2's own live-tick-rate default to this project's real Daily candle spacing.

ConfigurationHoldout return ratioTradesGuard-rail firesFinal exposure
Cohort, no guard rail1.5062large, still open (unrealized)
Cohort + guard rail1.13227 (cheap top-ups, 0.31% fee drag)1small, mostly realized

Wired into evolution/cohort_selector_v1.py as a genuinely independent override sitting ABOVE the router: it never influences which primitive is selected, it only has the power to force a real, fully-flattening close at the very end of the pipeline. Retrained the router from scratch with the guard rail present (its own optimal behavior could change once a real backstop exists). Result: a real, legitimately different equilibrium, not simply better or worse. The guard rail fired exactly ONCE (candle 34, a genuine drawdown from a real local peak) then stayed silent the rest of the episode — rare, decisive, non-intrusive, exactly the backstop role it was asked to play. What actually changed the headline number wasn't the guard rail intervening often; training itself found a different, more risk-averse strategy once a backstop existed at all — leaning mostly on halving-cycle's small, cheap, steady exposure (62 of 90 candles) as a backbone, borrowing the discovery trainer for real position growth and its own learned defensive close along the way, ending with realized gains and a much smaller open position instead of one large unrealized mark. Never liquidated.

Next-Primitive Predictor (Cascade)

evolution/candle_primitive_predictor_v1.py

Empirical transition table, cascading 1→2→3→5→8→13→21→34→55→89 candles — given the realized shape at one scale, predict the shape at the next. Tested at both small (Daily, ~700 pairs/stage) and huge (1m, ~1.1M pairs/stage) sample sizes to rule out "not enough data."

Stage 1m accuracy 1m baseline Daily accuracy Daily baseline

Conclusive negative result — accuracy tracks the baseline almost exactly at every stage, at both sample sizes. Bare shape-cluster identity carries no detectable sequential structure at any scale.

Escalated to a trained model, same conclusion — a sharper negative. The transition table above only ever sees the DISCRETE "from" cluster label; it discards the continuous shape vector entirely. Operator's own follow-up: "reassemble the wave attention transformer selector to pick which candle it anticipates... we're not going to enter any trades, we're simply trying to predict candle patterns." Reassembled the SAME wave-attention router this project uses to pick among trading primitives elsewhere (evolution/candle_pattern_attention_predictor_v1.py, scripts/run_candle_pattern_attention_predictor_train_v1.py) — GA-trained, not counted — fed the full continuous shape vector instead of the discretized label, same causal window pairing, same baseline definition, first Fibonacci stage (1→2 candles) at both scales for a fair rematch.

Timeframe Training samples Training accuracy Holdout accuracy Holdout baseline
Daily68718.3%6.8%18.2%
1 Minute1,106,1258.3%4.8%6.4%

Not just null — worse than the naive baseline, at both scales. A genuinely trained, more expressive model given strictly MORE information than the transition table (the full continuous shape, not just its discrete label) still failed to beat "always guess the single most common outcome," and actually scored below it on holdout both times — real overfitting to the per-generation training batch, not a hint of exploitable structure. Third independent confirmation this session (after the transition table and every raw-direction check) that candle-shape identity alone carries no predictable sequential signal, at any scale tried.

Oscillator Decay Projection Validated

evolution/oscillator_extrema_predictor_v1.py · magnitude/timing, not direction or shape

Direct pivot after the candle-pattern predictor's negative result: "following the price movement, not predicting candles with oscillator... check how much move at what time... predict the high and lows of the oscillation." Finds every REAL local peak in a timeframe's |returns| series, then projects the ALREADY-FITTED damping rate (evolution/multiscale_oscillator_v1.py's own real fit) forward across a half-cycle horizon and compares it against what actually happened next — the exact same exponential-decay math the take-profit target already uses, now tested in aggregate across thousands of real peaks instead of verified on a single trade. Compared against a flat/no-decay baseline (same peak value carried forward, no model at all).

Timeframe Real peaks Forecast points Model MAE Baseline MAE Improvement
1 Hour2,561107,5620.003240.0058044.2%
Daily874,9590.018230.0304040.0%

A real, substantial improvement — the first clearly validated predictive signal beyond direction-vs-volatility itself. At 1h, a huge real sample (107,562 forecast points from 2,561 genuine peaks) shows the fitted decay curve cuts prediction error by 44% versus assuming a spike just holds flat. Daily shows the same pattern (40%) on a thinner but real sample. Both scales agree: after a real volatility spike, the fitted damping rate genuinely tracks how far and how fast it comes back down.

Honest scope: this predicts the DECAY PATH given you're already at a known peak — not when the next peak begins. Raw correlation between model predictions and realized values stayed weak in absolute terms (0.14 at 1h, 0.03 at Daily) — most of the win is precision (getting the right ballpark number instead of wildly overshooting), not the model explaining most of the variance. Not a crystal ball; a real, usable improvement over the naive alternative, nothing more. "When does the next peak start" is a genuinely different, not-yet-tested question.

Follow-up, operator's own hypothesis: "everyone is usually all concerned about the peak and no one ever really knows, but is the wavelength or the time duration to the peak more reliable? I notice usually the cycling wavelengths seem more consistent than the actual peak." Tested directly — coefficient of variation (std/mean, dimensionless so candles-vs-magnitude is a fair comparison) of real peak-to-peak TIMING gaps vs. real peak MAGNITUDES, plus a robust median-absolute-deviation version in case a few huge spikes were skewing the plain version.

Timeframe Mean gap (candles) Gap CV Magnitude CV Gap CV (robust) Magnitude CV (robust)
1 Hour7.20.3850.7570.2860.434
Daily7.10.3550.5230.2860.296

Confirmed — timing really is more consistent than magnitude, at both scales, both statistics. Peak-to-peak TIMING varies less than PEAK SIZE does, every way it was measured. The gap is clear-cut at 1h (CV 0.385 vs 0.757 plain, 0.286 vs 0.434 robust); at Daily it's real but much narrower once outlier-robust (0.286 vs 0.296 — nearly tied on the robust measure, only the plain std-based version shows a wide gap there, likely because Daily's thinner sample — 87 peaks vs. 2,561 — is more exposed to a handful of outsized magnitude spikes).

One honest caveat on scale: these real peaks average ~7 candles apart at BOTH timeframes — much more frequent than the fitted macro "period" (84 candles at 1h, 115 at Daily) used for the decay projection above. The simple peak detector (strict local max within ±3 candles) is catching a faster, higher-frequency rhythm than the slow macro cycle the ACF fit describes — a real, honestly-measured finding in its own right, just not literally "the same wavelength" as the period number quoted elsewhere on this page. Whether the SAME timing-more-than-magnitude pattern holds at the macro cycle scale too is a natural next check, not yet run.

Equation-of-Motion Forecast Mixed

evolution/oscillator_equation_of_motion_v1.py · reuses wave_core/complex_oscillator_v1.py's own recurrence, unchanged

Operator's own ask: "can we do some physics with the oscillatory equation of motion, knowing the angular velocity and angular acceleration, its initial point of price, and predict something forward in time? don't forget dampening and decay." Seeds the project's EXISTING damped-rotation recurrence (Z[t+1]=λZ[t], λ=r·e^{iθ}, unchanged from wave_core/complex_oscillator_v1.py) with a REAL instant's own momentum state (Z0 = velocity + i·acceleration, evolution/candle_shape_ clustering_v1.py's own phase convention) and the ALREADY-FITTED damping rate/period (r=e^{-damping}, θ=2π/period). Projects forward, compounds into a price path, tested against thousands of real anchor points — NOT just at real peaks this time, at EVERY sampled point in time — against a flat baseline (nothing moves) and a constant-velocity baseline (naive trend extrapolation, no oscillator physics at all).

Timeframe Anchors Price MAE (model) Price MAE (flat) Price MAE (const. velocity) |Return| MAE (model) |Return| MAE (flat)
1 Hour3,7036.08%1.52%7.33%0.002930.00318
Daily12667.1%9.33%144.8%0.015900.01669

Signed price: negative, again — consistent with everything else this session. The physics-projected price path is WORSE than simply assuming nothing moves, at both timeframes, dramatically so at Daily (67% error vs. 9% for flat). Direction remains unpredictable — a fifth independent confirmation (raw ACF, entry-primitive hit rates, the Markov cascade, the trained attention model, now this). One genuinely useful thing the physics DOES do correctly: it clearly beats a naive undamped constant-velocity extrapolation (6% vs 7% at 1h; 67% vs 145% at Daily) — the damping term is real and functioning exactly as intended, taming what would otherwise be runaway compounding error. It just isn't enough to beat doing nothing.

|Return| magnitude: a real but much weaker edge than the peak-anchored version above. The model beats the flat baseline at both scales (0.00293 vs 0.00318 at 1h, ~8% relative; 0.01590 vs 0.01669 at Daily, ~5% relative) — but nowhere near the 44%/40% improvement the Oscillator Decay Projection section found by anchoring specifically at REAL PEAKS. The honest lesson: the damping model's real value comes from knowing you're AT a peak and projecting the decay from there — applied blindly at every random moment regardless of where you are in the cycle, the edge nearly disappears. This sharpens rather than contradicts the earlier finding.

Root cause, confirmed on real numbers, then visualized — operator's own instinct: "sounds like we need to debug this... plot the position based on our oscillatory equation of motion with dampening, and visually see how it matched up with price." Hand-traced one real anchor step by step: |Z[t]| (the true complex magnitude) DOES decay correctly, geometrically, exactly as designed. But Re(Z[t]) — the value read out as the projected return — stayed positive for the first 29 consecutive steps, because the rotation rate (theta=2π/period, period 85–115 candles) is far too slow to flip sign within any reasonable forecast horizon. That's a timescale mismatch, not a coding bug: the rotation speed was borrowed from the SLOW volatility envelope's own real, validated rhythm, then applied to the FAST, memoryless return-direction phase — two different real things that don't share a clock. Confirmed across all 4 sampled real Daily anchors (every one showed 20+ consecutive same-sign projected returns) and visualized directly on the EoM Debug dashboard — one anchor projects a 4.5× price explosion while real price actually fell.

Fix confirmed — operator's own diagnostic question: "is the equation wrong or we're missing some pricing scale or normalization or angular velocity or angular acceleration units are off?" Measured the REAL angular velocity directly from the (velocity, acceleration) phase's own candle-to- candle rotation (proper wrapped circular difference) instead of borrowing the volatility envelope's fitted period. Result: the real rotation is 24–32× faster than what was used (1h: 1.81 vs 0.075 rad/candle; Daily: 1.75 vs 0.055 rad/candle) — a real, measured confirmation of the diagnosed mismatch, not a guess. Also found and corrected a smaller, secondary issue: differencing amplifies acceleration's variance relative to velocity (~1.4×), so acceleration was rescaled onto velocity's own scale before combining them into the initial state.

Timeframe Price MAE (original) Price MAE (corrected) Price MAE (flat baseline) |Return| MAE (original) |Return| MAE (corrected)
1 Hour6.08%1.53%1.52%0.002930.00275
Daily67.1%9.52%9.33%0.015900.01465

The fix works exactly as diagnosed. Corrected price MAE (1.53%/9.52%) now sits almost exactly ON TOP of the flat baseline (1.52%/9.33%) — the catastrophic 7-10x blowup is gone. Honest limit: it doesn't BEAT flat either — direction still carries no real edge, confirmed yet again, exactly as expected given everything else this session. What the fix actually bought: turning an actively harmful model back into a harmless, realistic one. The |return| magnitude side did pick up a genuine, if modest, extra improvement (now ~13% better than flat at both scales, up from ~5-8% before the fix) — a small real gain from getting the rotation rate right, on top of the much bigger gain from anchoring at real peaks (see above).

Does refitting LOCALLY (1-2 recent cycles) beat the global, whole-history fit? Tested on 1m, operator's own ask: "lets then start with micro movement 1m and lets just train on 1 or 2 cycles to try to get as accurate a prediction." First checked what "1 cycle" should even mean for 1m's own (velocity, acceleration) rotation — its empirical period came out to ~3.5 candles, suspiciously close to the ~3.5-candle figure already found at BOTH 1h and Daily despite those being wildly different real timescales. That consistency is itself evidence this rotation is an artifact of returns being close to memoryless noise at every scale, not a real cycle worth training a window around (3-7 data points isn't a training set). Used the one REAL validated cycle instead — 1m's own volatility-envelope period (~1440 candles, ~1 day) — and refit both angular velocity and damping from just the preceding 1 or 2 days of real data at 400 real anchors, 15-minute horizon, causal (no lookahead).

Variant Price MAE (400 anchors, 15-candle horizon)
Flat baseline0.1097%
Local fit, 1 cycle (preceding day)0.1146%
Local fit, 2 cycles (preceding 2 days)0.1146%
Global fit (whole history)0.1149%

Clean null — local refitting doesn't meaningfully help. All three model variants land within 0.005 percentage points of each other, and the flat baseline is still marginally the best of all four. Refitting from a fresher, more locally-relevant window doesn't unlock anything a global average was missing — at every granularity tried, direction has no real predictable structure at 1m's own short horizon either. An honest, complete answer, not a partial one.

Could a genuinely SIMPLER equation do better? Tested directly, operator's own question: "could we not do better to fit the line to price? should be a simple equation." Built the simplest possible forecast — a plain least-squares straight line fit to the last 10 real closes, extrapolated forward, no rotation, no damping, no (velocity, acceleration) phase at all — and compared it against the physics model and flat on the identical real anchors and horizon already used above.

Timeframe Flat baseline Physics model (corrected) Simple linear trend
1 Hour0.750%0.786%1.119%
Daily3.906%4.218%5.827%

Simpler is worse here, not better — the plain linear trend is the WORST of all three at both timeframes, clearly. The reason is structural, not a fluke: a straight line has NO damping at all — whatever the last 10 candles happened to be doing, it extrapolates that forever, with nothing pulling it back. The physics model's damping (even though it doesn't add real predictive power over flat either) at least prevents that unchecked runaway extrapolation. Ranking held at both scales: flat < physics model < linear trend. "Simple" and "accurate" aren't the same thing when the ingredient a naive line is missing — decay toward no-opinion — is exactly the one piece that matters once you commit to a directional guess at all.

Path Predictor  

evolution/candle_path_predictor_v1.py · NEAT/GPU, not Monte Carlo

A genuine NEAT/GPU-evolved network that autoregressively rolls forward a multi-step price path (each step's prediction feeds back in as the next step's input) — not a single-point forecast, and not a Monte Carlo / bootstrap scenario generator. Reuses this project's own proven GPU-batched masked-NEAT substrate (the same population/tournament/mutation pipeline already validated on the jump game); only the fitness (path accuracy vs. a "predict no movement" baseline) is new. Closes the exact gap evolive2's dormant "world model" (v3/runners/world_model_runner_v3.py) never did — that one only ever trained single-step prediction and stubbed multi-horizon outputs by reusing the t+1 value.

Training (in-sample)

Daily BTC, warmup=55 candles, horizon=13 candles, 200 generations
Best skill score
vs. "predict no movement" baseline0.0

Holdout (genuinely unseen)

Mean skill score
Segments scoring above baseline

Escalation: a deliberately OVERFIT version, 1 Minute, 2 real cycles only. Operator's own ask, after the physics-equation and simple-linear-trend forecasts both failed to beat flat: "lets build a overfitted neat network with gpu to do the same, and predict the path for a short 2 cycle 1 min candle." Reuses this exact same module UNCHANGED (`scripts/run_candle_path_predictor_1m_overfit_v1.py`) -- only the SCOPE changes: instead of training across the full multi-year history for generalization, the network only ever sees ~2,680 real 1-minute candles (2 real, already-validated volatility-envelope cycles, ~1.9 days), deliberately allowed to overfit that specific small window. A genuinely unseen real tail (200 candles, immediately after, never touched during training) is the honest holdout.

Training (in-sample, deliberately overfit)

1 Minute BTC, warmup=100, horizon=15, 2,680 real candles, 200 generations
Best skill score0.1376
vs. "predict no movement" baseline0.0

Holdout (200 genuinely unseen real candles)

29 overlapping real windows, same champion genome
Mean skill score-0.1912
Segments scoring above baseline3 / 29

Negative — overfitting to a real, recent, short window doesn't help either. The network genuinely learned SOMETHING in-sample (0.1376, clearly above zero, not a saturated metric) — it just didn't transfer even one step forward in real time. The first 3 holdout windows (chronologically earliest, closest to the training cutoff) score positive (0.20, 0.23, 0.03) — an initial-condition echo, not durable skill — before decaying to consistently negative for the remaining 26. Magnitude (-0.19) lands almost exactly on the original Daily NEAT path predictor's own holdout result (-0.213) above, despite a completely different timeframe, window length, and training scope. Sixth independent confirmation this session that price direction carries no learnable structure a model can actually generalize on, no matter how the training data is scoped.

Trend Confirmation Coin flip

evolution/trend_confirmation_v1.py · scripts/run_trend_confirmation_check_v1.py

A deliberately different question from every other entry above: instead of PREDICTING direction before it happens, wait for a simple moving-average crossover to CONFIRM it's already underway, then measure real forward continuation net of the honest cost of waiting. Operator's own framing, after six straight negative direction-prediction results: "is just unpredictable and gambling? there are obvious bull and bear runs, stay out until those are identified?" Plain 20-period SMA (untuned), 20-candle forward horizon, real Daily and 1h candles.

Timeframe Confirmations Mean signed forward return Hit rate Mean lag cost
1h2,154+0.00085%49.16%1.64%
Daily67-0.171%50.75%8.69%

Negative, cleanly. Both hit rates sit within noise of the 50% coin-flip baseline — confirmation identifies continuation no better than chance. 1h's mean forward return is indistinguishable from zero; Daily's is actually negative. Meanwhile the lag cost is real and non-trivial: by the time a 1h crossover confirms, 1.64% of the recent swing is already gone; on Daily, 8.69%. A genuine edge would need to clear that cost before it's worth anything, and there's no edge to clear it with. Waiting for a "confirmed" trend isn't a way out of the unpredictability — it's the same null result wearing a lagging disguise, and it adds a real, measured cost (missed entry) on top. Seventh independent confirmation this session (raw ACF, entry-primitive hit rates, the Markov cascade, the trained attention predictor, the physics equation/linear trend, the overfitted 1m NEAT path predictor, now this) that this market's direction carries no exploitable structure at any horizon tried, predicted in advance or confirmed after the fact.

Blind Driver / Observability Ceiling Proof, not a primitive

evolution/blind_driver_track_v1.py · scripts/run_blind_driver_track_check_v1.py

Every negative result above had to be INFERRED from real, noisy market data. This is a controlled proof with the ground truth under our own control — operator's own analogy: "turning left or right, it can see what is going on ahead of time... in our case going up or down, we can't see what is in front of us in the future? what happens if we made up the track randomly for the driver to follow, it knows to turn left and right but can't see what's in front of it until it gets there?"

A "track" is a sequence of i.i.d. random left/right turns — unpredictable turn-to-turn BY CONSTRUCTION. Two GPU-batched masked-NEAT populations, identical topology/population/generations, identical single-scalar input width — only the INFORMATION in that scalar differs. Sighted: input[t] = turns[t] itself, a forward sensor that sees the very turn it must decide. Blind: input[t] = turns[t−1] only — reactive, exactly like a candle feature built from real past closes. Since turns are i.i.d., nothing in the blind input encodes turns[t] — the information-theoretic ceiling on accuracy is exactly 0.5, for any architecture, however long trained.

Sighted (forward sensor)

150 generations, population 120, 60-turn tracks
Training best accuracy100%
Holdout accuracy (200 fresh tracks)100% (std 0.0)

Blind (reactive only)

Same settings, only the input's information differs
Training best accuracy59.4%
Holdout accuracy (200 fresh tracks)50.8% (std 6.7%)

Exactly as predicted. The sighted driver hits perfect accuracy and holds it on every unseen track — the substrate genuinely can learn when the answer is observable. The blind driver's training accuracy (59.4%) looks like it beat chance, but that's the GA overfitting to the small, specific training tracks it actually saw — on 200 genuinely fresh tracks, holdout accuracy is 50.8%, squarely within one standard deviation of exactly 0.5. Same shape as the deliberately-overfit 1m NEAT path predictor above (real in-sample signal, negative/chance holdout), but here the ceiling is known in advance, not inferred. This is the mechanism behind every negative result on this page, stated plainly: a model can only be as good as what its own input actually contains about the future. Price has no forward sensor the way a road has one — no architecture, training budget, or amount of overfitting changes that. Not counted as a primitive family — this is a proof of the limit, not a candidate signal.

Follow-up — operator's counter-hypothesis: "it still follows a path, that's gradual... your random generator still might have some pattern it could recognize and exploit... the driver might recognize your track favors left turns or rights." Two claims, tested for real: does more MEMORY let a blind driver exploit real bias/momentum, and was the original track secretly biased? Extended the module with a tunable continuation_probability (real first-order Markov bias — 0.5 provably reduces to the original i.i.d. track) and a memory_length window (several real past turns, not just one).

Condition Training best Holdout accuracy (200 fresh tracks)
Blind, memory=5, unbiased (p=0.5)59.4%50.4% (std 6.3%)
Blind, memory=5, biased (p=0.7)75.4%69.5% (std 6.0%)
Sighted, biased (p=0.7), control100%100% (std 0.0)

Both claims resolved, in opposite directions. More memory on the genuinely unbiased track changes nothing (50.4%, same as the single-lag original) — memory alone can't manufacture a pattern from a truly memoryless process. But on a track with REAL bias (p=0.7), the same architecture converges to 69.5% — almost exactly the information-theoretic ceiling for a first-order Markov chain ("predict same as last turn" is right with probability p itself). The substrate genuinely can find and exploit this class of pattern when it exists. This reframes every negative result on this page: the capability gap isn't the model's — it's proven to exist here. The open question is empirical, and this session already answered it several times over (raw ACF, the Markov transition table between candle-shape clusters, the trend-confirmation crossover) — real BTC direction measures at base rate, not at some real p≠0.5. Real price behaves like the unbiased track, not the biased one — not because the method can't find momentum, but because the measured momentum isn't there.

Fee-Aware Race Track (Conviction) Solved, after 5 real negatives

evolution/vertical_track_fee_aware_v1.py · scripts/run_vertical_track_fee_aware_train_v1.py

Operator: "let's see if we can solve an easy problem by considering each round to cost fee... it must learn the big moves and not react to the little ones and get scared out of its position, it needs to have conviction to its decision." Directly the same whipsaw/fee-churn lesson already learned the hard way on real BTC (the volatility straddle, the Donchian controller, the candle-road order selector) — tested on the controlled toy first. Same track, same 5 real primitives, same architecture as the already-validated vertical_track_v1.py combined controller (76.6% on-track) — the only change is a real cost for every unit of change in the network's own continuous steering output.

ConfigMarginFee weightOn-trackCorrection cost
Original baseline0.050.076.6%0.314
Attempt 10.050.274.0%0.356
Attempt 20.050.0573.4%0.347
Attempt 3 (+self-feedback)0.050.267.9%0.350
Attempt 4 (margin only)0.150.088.9%0.328
Attempt 5 (WIN)0.150.290.3%0.090

Five real, independent negatives before the fix, reported plainly rather than only shipping the win. Fee pressure alone (attempts 1–2, at two different weights) made correction cost go UP, not down. A specific, testable hypothesis — the network never sees its own previous decision as an input, so "stick with what I just did" has to emerge from implicit recurrence alone — led to attempt 3 (explicit self-feedback of the prior steering value, the same principle already validated in the discovery-trainer work). Also negative. The real confound, found by re-examining the mechanics rather than guessing again: the original margin (0.05) is tight enough that the on-track reward ITSELF already forces near-constant correction just to stay within a narrow band on a genuinely noisy track — avoiding a correction to save fee cost immediately loses the dominant reward anyway, so the fee penalty never had real room to matter. The same lesson the trailing-stop guard rail already taught this session. Attempt 4 confirmed margin alone isn't the answer either (better tracking, cost barely moved) — the fee pressure needed the room, not a replacement for it.

Attempt 5, margin AND fee together: real, decisive, best-of-all-runs on BOTH axes at once — 90.3% on-track (highest of any run, including the original) and a ~71% reduction in correction cost. s_curve (the real single reversal) stayed at 100% on-track and 100% turn-catch-rate — still reliably catches the real reversal — with one honest tradeoff: mean lag among caught turns rose to ~2.9 steps (from ~1.0 in every other run). Genuine conviction costs a little reaction speed, exactly what "don't react to the little ones" implies for confirming the big one too.

Visually verified before trusting the numbers: traced the real steering sequence on holdout tracks. On s_curve, steering locks decisively negative for the entire first real regime, then cleanly flips positive for the second — one clean commitment, not jitter, despite real per-step noise not aligning with the ideal direction every step. On straight_zigzag (the fastest real chop, alternating every single step by construction), the champion correctly does NOT chase every alternation — it smooths over long stretches instead of whipsawing, exactly the asked-for behavior.

Not done without being asked, per the operator's own explicit sequencing ("we also need to add leverage and liquidation after we solve this"): the next step is extending this track with real leverage sizing and a liquidation mechanic, mirroring trading_env_v1, to test whether genuine conviction survives real leverage risk — not started this round.

Trend-Following Controller Negative

evolution/trend_following_controller_v1.py · scripts/run_trend_following_controller_check_v1.py

The "race car" toy proved a general lesson: pure prediction from lookback drifted badly, but a controller that reacts to real, already-realized feedback tracked the road far better. Operator: "use what we learn from the race car, rebuild it to just follow the trend." This is the direct real-market version — a classic Donchian-channel breakout: go long on a new real high, short on a new real low, hold otherwise. No prediction anywhere, purely reactive, exactly like the race car's feedback controller. Two lookback choices: a standard fixed 20-candle window, and an ADAPTIVE window set from this timeframe's own real, already-validated volatility-envelope period (84 candles at 1h, 115 at Daily, from the earlier oscillator fit on |returns| — not the fast, already-falsified raw-return rotation).

Timeframe Fixed (20) return Adaptive return Adaptive lookback Buy & hold
1h−37.7%−37.4%21−4.2%
Daily−54.0%−41.1%29+20.5%

Negative, and instructive about WHY. Both lookback choices lose badly and badly underperform buy-and-hold at both timeframes — the oscillator-informed lookback barely differs from the plain default (21 vs. 20 at 1h, 29 vs. 20 at Daily), since a quarter of the real envelope period lands close to the textbook value anyway. At 1h, 386-402 flips ate 38-40% of the account in pure fee drag — the same whipsaw signature this project found at the very start of this session. The race car's lesson still holds (react to real feedback, don't predict), but a breakout system's whole premise is that a new real extreme tends to keep extending — and this session has repeatedly found no such real persistence in raw BTC direction (ACF, trend confirmation, the Markov cascade). Reactive correction fixed a real problem in the toy (it had a real road to track); it can't manufacture a trend that isn't there in the market. Consistent with, not contradicting, every other result on this page.

Candle Road Controller Mechanism demo, real data

evolution/candle_road_feedback_controller_v1.py · live replay

Operator: "the High and Low is the boundary, the average open and close is the center of the road, now use the same car to drive the BTC candle road." The exact same closed-loop feedback mechanism as the toy race car (evolution/blind_driver_track_v1.py's feedback controller), now on real Daily BTC candles instead of a synthetic track: the road's center is log((open+close)/2), the lane boundary is the candle's own real high/low range — a genuine, market-defined tolerance, not an arbitrary percentage. Trained on random real windows, tested on a genuinely unseen 90-candle holdout tail.

First pass trained on Daily candles only (689 total, heavy window overlap). Operator asked how much training this really was and whether it held up — re-trained on real 1h candles instead (18,057 training candles, 500-candle genuine holdout, 400 generations, doubled window count) for a far more statistically meaningful check. Numbers below are the 1h, better-trained result.

No correction (zero-weight baseline)

Never steers — the honest "do nothing" comparison
On-road fraction (500-candle holdout)4.8%

Feedback-only controller

Reacts to its own real tracking error every candle — no lookback, no prediction
On-road fraction (500-candle holdout)12.0%

Combined: lookback + oscillator prediction + feedback

Adds a real returns lookback window and a causal oscillator peak-decay projection
On-road fraction (500-candle holdout)13.8%

Same qualitative story as the smaller Daily run, now confirmed on far more data: feedback correction is a real ~2.5x lift over doing nothing (4.8% → 12.0%); adding a genuine lookback window and a causal oscillator peak-decay projection lifts it a further, real but modest amount (12.0% → 13.8%), nowhere near the dramatic jump the same combination produced in the toy (65.6% → 86.8%). Absolute on-road numbers are lower than the Daily run's (40.4%/43.8%) because 1h's holdout run is much longer (500 steps vs. 90) and its real per-candle ranges are proportionally tighter — a genuinely harder continuous-tracking task, not a worse result. Consistent with the rest of this session: the toy's 5 named primitives had strong, clean, deterministic structure to recognize; real BTC candle patterns and oscillator-projected volatility carry much weaker, noisier structure, so lookback+prediction has far less to add on top of pure reactive correction. Framed explicitly as a mechanism demonstration continuing the blind-driver work, not a new trading primitive — no position, no fees, no P&L here, just a real, honest test of how well "react to real feedback, plus genuine lookback/prediction" tracks the market's own actual geometry. No tally change.

Candle Road Order Selector Negative

evolution/candle_road_order_selector_v1.py · scripts/run_candle_road_order_selector_check_v1.py · live replay

Operator, watching the candle-road replay: "it looks like its kinda lagging, but some areas it predicted the correct direction... if we had another selector for placing orders long or short to open or close or flip and applying trailing stop when needed, what would we get." The combined controller's own per-step steering decision is a genuine, causal directional bet (decided before that candle's real close is known) -- recovered exactly from its already-computed tracked price, gated by a confidence threshold, and turned into real open/hold/flip/exit orders with an ATR-based trailing stop. Reuses the Trend-Following Controller's own fee-aware returns evaluator unchanged. Tested as a full, untuned sensitivity grid (3 confidence thresholds × 3 trailing-stop multiples = 9 cells) on the same real 500-candle 1h holdout, so the result can't be an artifact of picking whichever cell looked best.

Confidence threshold Real signals fired Return range across trailing-stop multiples Fee drag range
0.2396−11.9% to −11.5%12.8–13.2%
0.3334−12.7% to −11.7%10.2–10.6%
0.5219−6.2% to −3.3%7.4–7.6%

Negative across all 9 cells of the grid, cleanly. Every combination loses money against a real buy-and-hold return of essentially flat (+0.03%) on this holdout. Fee drag alone (7.4%–13.2%) explains most of the damage -- the same whipsaw signature this whole session found at its very first check (the original 11-primitive router's 36.7% holdout churn). The combined controller's own real, measured tracking improvement (12.0%→13.8% on-road) is real but far too small and far too noisy, once converted into discrete buy/sell decisions, to survive realistic round-trip fees at 1h granularity. Consistent with every other attempt this session to turn any real signal into actual trades at this timeframe.

Multi-Timeframe Signal Fusion Negative

evolution/multiframe_signal_fusion_v1.py · scripts/run_multiframe_signal_fusion_check_v1.py · live replay

Operator: "lets run 1 hour and Daily candle, 1m gets 1x weight, and 1 hour gets 60x weight, and 1 day gets 1440x weight, and the biggest weight get the decision... 1m will sum up its weight to compare with 1 hour weight, when 1 hour is establish it is then accounted and accumulated forward." Trained the combined lookback+oscillator-prediction+feedback controller at THREE real timeframes (1m, 1h, Daily) on a calendar-aligned real 30-day holdout (2026-04-04 to 2026-05-04, verified present in all 3 pools: 43,200 real 1m candles, 720 real 1h candles, 30 real Daily candles). Each real 1m candle casts a ±1 vote; once its real hour closes, the hour's own steering casts a native ±60 vote, and the SIGN of their sum becomes that hour's fixed ±60 contribution forward. The same rule repeats one level up against the Daily controller's own ±1440 native vote. "Biggest weight wins" falls out of the arithmetic itself -- a confident higher timeframe can only be overridden by a genuinely unanimous lower one.

Honest caveat surfaced during training: the 1m controller's own real on-road tracking fraction over a full continuous 43,200-step holdout run was effectively zero (0.005%) -- compounding error over that many sequential steps is catastrophic at 1m's own tiny step scale. Its per-step steering SIGN is still well-defined and used as a real vote, but this is flagged plainly, not hidden.

MetricValue
Daily decisions (30-day holdout)22 long / 8 short / 0 flat
Total return, WITH trailing stop−8.08%
Total return, trailing stop DISABLED (pure signal)−6.24%
Real buy & hold, same window+17.44%
Fee drag (with trailing stop)2.55%

Negative, and the diagnosis was checked directly, not assumed. Real BTC rallied +17.4% over this 30-day window, and the fused signal was net long most of it (22 of 30 days) -- yet still lost money. Isolated the cause before reporting it: re-ran with the trailing stop disabled entirely, and the pure directional signal STILL lost 6.24% against the real rally -- a ~24-point gap that has nothing to do with execution or stop-outs. The trailing stop adds a further, real but secondary drag (~1.8 points, plus extra fee churn). The core problem is the fused directional calls themselves getting the specific days wrong, not the mechanism that turns them into trades. Consistent with every other attempt this session to convert a real, measured tracking improvement into actual trading edge -- being right about the market's own geometry more often than not (the on-road results) doesn't automatically mean being right about DIRECTION on the days that matter most.

Halving-Cycle Timing Theory Validated

evolution/halving_cycle_strategy_v1.py · scripts/run_halving_cycle_strategy_check_v1.py · live replay

A genuinely different hypothesis, sourced from a separate local project (tradebotsecrets_anvil_app), not derived from candle geometry at all. Operator: "we timed btc high and low base on halving duration... if we use this idea and rerun last 6 months simulation would we be right?" then "lets build it out and simulate it to the beginning of bitcoin birth." The rule (read directly from that project's own strategy_cycle_module.py): BTC's cycle top falls a fixed 526 real days after each halving, and its cycle bottom a fixed 889 real days after each halving — pure calendar arithmetic, computable the day of the halving with zero price information. Go SHORT from a cycle's own predicted top until that same cycle's predicted bottom, LONG otherwise.

Real data sourced for a genuinely non-hindsight, full-history test. CoinGecko's free API restricts historical queries to the past 365 days; switched to blockchain.info's full-history daily price API, saved as candles/btc_full_history_daily/btc_full_history_daily.json — 6,424 real daily points from 2009-01-03 onward. The leading zero-price stretch (2009-01-03 to 2010-08-17, before any real exchange existed) is dropped before backtesting.

Timing accuracy checked against real historical extrema first, searching a bounded window around each predicted date (±365 days for tops, ±200 for bottoms, to avoid catching the NEXT cycle's own pre-halving pump instead) across all 4 real halvings:

HalvingPredicted topReal topDiffPredicted bottomReal bottomDiff
2012-11-282014-05-082013-12-05 ($1,136.90)-154d2015-05-062015-01-15 ($172.00)-111d
2016-07-092017-12-172017-12-17 ($19,279.90)0d (exact)2018-12-152018-12-16 ($3,231.91)+1d
2020-05-112021-10-192021-11-09 ($67,562.17)+21d2022-10-172022-11-22 ($15,759.61)+36d
2024-04-202025-09-282025-10-07 ($124,776.68)+9d2026-09-26not yet occurred

One exact-day match and two more within the theory's own stated ±113-day tolerance, across all 3 fully completed real cycles — explicitly flagged as n=3-4 macro samples, not a large-N statistical proof, but the closest any timing model has matched real history all session.

Full-history backtest, reusing trend_following_controller_v1_evaluate_returns UNCHANGED for fee-aware P&L (5bps taker fee), across the entire real, non-zero-price BTC history (2010-08-17 to 2026-08-04, 5,832 real daily candles). Positions decided from halving dates alone, never from price — a genuinely non-hindsight test.

MetricValue
Outperformance multiple (strategy equity ÷ buy-and-hold equity)595.5×
Total flips, whole history7
Total fee drag, whole history0.007%
Real price change, cycle 1 SHORT phase (predicted top→bottom)-46.5%
Real price change, cycle 2 SHORT phase-83.2%
Real price change, cycle 3 SHORT phase-68.9%

Positive, and checked per-cycle rather than as one aggregate number. Because BTC's own baseline appreciation is astronomical (~915,000× over this window), raw total-return percentages are unreadable on either curve — the outperformance MULTIPLE is the real signal, and it isn't one lucky early-history fluke: on each of the 3 fully-completed real cycles, going short on the model's own predicted (non-hindsight) top date and covering on its own predicted bottom date captured a real -46% to -83% drawdown. Only 7 flips across Bitcoin's entire real history means fee drag is negligible — structurally the opposite problem from every high-frequency signal in this registry, which was consistently eaten alive by round-trip fees before direction even mattered. Caveat stated plainly: n=3 completed cycles is a small sample for a macro claim, and the theory's current live call (predicted bottom 2026-09-26) has not yet resolved — a strong historical backtest, not a forward guarantee.

DCA Ladder Spacing Optimization Mixed / Validated

evolution/halving_cycle_dca_ladder_v1.py · evolution/halving_cycle_dca_ladder_optimizer_v1.py · live replay

Operator has a real, live DCA ladder (entry/exit price-distance and size-fraction arrays, biased by the halving-cycle SHORT call) on two real accounts (1x and 5x leverage): "can we simulate some optimize dca position size, and spacing, to see what is best... use the tradebot_secret uplink folder algorithm to simulate it." Reproduced the REAL algorithm, not a guessed convention: entry prices chain MULTIPLICATIVELY off the previous entry (not a fixed reference); exit prices chain off min/max(avg_entry, current_price); entry sizing is independent fractions of the total budget; exit sizing is a fraction of the ORIGINAL position at every level (not the shrinking remainder), capped at whatever's still open so it can never overshoot; the whole ladder re-centers every real autorun tick, and the entry budget each tick tracks the real "Available Margin" formula (a winning position's unrealized P&L genuinely grows the NEXT tick's budget, a losing one shrinks it), not a static allocation. Two real corrections landed mid-thread, both caught by the operator directly, not found independently: the exit-sizing formula was first read from a different, non-authoritative sibling repo whose logic looked cascading ("in my dca.tradebotsecrets.com its calculated as the original position... its teh 24hour force autorun that calculates it" pointed at the actual live uplink repo); then a second gap surfaced when asked directly ("i also have available margin setting set, so it recalculates base on whats available, is that the same simulation you have?") -- it wasn't, fully, until the compounding fix above. All results below are from the fully corrected mechanics. Verified with hand-computed proof tests (including a real forced-liquidation scenario and a test proving a winning position measurably grows the next entry's budget) before trusting any search built on top of it.

Backtested on real 1h BTC candles (7,280 candles) for the CURRENT halving cycle's own SHORT window (2025-09-28 predicted top through the latest real data, real price -42.1%). A first single-split GA search found that 3 of 4 "optimized" configs underperformed the operator's own baseline out-of-sample -- a real overfitting result, reported honestly rather than discarded. Fixed with 3 real walk-forward folds seeded by the operator's OWN real ladder (so the search can only match or beat it, never wander off), median-aggregated per account, validated on a 4th, completely untouched final slice.

MetricYour ladder1x-optimized (drawdown-adjusted)
Final untouched holdout, 1x+1.39%+2.74%
Final untouched holdout, 5x+5.06%no reliable improvement
Full window, 1x+47.74%+73.20%
Full window, 5x+469.13%(same, no fabricated win)

Only a real, modest win at 1x -- and the available-margin fix specifically is what took the 5x "win" away. Before that fix, 5x looked like the strongest result on this whole page (drawdown- adjusted spacing beating baseline by 2x); after fixing the compounding dynamics, that same search's full-window number got EVEN BIGGER (+756.6% at one point) while its genuinely held-out result went NEGATIVE (-6.06%, worse than just keeping the operator's own ladder). Fold-to-fold holdout returns for 5x swung from -34% to +84% -- real instability from leverage-amplified compounding exposure, not noise to average away. Practical takeaway surfaced by this, independent of any spacing search: "Available Margin" mode means a winning position's own gains grow its next entries -- at 5x that compounds real exposure deeper into whatever trend is happening, which cuts both ways; "Total Wallet Value" (a fixed budget) may be the structurally safer choice for a leveraged account regardless of spacing. The entry/exit COUNT check (11 entries/2 exits beat 7/3 for 1x) was run BEFORE this fix and has not been re-confirmed under the corrected mechanics -- flagged as open, not re-claimed.

Safe leverage, grounded in the real simulator (not a back-of-envelope estimate), re-verified under the corrected mechanics: sweeping the operator's own ladder shape up in real leverage against the full real window, 5x survived (no liquidation) but 7x was liquidated partway through the holdout period -- meaning the operator's existing 5x account was already sitting at the edge of what this one real historical path would tolerate, not a comfortable margin. The liquidation formula ignores maintenance margin/funding, so real safety is somewhat lower than this number -- 3-4x is the more genuinely conservative range.