Full-history pull off production signals_perf.db (263 calls, 2026‑07‑16 → 2026‑07‑21, read-only via the owner-gated /api/perception-desk API). Nothing in the live engine touched. This is source material for a backtest — not applied changes.
63% (165/263) are still tracking — these numbers can only get BETTER as peaks are still forming, never worse. Every stat below is honest against real stored outcomes as of pull time.
| Threshold | Count | % of 263 |
|---|---|---|
| 10x+ | 11 | 4.2% |
| 5x+ | 26 | 9.9% |
| 3x+ | 51 | 19.4% |
| 2x+ ("win") | 95 | 36.1% |
| 1.5x+ | 131 | 49.8% |
| <1.5x (tail) | 132 | 50.2% |
| rugged (eventually) | 156 | 59.3% |
avg_multiple disagreeing between PERCEPTION's own DB (1.84x) and the public leaderboard's per-caller stat (5.26x) for the same 263ish rows — likely a computation-window or dedup difference between the two code paths, not a live bug I'm touching. Flagging so it gets reconciled; I used the raw per-row DB data throughout this report, not either pre-aggregated number.Correlation of eyezon_score to log(peak_multiple): r = ‑0.06 — statistically flat, if anything the wrong sign. The score, as currently weighted, carries almost no information about how big a call will run.
| Tier | n | avg mult | median | ≥2x% | ≥5x% | rug% |
|---|---|---|---|---|---|---|
| STRONG (≥78) | 1 | 4.94 | 4.94 | 100% | 0% | 0% |
| SOLID (62‑77) | 16 | 2.45 | 1.88 | 43.8% | 6.2% | 75% |
| OK (45‑61) | 38 | 2.33 | 1.50 | 36.8% | 10.5% | 68.4% |
| WEAK (<45, record-only, never posted) | 179 | 3.12 | 1.52 | 37.4% | 9.5% | 59.2% |
WEAK — the tier that structurally can never broadcast — has a higher average multiple and materially lower rug rate than SOLID. This isn't noise at the edges; it's the largest bucket in the dataset (179/263, 68%).
| n | avg mult | median | ≥1.5x% | ≥2x% | ≥5x% | rug% | |
|---|---|---|---|---|---|---|---|
| BROADCAST (score≥45, what gets posted) | 55 | 2.41 | 1.51 | 54.5% | 40.0% | 9.1% | 69.1% |
| RECORD-ONLY (score<45, logged but never posted) | 208 | 5.36 | 1.47 | 48.6% | 35.1% | 10.1% | 56.7% |
The 45-point broadcast floor is filtering ~79% of PERCEPTION's calls (208/263) for a hit-rate that is statistically indistinguishable from — and on this sample slightly worse than — what it silently discards. The all-time best call, $Jimothy at 412.7x, scored 63/tier=null and would still have cleared today's 45 floor, but dozens of other real winners sit below it: the WEAK bucket alone contains 9.5% of its population hitting 5x+.
| feature | winners(5x+) med | duds(<1.5x) med | corr w/ log(mult) | read |
|---|---|---|---|---|
| authenticity (62% of score weight) | 80 | 78 | ‑0.037 | flat/noise |
| buyers | 229 | 287 | ‑0.149 | inverted — see §3 |
| sellers | 101 | 166 | ‑0.160 | inverted |
| buy_pct (dominance) | 64.1 | 61.5 | — | weak alone, strong combined (§4) |
| liquidity ($) | 14,628 | 14,618 | — | flat |
| whale_count | 0 | 0 | — | dead — see §3 |
| crowd_count | 0 | 0 | — | dead — see §3 |
| top10_pct | 36.5 | 39.4 | +0.028 | flat, wrong-signed vs the penalty |
| bundle_score | 55 | 55 | ‑0.078 | weak alone |
| liq_mcap_ratio | 0.34 | 0.33 | ‑0.038 | flat alone |
| lead_time_sec (NOT scored today) | 254 | 311 | ‑0.152 | real, unused |
auth_min already hard-rejects the worst readings before evaluate() ever computes a conviction score. Among survivors, authenticity clusters tight (72‑80 across both winners and duds) — the single heaviest-weighted term in the formula (0.62×) carries almost no residual separating power on the population that actually reaches scoring. Most of the real information is riding on the other 38% of the weight — several legs of which are dead (whale/crowd) or wrong-signed (top10 penalty).
| buyers at entry | n | avg mult | ≥2x% |
|---|---|---|---|
| 0‑60 | 22 | 6.40 | 27.3% |
| 60‑150 (sweet spot) | 31 | 5.68 | 64.5% |
| 150‑300 | 86 | 3.20 | 36.0% |
| 300+ (crowded) | 112 | 2.07 | 33.0% |
60‑150 buyers at entry is the best bucket in the whole dataset (64.5% hit 2x+, more than 3x the win-rate of the 300+ "most active-looking" bucket). The current authenticity sub-term (_lin(buyers,8,60)) treats 60 buyers and 1,000 buyers identically — it never penalises an entry that's already crowded/late, which is exactly the failure mode a "quiet real winner" (Pete's flag) falls into: fewer buyers at entry reads as lower authenticity today when it's often the more room-to-run entry.
| top10 concentration | n | avg mult | ≥2x% |
|---|---|---|---|
| 0‑20% | 2 | 1.31 | 0% |
| 20‑40% | 118 | 3.59 | 37.3% |
| 40‑60% | 87 | 2.01 | 34.5% |
| 60%+ (flagged as risk, +20 to compound gate) | 19 | 3.56 | 57.9% |
Consistent with the codebase's own prior notes ($Himgajria at 87.5% concentration was a real winner) — concentration alone is not the risk signal the compound gate treats it as. See §4C for the more useful, interaction-based version of this finding: concentration only turns genuinely bad when it stacks with bundle evidence.
Pete's ask: what data points don't currently talk to each other but would help if they did. Tested directly against the 263-call history wherever the fields exist; flagged as hypothesis where they don't.
| combo | n | avg mult | ≥2x% | ≥5x% |
|---|---|---|---|---|
| buyers 60‑150 + entry mcap <$50K | 26 | 6.45 | 69.2% | 26.9% |
| buyers 60‑150 + mcap $50‑100K | 5 | 1.68 | 40.0% | 0% |
| buyers 300+ (crowded) + mcap <$50K | 57 | 1.88 | 26.3% | 7.0% |
Neither feature is this strong alone. Cheap mcap alone gives 5x+ at 13.5% (§5); the buyer sweet-spot alone gives 2x+ at 64.5%. Together (n=26): 69.2% hit 2x+ and 26.9% hit 5x+ — roughly 2‑3x the dataset baseline. Low mcap alone is not sufficient — the same low-mcap population with crowded (300+) buyers is actually below-baseline (26.3%). Low mcap only pays off when it's still quiet.
| combo | n | avg mult | ≥2x% | ≥5x% |
|---|---|---|---|---|
| fast call (<290s) + sweet-spot buyers | 16 | 7.65 | 56.2% | 25.0% |
| fast call + crowded buyers (300+) | 35 | 2.19 | 40.0% | 5.7% |
| slow call (≥290s), all buyer counts | 81 | 1.94 | 28.4% | 7.4% |
lead_time_sec (currently unscored — §3/§6) and buyer-count are individually modest (r=‑0.15 and non-linear respectively) but the intersection is one of the cleanest signals in the whole dataset. "We caught it fast AND the crowd hasn't arrived yet" is a genuinely different, stronger claim than either alone.
| risk-leg combination | n | avg mult | ≥2x% | ≥5x% | rug% |
|---|---|---|---|---|---|
| BOTH top10≥60% + bundle≥35 elevated | 4 | 6.03 | 75.0% | 50.0% | 100% |
| only top10≥60% | 15 | 2.90 | 53.3% | 6.7% | 53.3% |
| only bundle≥35 | 162 | 2.88 | 34.0% | 8.6% | 58.6% |
| neither elevated | 45 | 3.00 | 42.2% | 8.9% | 64.4% |
n=4 is too thin to act on, but directionally interesting and worth a flag: when BOTH risk legs stack, the survivors that clear the hard-reject threshold anyway look like a distinct extreme-variance regime — big pumps AND a 100% eventual rug, i.e. "real pump, guaranteed rug" rather than "just bad." That's a genuinely different trading instruction (fast in/fast out) than what a single elevated leg implies. Worth deliberately growing this bucket's sample via a shadow (record-only, non-broadcast) log before touching the live gate.
| combo | n | avg mult | ≥2x% | ≥5x% |
|---|---|---|---|---|
| buy-dominance ≥65% + liq/mcap ≥30% (healthy) | 43 | 4.41 | 48.8% | 20.9% |
| buy-dominance ≥65% + liq/mcap <15% (thin) | 2 | 1.31 | 0% | 0% |
Both features are individually near-zero-correlated with outcome (§2). Combined, "heavy one-sided accumulation with real liquidity behind it" clears baseline meaningfully (48.8% vs 36.1% overall). The thin-liquidity leg is too small (n=2) to trust but is directionally exactly what you'd expect — heavy buying against thin liquidity is a fragile, easily-rugged spike, not organic accumulation.
| combo | n | avg mult | ≥2x% |
|---|---|---|---|
| authenticity≥80 + sweet-spot buyers | 25 | 6.22 | 68.0% |
| authenticity≥80 + crowded buyers (300+) | 32 | 2.14 | 31.2% |
The single highest-leverage fix in this whole report: the same high authenticity score (≥80) means a 68% win-rate in one crowding state and a 31% win-rate in another — more than 2x apart on the SAME feature. Since authenticity is 62% of conviction, this interaction (not currently modelled at all) is likely costing more accuracy than any other single item here.
agree ratio is computed live (used as a hard gate at <60%) but the ratio itself is discarded, never stored per call. A call that passed at 61% agreement and one that passed at 99% look identical in the DB today — the gate uses it as a boolean, not a graded confidence signal.| entry mcap | n | avg mult | ≥2x% | ≥5x% |
|---|---|---|---|---|
| <$50K | 163 | 6.43 | 38.0% | 13.5% |
| $50‑100K | 80 | 1.91 | 30.0% | 3.8% |
| $100‑150K | 15 | 2.47 | 46.7% | 6.7% |
| $150‑300K | 5 | 2.03 | 40.0% | 0% |
The code's own prior advisory (compute_optimal_mcap_band) flagged this direction on n=6/n=1 buckets — too thin to trust. At n=163/n=80 the same direction now holds with real weight: sub‑$50K is unambiguously the strongest band (3.5x the 5x+ rate of $50‑100K), and $50‑100K is the clear laggard.
Every source the scoring path (signals_scanner.evaluate()) actually touches, read directly from the code + its own inline incident notes — not guessed.
Free-RPC watch of the pump.fun migration authority, near-zero lag, self-owned — no third-party dependency. Live 10-min head-to-head: GeckoTerminal's new_pools (the old primary) surfaced 0 of 5 real graduations against a true ~25/hr chain-wide rate; GT is now demoted to a fallback/reconciliation pass only. This is the single biggest documented source upgrade in the codebase and it's already shipped — good.
Primary feed for cand.mcap/liq/vol_h1/chg_h1 and the buyers/sellers counts that drive authenticity — the single heaviest-weighted score component (62%). Two documented, currently-unfixed issues: (1) rate-limit/429 storms serious enough to need a dedicated pacer/breaker module (gt_pacer.py); (2) GT's buyers/sellers counts double-count wallets that both buy AND sell in the window ("flips") — live-verified an ~11.5% inflation on a real sample, feeding directly into the unique-wallet-ratio term of authenticity. Documented in the code as a known, accepted gap (no cheap fix available without a second RPC-level walk). This is the clearest case in the whole audit of a slightly-noisy input sitting underneath the most heavily-weighted output.
Direct-source (complete + pump_swap_pool), not an aggregator's inference. GT launchpad_details as fallback only. No documented issues.
Purpose-built after a real incident ($Yuna: entry landed inside an unstable migration-snipe spike). GT-vs-DS agreement gate + MIN(GT,DS) liquidity is good defense-in-depth. Only gap: the agreement ratio itself is discarded after the pass/fail check (§4F) — a graded confidence feature is sitting right there, unused.
Feeds the compound risk gate directly (§3C's top10_pct penalty). The codebase has ALREADY had to add a hard sanity guard (_top_holders_sane) after live-confirmed corrupt reads: topHolders summing to 114.4% ($ANSEMIO), the null system-address listed 5x as a "holder" ($PUMPGU). Also has real coverage gaps — HTTP 400/empty on some brand-new mints (fell back silently to on-chain pnl.top_holders for $AnsemAI). This is the one source where a documented data-quality problem and a documented scoring-accuracy problem (top10_pct is flat/wrong-signed in §3C) plausibly compound each other.
No third-party cost or rate-limit exposure. Two real false-positive bugs already found and fixed live (raw-tx-count vs distinct-wallet-count on $Himgajria; a noisy ==2 sub-case removed after the Agamemnon/RACCTARD backtest) — a good sign of active calibration, but also a sign it's not yet mature. Only computed for candidates that already cleared every cheaper gate (86% coverage, 226/263) — by design, not a gap.
No pagination/history on GT's free trades endpoint — the original case-study wash-burst ($ACT:S) had already rolled off the feed by the time the fix was validated, so thresholds were calibrated from a fresh live re-pull, not the actual incident data. Structurally can only catch a wash campaign that's still active at scan time, not one that already finished.
Was a hard-reject, downgraded to a penalty after live evidence it filtered real 5‑9x winners ($Sagawa class). Good example of the codebase correcting itself against real outcomes — the standard the rest of the matrix should be held to.
Nonzero on 1/251 and 2/251 calls respectively. Not a data bug — a tracked whale or a second EyezOn scanner genuinely hasn't had time to arrive 3‑10 minutes post-migration, which is exactly when PERCEPTION fires by design. The score formula still reserves a flat +15/+15 bonus slot for each that essentially never pays out. These fusions are real and validated on OTHER surfaces (Radar/Aura) — just not at PERCEPTION's entry point.
One of the only bonus terms with an actual documented backtest behind it: n=4 KOL-creator cases, 75% win-rate (≥2x) vs 38.7% for anon creators, p=0.012, zero rugs among the KOL cases at the time. Genuinely the strongest-evidenced single bonus in the formula — also the thinnest (n=4) and not currently re-testable from this pull's data (component not broken out per call, see §6-logging gap below).
Both have 3-strike/5-min circuit breakers — under X rate-limiting they silently degrade to zero contribution with no flag left in the stored row. A call scored during an X outage is indistinguishable, in history, from a call that genuinely had zero chatter. One real bug already found and fixed: raw X "verified" badge was tested and disproved as a trust signal (bot/caller accounts buy verification too) and correctly removed. Not separately logged per call — can't backtest its real contribution today.
Not part of the entry score, but it IS what computes peak_mcap/peak_multiple — the literal ground truth this entire report is built on. This is its own active workstream (the EyezOn Data indexer/persistent-candle-store build) and wasn't deep-audited again here to avoid duplicating that effort. Worth a standing reminder: any gap in this feed silently propagates into every number above it.
signals_min_conviction live — and consider pairing with fix #4/#5 below FIRST, since they're likely why the floor currently has to sit where it does.signals_scanner_max_mcap toward the $50‑80K range
Extends §1-8 above (unchanged, still stands) with an exhaustive pass: every data point EyezOn's marketing/Aura dossier/memory specs claim to use, cross-checked against signals_scanner.py (PERCEPTION's live composite), site/wsgi.py's Aura dossier code (eyezonaura_preview.html's dsxRenderPanel* functions — which self-report live/maturing honestly in their own client code), and a fresh live pull of /api/perception-desk (264 calls, re-confirms §1-8's numbers still hold at n+1). Nothing deployed; nothing in the live engine touched.
| claimed data point | wired into a live score? | shown on Aura? | logged/auditable? | status |
|---|---|---|---|---|
| Attention Graph (crowd scans) | yes — PERCEPTION att_bonus + Radar aura_score_v1 w_attention=0.40 | no (Aura "Attention" panel ≠ this, see 9C) | yes | wired, but sparse at entry (2/264) |
| Whale / smart-money moves | yes — 3 places (PERCEPTION whale_bonus, Radar w_whale=0.10, Aura Smart Money panel) | yes, live | yes | wired, sparse at entry (1/264) — real elsewhere (Radar/Directory) |
| X/Twitter ticker virality | yes — up to 30pts, live | no dedicated panel | NO — zero persistence, no column exists | wired but unauditable |
| Source-virality (viral seed post) | yes — up to 30pts, live | no dedicated panel | columns exist, but desk API doesn't surface them | wired, logged, currently unpullable read-only |
| KOL-cluster (≥3 tracked KOLs, same token) | NO — zero implementation found | no | n/a | MISSING — spec only (spec_eyezon_kol_cluster_signal) |
| KOL-in-card names ("👑 In: theo, Cupsey…") | display-tied to whale_bonus gate | yes, card only | yes (via whale_count) | wired, real, same sparsity as whale |
| Creator / handle reputation | yes — up to 30pts, live | not surfaced on Aura | NO — no creator_bonus column | wired, evidenced (n=4, p=0.012) but unauditable at scale |
| Bundle / sniper (same-slot buyer) detection | yes — penalty + compound risk gate | Holder X-Ray panel explicitly says "maturing" | yes (bundle_score column) | built + scored, NOT yet shown on the sellable dossier |
| Fresh-wallet / equal-split forensics | yes — penalty, live | Holder X-Ray panel explicitly says "maturing" | no dedicated column (folded into reasons only) | built + scored, NOT yet shown on Aura |
| Dedicated per-wallet "sniper" flag | no — folded into bundle_score only | no | n/a | MISSING — devs' own note: "GMGN sniper-flag integration (planned but not yet built)" |
| Holder concentration (top10_pct) | yes — compound risk gate | yes, live (Holder X-Ray) | yes | flat/wrong-signed (§3C) — a source-quality problem, not a wiring gap |
| Wash-trade / manufactured-volume screen | yes — PERCEPTION hard-reject; Radar fake_vol factor | Volume Authenticity panel shows only basic ratios, not the full wash fingerprint | partially (reasons text only) | built, PARTIALLY shown |
| Smart-money reverse-lookup (proven early wallets, Lever 2) | computed always, only SCORES when early_entry_v2_live=1 | no | attached to internal dict, not DB-persisted per call | shadow-only by design — devs' own backtest (n=13) found no discriminative power yet |
| Behavioral leading score (Lever 1) | same flag as above — shadow only | no | no | shadow-only, same as Lever 2 |
| Dev / insider wallet clustering | no — not a systematized feature | no | n/a | MISSING as a product feature — exists only as a one-off manual dossier (YOLOkol) |
| Graduation ground-truth (pump.fun direct) | yes — hard gate | indirect (via mcap/chart) | yes | healthy — examined in §6 |
| lead_time_sec (speed-to-call) | computed, NOT in composite score | no | yes | real signal (§2/§4B), simply unused — cheapest fix on the list |
x_reader.search_ticker_virality)Wired: yes — feeds virality_bonus directly into the composite (signals_scanner.py:1817), capped at 30pts, weighted 0.8× a shill-bot-filtered velocity score. Has a real, already-fixed bug history (verified-badge signal tested and correctly disproved/removed — bot callers buy X Premium too). Rate-limit behaviour: a proper 3-strike/5-min circuit breaker degrades silently to velocity_score=0 on failure — a genuine X outage and genuine zero chatter are indistinguishable in the score AND in history, because the schema has no virality_bonus column at all — the number is computed, added into conviction, then thrown away. Verdict: real mechanic, structurally impossible to backtest today. This is the single biggest "we can't tell if this actually helps" gap in the whole social layer — fix is one column + one write, zero scoring risk.
source_virality.py)Wired: yes — up to 30pts, log10-scaled seed-post view count (dominant term) + follower-tier + recency decay. Logged: yes — unlike X-virality, this ships with its own columns (source_virality_bonus, source_seed_handle, source_seed_views), confirmed live in perf_tracker.log_signal(). But: the owner-gated /api/perception-desk endpoint — my only read-only production window — deliberately excludes these three columns from its breakdown (comment in the code: "NOT persisted historically: a labelled 5-way score breakdown... only the raw component reads"). Local dev copies of signals_perf.db predate this feature's 2026-07-19 ship date (0/43 rows, all older). I could not pull its live nonzero rate or predictiveness this pass without either a code deploy (out of scope) or direct prod DB access (not available read-only). Recommend as the fastest next step: a 3-field addition to the desk API's breakdown dict (additive, zero scoring risk) — then this becomes fully backtestable next pass.
This is a named priority in memory (spec_eyezon_kol_cluster_signal, explicitly framed as the fix for missing Quantix-style pumps like $URKL) — but a full-repo grep for kol_cluster/KOL_CLUSTER/a cluster-count trigger returns zero implementation anywhere in signals_scanner.py, pre_migration_tracker.py, or whale_db.py. What DOES exist and is live: (1) a single binary is_kol check on the TOKEN'S CREATOR wallet only (not a cluster of buyers), and (2) proven_early_wallets.py — a related-but-different "our own winners' early buyers recur" mechanic, whose own docstring calls out this exact gap by name: "our existing whale_db cohort does NOT overlap real pump.fun ignition (the spec_eyezon_kol_cluster_signal-identified gap)." Verdict: MISSING. The idea is sound and specced; nobody has shipped the ≥3-KOL trigger itself.
Wired: yes, twice — PERCEPTION's att_bonus (store.calls_by_ca, cap 15pts at 3+ scans) and Radar's board-level aura_score_v1 (w_attention=0.40, the single largest weight in that formula, decayed scanner-arrival mass). Coverage at PERCEPTION's entry point: reconfirmed on the FRESH 264-row live pull (grew from 263→264 since the prior report, numbers durable): nonzero on only 2/264 (0.8%). This is not a bug — PERCEPTION fires within minutes of graduation, before a second EyezOn scanner or the crowd graph has had time to arrive. The moat signal is real and IS the largest weight on the Radar board — it's just structurally near-silent at PERCEPTION's very-early entry point specifically, which is a different surface than Radar/Aura where it's genuinely load-bearing.
Wired: yes, up to 30pts, unconditional (runs even on candidates that later fail other gates). Evidence: the single best-documented bonus in the formula — n=4 KOL-creator cases (Elfie/Himothy/Tatesjail/Agamemnon) vs anon, 75% win-rate vs 38.7%, Mann-Whitney p=0.012 — a real, statistically significant result, just thin (n=4). Gap: no creator_bonus column exists in signals_perf.db either — same "computed then discarded" pattern as X-virality — so this genuinely strong early result has never been re-verified at today's n=264. This is the #1 candidate for "add one column, re-run the exact same backtest at 66× the sample size."
Read directly from eyezonaura_preview.html's dsxRenderPanel* functions — the frontend already tags each panel "live" vs "maturing" itself and never fabricates a reading. This is the most honest piece of code in the whole audit; the table below just makes its findings explicit.
| panel | data source | real? | gap vs marketing name |
|---|---|---|---|
| 📈 Attention & Heat | DexScreener tx buckets (5m/1h/6h/24h buy/sell counts) | live | named "Attention" but is generic on-chain tx volume — the code's own comment admits the real Attention Graph (crowd scans) "joins this panel next," i.e. not yet |
| 🐳 Smart Money | whale_db match on this exact CA's real trades | live, honest ("0 whales" shown as a real read) | none — this one does what it says |
| 🔬 Volume Authenticity | DexScreener vol/liq churn + buy/sell ratios | live ratios | basic ratios only — the DEEPER wash-trade fingerprint (per-wallet loop detection) already runs live inside PERCEPTION's own scanner (_detect_wash_trades_sync) but is not surfaced here |
| 🧬 Holder X-Ray | pnl.top_holders (same as whale_db) — concentration, smart-money-in-holders, LP/burn | live for concentration/LP/burn | explicitly labelled "maturing" for per-wallet sniper/fresh-wallet/bundled-buyer forensics — even though fresh_wallet_forensics.py and bundle detection are ALREADY live and scoring inside PERCEPTION. Built in the engine, not yet shipped to the product. |
| 🌀 Fingerprint Read | composite: fake_vol + holder concentration + whale count + scanner count | live where inputs exist | the "scanner count" (crowd) half of its "organic backing" read only populates for tokens already in call_log — a cold Aura lookup on an arbitrary CA usually has this at 0, silently dropping half the read (not fabricated, just often absent) |
The clearest single finding in this section: the fastest, lowest-risk upgrade available to the sellable Aura product is not new data collection — it's wiring fresh_wallet_forensics + _detect_bundling_sync (both already live and running inside PERCEPTION on every candidate) into the Holder X-Ray panel, and the wash-trade fingerprint into Volume Authenticity. Zero new sources, zero new cost — just surfacing what the engine already computes.
signals_scanner.py's composite formula, the Aura client's own live/maturing self-tags, and a fresh 264-row live production pull (up from 263 in the prior pass, confirming §1-8's findings are durable, not a one-time snapshot).Standalone, read-only harness (perception_backtest.py, scratchpad-only) re-pulled the FULL history fresh via /api/perception-desk?tf=all (now 265 calls, up from 263/264 in the audit above — same source, numbers durable) and backtested each proposed change (a)–(g) as a toggleable selection/re-rank rule, solo and combined, with a chronological 70/30 train/test split (first 185 calls vs last 80) to catch overfitting. Nothing in the live engine, DB, or deploy touched — see §11 for the full read-only methodology.
| n | win≥2x | win≥5x | sub‑1.5x | rug% | avg mult | median mult | |
|---|---|---|---|---|---|---|---|
| full population | 265 | 35.8% | 9.8% | 49.8% | 58.9% | 4.72 | 1.50 |
| train (chrono first 70%, n=185) | 185 | 38.4% | 11.4% | 46.5% | 74.1% | 5.25 | 1.55 |
| test (chrono last 30%, n=80) | 80 | 30.0% | 6.2% | 57.5% | 23.8% | 3.50 | 1.24 |
Train/test rug% is NOT comparable directly — test rows are younger (less time to rug/resolve), not evidence a rule changed rug behaviour. All Δrug figures below should be read the same way; Δwin2x/Δwin5x/Δsub1.5x are the load-bearing numbers.
| variant | n (coverage) | win≥2x | Δ win2x | win≥5x | sub‑1.5x | train Δ | test Δ | sign-stable? | read |
|---|---|---|---|---|---|---|---|---|---|
| (a1) buyers 60‑150 only | 31 (11.7%) | 64.5% | +28.7 | 22.6% | 29.0% | +31.2 | +20.0 | yes | strong, robust |
| (a2) exclude buyers≥300 | 140 (52.8%) | 40.7% | +4.9 | 13.6% | 46.4% | +2.8 | +9.5 | yes | modest but keeps HALF the volume |
| (a1) + buyer-count deflated 11.5% (f) | 40 (15.1%) | 60.0% | +24.2 | 20.0% | 35.0% | +21.6 | +30.0 | yes | correction widens the bucket (31→40) but doesn't change the verdict |
| (a2) + deflated (f) | 159 (60.0%) | 39.6% | +3.8 | 13.2% | 47.8% | +3.4 | +4.7 | yes | same — (f) is low-impact on THIS decision |
| (b) fusion: buyers 60‑150 + mcap<$50K | 26 (9.8%) | 69.2% | +33.4 | 26.9% | 23.1% | +31.6 | +36.7 | yes | BEST single lever — high-precision, low-recall |
| (b) + deflated (f) | 29 (10.9%) | 72.4% | +36.6 | 27.6% | 20.7% | +34.3 | +41.4 | yes | marginally better, thin n either way |
| (c) lead_time_sec <290s | 83 (31.3%, logging gap) | 42.2% | +6.4 | 10.8% | 45.8% | +10.4 | +5.0 | yes | real, cheap (already computed, unused today) |
| (e) mcap ceiling ≤$50K | 165 (62.3%) | 37.6% | +1.8 | 13.3% | 49.7% | +2.6 | ‑0.8 | NO | sign flips — do not ship alone |
| (e) mcap ceiling ≤$80K | 231 (87.2%) | 35.5% | ‑0.3 | 10.4% | 50.6% | ‑0.1 | ‑1.0 | yes (flat/negative) | net-zero to slightly negative vs true baseline |
| (e) mcap ceiling ≤$100K | 245 (92.5%) | 35.1% | ‑0.7 | 10.2% | 50.6% | ‑1.2 | +0.1 | NO | no real effect — barely trims the population |
Correction to §5/§7-rec-6 of the audit above: that section compared the sub‑$50K band against the $50‑100K band in isolation (real, still true) — but a mcap-ceiling RULE compared against the TRUE full-population baseline (which also contains strong $100‑150K performers) nets out to roughly flat, and the effect isn't even consistently signed across the train/test split. Recommendation downgraded from "ship" to "do not ship as a standalone ceiling change."
| n | win≥2x | win≥5x | rug% | |
|---|---|---|---|---|
| top10 ≥60% | 19 | 57.9% | 15.8% | 63.2% |
| top10 <60% | 209 | 35.4% | 8.6% | 59.3% |
Direction matches the earlier audit (§3C) — but n=19 is DIRECTIONAL ONLY, not decision-grade alone. Simulating "un-apply the ‑5pt soft penalty" (the code's real formula: conviction −= risk_score×0.25, risk_score=20 for top10≥60% alone) finds only 3 record-only rows in [40,45) that would newly cross the 45 broadcast floor — too thin to read anything into. Structural limit, stated plainly: the live compound gate HARD-REJECTS at risk_score≥70 (needs top10≥60 STACKED with bundle≥55 or thin-liq, not top10 alone) — those candidates are never logged at all, so full removal of the gate can never be backtested from this dataset; only a live shadow/record-only log (Pete's own prescribed method, §4C above) can close this gap.
| n (closed only) | win≥2x | avg mult | |
|---|---|---|---|
| peak-based (ground truth used everywhere above) | 104 | 36.5% | 7.12x |
| held to last-tracked price (real trader outcome) | 104 | 3.8% | 2.39x |
A 32.7-point gap between "what the token did" and "what a trader who didn't sell at the exact top actually captured," on the 104 calls whose tracking window has fully closed (161/265 still tracking, so this will keep filling in — not yet the full picture). This is not a scoring-formula problem — it argues for labelling every future headline number explicitly as "peak-capturable" vs "hold-through," not silently leading with the flattering one. The bigger unlock this points to: a real EXIT signal (when to sell) may matter more to Pete's actual P&L than any entry-side scoring tweak on this list.
| combo | n (coverage) | win≥2x | Δ | win≥5x | train Δ | test Δ | read |
|---|---|---|---|---|---|---|---|
| fusion (b) + fast lead-time (c) | 14 (5.3%) | 57.1% | +21.3 | 28.6% | +17.2 | +30.0 | DIRECTIONAL (n<20) — too thin to trust despite consistent sign |
| sweet-spot (a1) + mcap≤$50K | 26 (9.8%) | 69.2% | +33.4 | 26.9% | +31.6 | +36.7 | identical set to (b) — same rule, confirms (b) is well-formed |
| exclude-crowded (a2) + mcap≤$80K | 134 (50.6%) | 39.6% | +3.8 | 13.4% | +1.4 | +9.0 | best VOLUME-preserving combo — half the calls, modest but real lift, sign-stable |
| fusion (b) + fast (c) + mcap≤$50K | 14 (5.3%) | 57.1% | +21.3 | 28.6% | +17.2 | +30.0 | same n=14 set as row 1 — mcap≤50K adds nothing extra here (already implied by fusion) |
perception_backtest.py (scratchpad) · @sudiyasa_'s desk only