EyezOn Data · Internal, Owner-Only

PERCEPTION Signal Matrix — Deep Audit

Full-history pull off production signals_perf.db (263 calls, 2026‑07‑16 → 2026‑07‑21, read-only via the owner-gated /api/perception-desk API). Nothing in the live engine touched. This is source material for a backtest — not applied changes.

Verdict: Pete is right. The score barely predicts outcome, and record‑only "WEAK" calls the matrix silently drops perform statistically the SAME as (or better than) the calls it actually broadcasts — including the single best call in PERCEPTION's history.
263
total calls
36%
win rate (≥2x)
50%
sub‑1.5x tail
59%
rugged

63% (165/263) are still tracking — these numbers can only get BETTER as peaks are still forming, never worse. Every stat below is honest against real stored outcomes as of pull time.

1 · Baseline — what PERCEPTION actually produces

ThresholdCount% of 263
10x+114.2%
5x+269.9%
3x+5119.4%
2x+ ("win")9536.1%
1.5x+13149.8%
<1.5x (tail)13250.2%
rugged (eventually)15659.3%
Data-integrity flag surfaced honestly: the desk API's own crosscheck shows avg_multiple disagreeing between PERCEPTION's own DB (1.84x) and the public leaderboard's per-caller stat (5.26x) for the same 263ish rows — likely a computation-window or dedup difference between the two code paths, not a live bug I'm touching. Flagging so it gets reconciled; I used the raw per-row DB data throughout this report, not either pre-aggregated number.
Rug ≠ loss. 43 of the 156 rugged tokens (27.6%) peaked ≥2x before rugging — 8 peaked ≥5x. Top examples: TABLEPEDE 35.3x, BULLZARD 22.6x, GHOSTI 11.8x, Tatesjail 11.3x, PTBL 9.3x, TRACCOON 8.8x. The current "rugged" outcome bucket hides these as pure losses. If win-rate is reported off peak‑multiple (what a trader could actually have captured) rather than "survived," several headline numbers above are understated.

2 · Score vs outcome — the core misalignment, quantified

Correlation of eyezon_score to log(peak_multiple): r = ‑0.06 — statistically flat, if anything the wrong sign. The score, as currently weighted, carries almost no information about how big a call will run.

By tier

Tiernavg multmedian≥2x%≥5x%rug%
STRONG (≥78)14.944.94100%0%0%
SOLID (62‑77)162.451.8843.8%6.2%75%
OK (45‑61)382.331.5036.8%10.5%68.4%
WEAK (<45, record-only, never posted)1793.121.5237.4%9.5%59.2%

WEAK — the tier that structurally can never broadcast — has a higher average multiple and materially lower rug rate than SOLID. This isn't noise at the edges; it's the largest bucket in the dataset (179/263, 68%).

Broadcast floor test (the actual gate Pete's channel experiences)

navg multmedian≥1.5x%≥2x%≥5x%rug%
BROADCAST (score≥45, what gets posted)552.411.5154.5%40.0%9.1%69.1%
RECORD-ONLY (score<45, logged but never posted)2085.361.4748.6%35.1%10.1%56.7%

The 45-point broadcast floor is filtering ~79% of PERCEPTION's calls (208/263) for a hit-rate that is statistically indistinguishable from — and on this sample slightly worse than — what it silently discards. The all-time best call, $Jimothy at 412.7x, scored 63/tier=null and would still have cleared today's 45 floor, but dozens of other real winners sit below it: the WEAK bucket alone contains 9.5% of its population hitting 5x+.

Feature predictiveness (5x+ winners vs sub‑1.5x duds)

featurewinners(5x+) medduds(<1.5x) medcorr w/ log(mult)read
authenticity (62% of score weight)8078‑0.037flat/noise
buyers229287‑0.149inverted — see §3
sellers101166‑0.160inverted
buy_pct (dominance)64.161.5weak alone, strong combined (§4)
liquidity ($)14,62814,618flat
whale_count00dead — see §3
crowd_count00dead — see §3
top10_pct36.539.4+0.028flat, wrong-signed vs the penalty
bundle_score5555‑0.078weak alone
liq_mcap_ratio0.340.33‑0.038flat alone
lead_time_sec (NOT scored today)254311‑0.152real, unused

3 · The three concrete "why" findings behind the misalignment

A) Authenticity is gated and weighted 62% — so it's double-counted and low-information

auth_min already hard-rejects the worst readings before evaluate() ever computes a conviction score. Among survivors, authenticity clusters tight (72‑80 across both winners and duds) — the single heaviest-weighted term in the formula (0.62×) carries almost no residual separating power on the population that actually reaches scoring. Most of the real information is riding on the other 38% of the weight — several legs of which are dead (whale/crowd) or wrong-signed (top10 penalty).

B) Buyer-count is scored as "more = better" (saturating at 60) — the data says it's U‑shaped

buyers at entrynavg mult≥2x%
0‑60226.4027.3%
60‑150 (sweet spot)315.6864.5%
150‑300863.2036.0%
300+ (crowded)1122.0733.0%

60‑150 buyers at entry is the best bucket in the whole dataset (64.5% hit 2x+, more than 3x the win-rate of the 300+ "most active-looking" bucket). The current authenticity sub-term (_lin(buyers,8,60)) treats 60 buyers and 1,000 buyers identically — it never penalises an entry that's already crowded/late, which is exactly the failure mode a "quiet real winner" (Pete's flag) falls into: fewer buyers at entry reads as lower authenticity today when it's often the more room-to-run entry.

C) top10_pct concentration is scored as risk — on this data it's the opposite

top10 concentrationnavg mult≥2x%
0‑20%21.310%
20‑40%1183.5937.3%
40‑60%872.0134.5%
60%+ (flagged as risk, +20 to compound gate)193.5657.9%

Consistent with the codebase's own prior notes ($Himgajria at 87.5% concentration was a real winner) — concentration alone is not the risk signal the compound gate treats it as. See §4C for the more useful, interaction-based version of this finding: concentration only turns genuinely bad when it stacks with bundle evidence.

4 · Untapped synergies — signals that are weak alone, strong combined

Pete's ask: what data points don't currently talk to each other but would help if they did. Tested directly against the 263-call history wherever the fields exist; flagged as hypothesis where they don't.

A) Low mcap × moderate (non-crowded) buyer-count — the strongest fusion found

combonavg mult≥2x%≥5x%
buyers 60‑150 + entry mcap <$50K266.4569.2%26.9%
buyers 60‑150 + mcap $50‑100K51.6840.0%0%
buyers 300+ (crowded) + mcap <$50K571.8826.3%7.0%

Neither feature is this strong alone. Cheap mcap alone gives 5x+ at 13.5% (§5); the buyer sweet-spot alone gives 2x+ at 64.5%. Together (n=26): 69.2% hit 2x+ and 26.9% hit 5x+ — roughly 2‑3x the dataset baseline. Low mcap alone is not sufficient — the same low-mcap population with crowded (300+) buyers is actually below-baseline (26.3%). Low mcap only pays off when it's still quiet.

B) Speed × crowding — the "fast AND still quiet" double-confirmation

combonavg mult≥2x%≥5x%
fast call (<290s) + sweet-spot buyers167.6556.2%25.0%
fast call + crowded buyers (300+)352.1940.0%5.7%
slow call (≥290s), all buyer counts811.9428.4%7.4%

lead_time_sec (currently unscored — §3/§6) and buyer-count are individually modest (r=‑0.15 and non-linear respectively) but the intersection is one of the cleanest signals in the whole dataset. "We caught it fast AND the crowd hasn't arrived yet" is a genuinely different, stronger claim than either alone.

C) The compound risk gate's own premise, tested directly

risk-leg combinationnavg mult≥2x%≥5x%rug%
BOTH top10≥60% + bundle≥35 elevated46.0375.0%50.0%100%
only top10≥60%152.9053.3%6.7%53.3%
only bundle≥351622.8834.0%8.6%58.6%
neither elevated453.0042.2%8.9%64.4%

n=4 is too thin to act on, but directionally interesting and worth a flag: when BOTH risk legs stack, the survivors that clear the hard-reject threshold anyway look like a distinct extreme-variance regime — big pumps AND a 100% eventual rug, i.e. "real pump, guaranteed rug" rather than "just bad." That's a genuinely different trading instruction (fast in/fast out) than what a single elevated leg implies. Worth deliberately growing this bucket's sample via a shadow (record-only, non-broadcast) log before touching the live gate.

D) Buy-dominance × liquidity depth

combonavg mult≥2x%≥5x%
buy-dominance ≥65% + liq/mcap ≥30% (healthy)434.4148.8%20.9%
buy-dominance ≥65% + liq/mcap <15% (thin)21.310%0%

Both features are individually near-zero-correlated with outcome (§2). Combined, "heavy one-sided accumulation with real liquidity behind it" clears baseline meaningfully (48.8% vs 36.1% overall). The thin-liquidity leg is too small (n=2) to trust but is directionally exactly what you'd expect — heavy buying against thin liquidity is a fragile, easily-rugged spike, not organic accumulation.

E) Authenticity only means something when paired with crowding state

combonavg mult≥2x%
authenticity≥80 + sweet-spot buyers256.2268.0%
authenticity≥80 + crowded buyers (300+)322.1431.2%

The single highest-leverage fix in this whole report: the same high authenticity score (≥80) means a 68% win-rate in one crowding state and a 31% win-rate in another — more than 2x apart on the SAME feature. Since authenticity is 62% of conviction, this interaction (not currently modelled at all) is likely costing more accuracy than any other single item here.

F) Signals that are too siloed in this dataset to test — need logging first (honest gaps)

5 · Entry-mcap band (re-run at 3x the old sample size)

entry mcapnavg mult≥2x%≥5x%
<$50K1636.4338.0%13.5%
$50‑100K801.9130.0%3.8%
$100‑150K152.4746.7%6.7%
$150‑300K52.0340.0%0%

The code's own prior advisory (compute_optimal_mcap_band) flagged this direction on n=6/n=1 buckets — too thin to trust. At n=163/n=80 the same direction now holds with real weight: sub‑$50K is unambiguously the strongest band (3.5x the 5x+ rate of $50‑100K), and $50‑100K is the clear laggard.

6 · Source-by-source audit — what feeds the matrix, and how healthy is each

Every source the scoring path (signals_scanner.evaluate()) actually touches, read directly from the code + its own inline incident notes — not guessed.

healthy

Discovery — Iris (own chain-watch, primary)

Free-RPC watch of the pump.fun migration authority, near-zero lag, self-owned — no third-party dependency. Live 10-min head-to-head: GeckoTerminal's new_pools (the old primary) surfaced 0 of 5 real graduations against a true ~25/hr chain-wide rate; GT is now demoted to a fallback/reconciliation pass only. This is the single biggest documented source upgrade in the codebase and it's already shipped — good.

limited / known bug, unfixed

Market data — GeckoTerminal (mcap / liquidity / volume / tx counts)

Primary feed for cand.mcap/liq/vol_h1/chg_h1 and the buyers/sellers counts that drive authenticity — the single heaviest-weighted score component (62%). Two documented, currently-unfixed issues: (1) rate-limit/429 storms serious enough to need a dedicated pacer/breaker module (gt_pacer.py); (2) GT's buyers/sellers counts double-count wallets that both buy AND sell in the window ("flips") — live-verified an ~11.5% inflation on a real sample, feeding directly into the unique-wallet-ratio term of authenticity. Documented in the code as a known, accepted gap (no cheap fix available without a second RPC-level walk). This is the clearest case in the whole audit of a slightly-noisy input sitting underneath the most heavily-weighted output.

healthy

Graduation ground-truth — pump.fun's own API

Direct-source (complete + pump_swap_pool), not an aggregator's inference. GT launchpad_details as fallback only. No documented issues.

healthy

Cross-check — DexScreener (mcap agreement + liquidity floor)

Purpose-built after a real incident ($Yuna: entry landed inside an unstable migration-snipe spike). GT-vs-DS agreement gate + MIN(GT,DS) liquidity is good defense-in-depth. Only gap: the agreement ratio itself is discarded after the pass/fail check (§4F) — a graded confidence feature is sitting right there, unused.

buggy / documented data corruption

Safety — RugCheck API (top10%, LP-locked, mint/freeze authority)

Feeds the compound risk gate directly (§3C's top10_pct penalty). The codebase has ALREADY had to add a hard sanity guard (_top_holders_sane) after live-confirmed corrupt reads: topHolders summing to 114.4% ($ANSEMIO), the null system-address listed 5x as a "holder" ($PUMPGU). Also has real coverage gaps — HTTP 400/empty on some brand-new mints (fell back silently to on-chain pnl.top_holders for $AnsemAI). This is the one source where a documented data-quality problem and a documented scoring-accuracy problem (top10_pct is flat/wrong-signed in §3C) plausibly compound each other.

healthy, self-owned, still young

Bundle/sniper detection — own free-RPC walk

No third-party cost or rate-limit exposure. Two real false-positive bugs already found and fixed live (raw-tx-count vs distinct-wallet-count on $Himgajria; a noisy ==2 sub-case removed after the Agamemnon/RACCTARD backtest) — a good sign of active calibration, but also a sign it's not yet mature. Only computed for candidates that already cleared every cheaper gate (86% coverage, 226/263) — by design, not a gap.

limited by design (short window)

Wash-trade screen — GeckoTerminal pool trades (last ~300 only)

No pagination/history on GT's free trades endpoint — the original case-study wash-burst ($ACT:S) had already rolled off the feed by the time the fix was validated, so thresholds were calibrated from a fresh live re-pull, not the actual incident data. Structurally can only catch a wash campaign that's still active at scan time, not one that already finished.

healthy, appropriately humble

Fresh-wallet / equal-split forensics — own on-chain RPC

Was a hard-reject, downgraded to a penalty after live evidence it filtered real 5‑9x winners ($Sagawa class). Good example of the codebase correcting itself against real outcomes — the standard the rest of the matrix should be held to.

technically correct, practically dead for PERCEPTION

Whale presence & crowd attention — own DBs (whale_db 1,844 wallets · store.py scan log)

Nonzero on 1/251 and 2/251 calls respectively. Not a data bug — a tracked whale or a second EyezOn scanner genuinely hasn't had time to arrive 3‑10 minutes post-migration, which is exactly when PERCEPTION fires by design. The score formula still reserves a flat +15/+15 bonus slot for each that essentially never pays out. These fusions are real and validated on OTHER surfaces (Radar/Aura) — just not at PERCEPTION's entry point.

healthy, validated (small n)

Creator reputation — whale_db lookup on the pump.fun creator wallet

One of the only bonus terms with an actual documented backtest behind it: n=4 KOL-creator cases, 75% win-rate (≥2x) vs 38.7% for anon creators, p=0.012, zero rugs among the KOL cases at the time. Genuinely the strongest-evidenced single bonus in the formula — also the thinnest (n=4) and not currently re-testable from this pull's data (component not broken out per call, see §6-logging gap below).

plausible, currently un-auditable

X-virality & source-virality — x_reader.py / source_virality.py (burner-account scrape + FxTwitter)

Both have 3-strike/5-min circuit breakers — under X rate-limiting they silently degrade to zero contribution with no flag left in the stored row. A call scored during an X outage is indistinguishable, in history, from a call that genuinely had zero chatter. One real bug already found and fixed: raw X "verified" badge was tested and disproved as a trust signal (bot/caller accounts buy verification too) and correctly removed. Not separately logged per call — can't backtest its real contribution today.

out of scope this pass, but the ground truth for everything above

Outcome pricing — candle/OHLCV cascade (data.py, candle_store.py, Birdeye/GT)

Not part of the entry score, but it IS what computes peak_mcap/peak_multiple — the literal ground truth this entire report is built on. This is its own active workstream (the EyezOn Data indexer/persistent-candle-store build) and wasn't deep-audited again here to avoid duplicating that effort. Worth a standing reminder: any gap in this feed silently propagates into every number above it.

7 · Prioritized recommendations — proposals to backtest, not changes applied

1 Report win-rate off peak multiple, not eventual survival
Evidence: 43/156 rugged tokens (27.6%) peaked ≥2x, 8 peaked ≥5x (up to 35x) before rugging.
Impact: corrects the headline win-rate/2x+ number Pete already suspects is too low — likely a reporting fix more than a matrix fix.
Caveat: doesn't change what actually happens to a real trader's bag if they hold through the rug — needs a clear "peak-capturable win-rate" vs "hold-to-now win-rate" split, not a silent swap.
2 Re-test lowering (or restructuring) the 45-point broadcast floor
Evidence: record-only WEAK calls (208/263) hit 2x+ at 35.1% vs 40.0% for broadcast calls, 5x+ at 10.1% vs 9.1%, rug rate 56.7% vs 69.1% — statistically indistinguishable to slightly BETTER, and includes the all-time best call (Jimothy, 412x, scored 63/tier=null).
Impact: potentially ~3-4x more postable call volume at equal or better quality, on this sample.
Caveat: n=55 broadcast is thin; backtest properly (not just re-run this exact query) before touching signals_min_conviction live — and consider pairing with fix #4/#5 below FIRST, since they're likely why the floor currently has to sit where it does.
3 Make buyer-count non-monotonic (U-shaped, not saturating) in the authenticity formula
Evidence: buyers 60‑150 → 64.5% hit 2x+ (best bucket); buyers 300+ → 33.0% (worst non-thin bucket). Interaction with authenticity (§4E) shows the SAME auth≥80 score means 68% vs 31% win-rate depending on crowding.
Impact: likely the single highest-leverage fix in this report — directly targets Pete's "quiet real winners scoring low" complaint. Zero extra data cost (buyers already fetched).
Caveat: buyer-count naturally rises with token age — control for age/lead-time before concluding the effect is purely about crowding vs just "later snapshot."
4 Add lead_time_sec into the score — it's logged but currently NOT used at all
Evidence: fast calls (<290s) hit 2x+ at 42.2% vs 28.4% for slow (≥290s); fast+sweet-spot-buyers combo hits 56.2%/25.0% (2x+/5x+).
Impact: real, already-computed, currently wasted signal — cheapest fix to ship of the five.
Caveat: only populated on 62% of rows (164/263) — confirm why the other 38% are missing it before wiring it into scoring (could be a logging gap, not a "no lead time" case).
5 Drop or heavily discount the top10_pct≥60% leg of the compound risk gate
Evidence: top10≥60% bucket has the BEST 2x+ rate of any concentration bucket (57.9%, n=19) — opposite of the risk-gate's assumption, consistent with the code's own $Himgajria prior note (87.5% concentration, real winner).
Impact: stops penalizing a feature that isn't actually bad on this data; keep the bundle_score leg (weak but at least correctly-signed) and the thin-liquidity leg.
Caveat: RugCheck (the source for this feature) has documented corruption incidents (§6) — worth also asking whether top10_pct itself is trustworthy enough to score at all before deciding how to weight it.
6 Tighten signals_scanner_max_mcap toward the $50‑80K range
Evidence: sub-$50K entries (163/263, 62% of all calls) hit 5x+ at 13.5% vs 3.8% for $50‑100K (n=80) — same direction as the code's own old advisory, now on 15‑25x the sample size.
Impact: cuts exposure to the demonstrably weakest band without touching the strongest one.
Caveat: $100‑150K (n=15) actually shows a decent 46.7% 2x+ rate — don't cut the ceiling further than $50‑100K without more data there; sample still thin above $100K.
7 Log the score sub-components per call, not just the final composite
Evidence: momentum, virality_bonus, source_virality_bonus, creator_bonus, and v2_bonus are all computed live but discarded after the composite conviction number is calculated — several findings in §4/§6 (creator-reputation re-test, virality×crowding, degraded-source flagging) are blocked purely by this gap.
Impact: enables every future backtest pass to be sharper — this is the "we should log X" finding the brief specifically asked to surface honestly.
Caveat: pure logging addition, zero scoring-behavior risk — lowest-risk item on this list, but needs a schema change (not touched this pass, per the read-only constraint).

8 · Honest caveats

9 · Completeness & Social Audit — every claimed data point, checked

Extends §1-8 above (unchanged, still stands) with an exhaustive pass: every data point EyezOn's marketing/Aura dossier/memory specs claim to use, cross-checked against signals_scanner.py (PERCEPTION's live composite), site/wsgi.py's Aura dossier code (eyezonaura_preview.html's dsxRenderPanel* functions — which self-report live/maturing honestly in their own client code), and a fresh live pull of /api/perception-desk (264 calls, re-confirms §1-8's numbers still hold at n+1). Nothing deployed; nothing in the live engine touched.

Verdict: the ENGINE is more complete than the PRODUCT shows. Every data point EyezOn actually claims as its moat is genuinely computed somewhere in the code — but three of the five Aura dossier panels display less than PERCEPTION's own scanner already knows, and two real-time social signals (X-virality, creator-reputation) are wired into the live score yet ZERO-persisted, making them structurally unauditable. One flagship claim — the KOL-cluster trigger — is spec-only, never built.

9A — Full claim vs reality table

claimed data pointwired into a live score?shown on Aura?logged/auditable?status
Attention Graph (crowd scans)yes — PERCEPTION att_bonus + Radar aura_score_v1 w_attention=0.40no (Aura "Attention" panel ≠ this, see 9C)yeswired, but sparse at entry (2/264)
Whale / smart-money movesyes — 3 places (PERCEPTION whale_bonus, Radar w_whale=0.10, Aura Smart Money panel)yes, liveyeswired, sparse at entry (1/264) — real elsewhere (Radar/Directory)
X/Twitter ticker viralityyes — up to 30pts, liveno dedicated panelNO — zero persistence, no column existswired but unauditable
Source-virality (viral seed post)yes — up to 30pts, liveno dedicated panelcolumns exist, but desk API doesn't surface themwired, logged, currently unpullable read-only
KOL-cluster (≥3 tracked KOLs, same token)NO — zero implementation foundnon/aMISSING — spec only (spec_eyezon_kol_cluster_signal)
KOL-in-card names ("👑 In: theo, Cupsey…")display-tied to whale_bonus gateyes, card onlyyes (via whale_count)wired, real, same sparsity as whale
Creator / handle reputationyes — up to 30pts, livenot surfaced on AuraNO — no creator_bonus columnwired, evidenced (n=4, p=0.012) but unauditable at scale
Bundle / sniper (same-slot buyer) detectionyes — penalty + compound risk gateHolder X-Ray panel explicitly says "maturing"yes (bundle_score column)built + scored, NOT yet shown on the sellable dossier
Fresh-wallet / equal-split forensicsyes — penalty, liveHolder X-Ray panel explicitly says "maturing"no dedicated column (folded into reasons only)built + scored, NOT yet shown on Aura
Dedicated per-wallet "sniper" flagno — folded into bundle_score onlynon/aMISSING — devs' own note: "GMGN sniper-flag integration (planned but not yet built)"
Holder concentration (top10_pct)yes — compound risk gateyes, live (Holder X-Ray)yesflat/wrong-signed (§3C) — a source-quality problem, not a wiring gap
Wash-trade / manufactured-volume screenyes — PERCEPTION hard-reject; Radar fake_vol factorVolume Authenticity panel shows only basic ratios, not the full wash fingerprintpartially (reasons text only)built, PARTIALLY shown
Smart-money reverse-lookup (proven early wallets, Lever 2)computed always, only SCORES when early_entry_v2_live=1noattached to internal dict, not DB-persisted per callshadow-only by design — devs' own backtest (n=13) found no discriminative power yet
Behavioral leading score (Lever 1)same flag as above — shadow onlynonoshadow-only, same as Lever 2
Dev / insider wallet clusteringno — not a systematized featurenon/aMISSING as a product feature — exists only as a one-off manual dossier (YOLOkol)
Graduation ground-truth (pump.fun direct)yes — hard gateindirect (via mcap/chart)yeshealthy — examined in §6
lead_time_sec (speed-to-call)computed, NOT in composite scorenoyesreal signal (§2/§4B), simply unused — cheapest fix on the list

9B — Social metrics, individually, the honest read (Pete's specific ask)

wired, live-only, unauditable

X/Twitter ticker virality (x_reader.search_ticker_virality)

Wired: yes — feeds virality_bonus directly into the composite (signals_scanner.py:1817), capped at 30pts, weighted 0.8× a shill-bot-filtered velocity score. Has a real, already-fixed bug history (verified-badge signal tested and correctly disproved/removed — bot callers buy X Premium too). Rate-limit behaviour: a proper 3-strike/5-min circuit breaker degrades silently to velocity_score=0 on failure — a genuine X outage and genuine zero chatter are indistinguishable in the score AND in history, because the schema has no virality_bonus column at all — the number is computed, added into conviction, then thrown away. Verdict: real mechanic, structurally impossible to backtest today. This is the single biggest "we can't tell if this actually helps" gap in the whole social layer — fix is one column + one write, zero scoring risk.

wired + logged, but unpullable this pass

Source-virality — viral seed post reach (source_virality.py)

Wired: yes — up to 30pts, log10-scaled seed-post view count (dominant term) + follower-tier + recency decay. Logged: yes — unlike X-virality, this ships with its own columns (source_virality_bonus, source_seed_handle, source_seed_views), confirmed live in perf_tracker.log_signal(). But: the owner-gated /api/perception-desk endpoint — my only read-only production window — deliberately excludes these three columns from its breakdown (comment in the code: "NOT persisted historically: a labelled 5-way score breakdown... only the raw component reads"). Local dev copies of signals_perf.db predate this feature's 2026-07-19 ship date (0/43 rows, all older). I could not pull its live nonzero rate or predictiveness this pass without either a code deploy (out of scope) or direct prod DB access (not available read-only). Recommend as the fastest next step: a 3-field addition to the desk API's breakdown dict (additive, zero scoring risk) — then this becomes fully backtestable next pass.

claimed, not built

KOL-cluster signal (≥3 tracked KOLs buying the same fresh token)

This is a named priority in memory (spec_eyezon_kol_cluster_signal, explicitly framed as the fix for missing Quantix-style pumps like $URKL) — but a full-repo grep for kol_cluster/KOL_CLUSTER/a cluster-count trigger returns zero implementation anywhere in signals_scanner.py, pre_migration_tracker.py, or whale_db.py. What DOES exist and is live: (1) a single binary is_kol check on the TOKEN'S CREATOR wallet only (not a cluster of buyers), and (2) proven_early_wallets.py — a related-but-different "our own winners' early buyers recur" mechanic, whose own docstring calls out this exact gap by name: "our existing whale_db cohort does NOT overlap real pump.fun ignition (the spec_eyezon_kol_cluster_signal-identified gap)." Verdict: MISSING. The idea is sound and specced; nobody has shipped the ≥3-KOL trigger itself.

wired, real, structurally rare

Telegram/crowd attention — "N eyes on this" (EyezOn's own scanner count)

Wired: yes, twice — PERCEPTION's att_bonus (store.calls_by_ca, cap 15pts at 3+ scans) and Radar's board-level aura_score_v1 (w_attention=0.40, the single largest weight in that formula, decayed scanner-arrival mass). Coverage at PERCEPTION's entry point: reconfirmed on the FRESH 264-row live pull (grew from 263→264 since the prior report, numbers durable): nonzero on only 2/264 (0.8%). This is not a bug — PERCEPTION fires within minutes of graduation, before a second EyezOn scanner or the crowd graph has had time to arrive. The moat signal is real and IS the largest weight on the Radar board — it's just structurally near-silent at PERCEPTION's very-early entry point specifically, which is a different surface than Radar/Aura where it's genuinely load-bearing.

wired, real, evidenced — but the weakest-audited of the "strong" signals

Creator / handle reputation

Wired: yes, up to 30pts, unconditional (runs even on candidates that later fail other gates). Evidence: the single best-documented bonus in the formula — n=4 KOL-creator cases (Elfie/Himothy/Tatesjail/Agamemnon) vs anon, 75% win-rate vs 38.7%, Mann-Whitney p=0.012 — a real, statistically significant result, just thin (n=4). Gap: no creator_bonus column exists in signals_perf.db either — same "computed then discarded" pattern as X-virality — so this genuinely strong early result has never been re-verified at today's n=264. This is the #1 candidate for "add one column, re-run the exact same backtest at 66× the sample size."

9C — Aura dossier, panel by panel (the client code's OWN honest self-report)

Read directly from eyezonaura_preview.html's dsxRenderPanel* functions — the frontend already tags each panel "live" vs "maturing" itself and never fabricates a reading. This is the most honest piece of code in the whole audit; the table below just makes its findings explicit.

paneldata sourcereal?gap vs marketing name
📈 Attention & HeatDexScreener tx buckets (5m/1h/6h/24h buy/sell counts)livenamed "Attention" but is generic on-chain tx volume — the code's own comment admits the real Attention Graph (crowd scans) "joins this panel next," i.e. not yet
🐳 Smart Moneywhale_db match on this exact CA's real tradeslive, honest ("0 whales" shown as a real read)none — this one does what it says
🔬 Volume AuthenticityDexScreener vol/liq churn + buy/sell ratioslive ratiosbasic ratios only — the DEEPER wash-trade fingerprint (per-wallet loop detection) already runs live inside PERCEPTION's own scanner (_detect_wash_trades_sync) but is not surfaced here
🧬 Holder X-Raypnl.top_holders (same as whale_db) — concentration, smart-money-in-holders, LP/burnlive for concentration/LP/burnexplicitly labelled "maturing" for per-wallet sniper/fresh-wallet/bundled-buyer forensics — even though fresh_wallet_forensics.py and bundle detection are ALREADY live and scoring inside PERCEPTION. Built in the engine, not yet shipped to the product.
🌀 Fingerprint Readcomposite: fake_vol + holder concentration + whale count + scanner countlive where inputs existthe "scanner count" (crowd) half of its "organic backing" read only populates for tokens already in call_log — a cold Aura lookup on an arbitrary CA usually has this at 0, silently dropping half the read (not fabricated, just often absent)

The clearest single finding in this section: the fastest, lowest-risk upgrade available to the sellable Aura product is not new data collection — it's wiring fresh_wallet_forensics + _detect_bundling_sync (both already live and running inside PERCEPTION on every candidate) into the Holder X-Ray panel, and the wash-trade fingerprint into Volume Authenticity. Zero new sources, zero new cost — just surfacing what the engine already computes.

9D — Completeness statement

10 · Backtest Harness — evidence for every proposed change, on 265 calls

Standalone, read-only harness (perception_backtest.py, scratchpad-only) re-pulled the FULL history fresh via /api/perception-desk?tf=all (now 265 calls, up from 263/264 in the audit above — same source, numbers durable) and backtested each proposed change (a)–(g) as a toggleable selection/re-rank rule, solo and combined, with a chronological 70/30 train/test split (first 185 calls vs last 80) to catch overfitting. Nothing in the live engine, DB, or deploy touched — see §11 for the full read-only methodology.

Verdict: two changes are strong AND robust — the low-mcap × moderate-buyers fusion trigger (b) and, more cautiously, excluding crowded (300+) buyer entries (a2). One widely-assumed fix — tightening the mcap ceiling (e) — does NOT hold up once tested against the true full-population baseline instead of a single weak band; its sign flips between train and test. The single most important finding isn't a scoring change at all: peak-based win-rate overstates what a real hold captures by ~33 points.

10A — Baseline (fresh pull, 265 calls)

nwin≥2xwin≥5xsub‑1.5xrug%avg multmedian mult
full population26535.8%9.8%49.8%58.9%4.721.50
train (chrono first 70%, n=185)18538.4%11.4%46.5%74.1%5.251.55
test (chrono last 30%, n=80)8030.0%6.2%57.5%23.8%3.501.24

Train/test rug% is NOT comparable directly — test rows are younger (less time to rug/resolve), not evidence a rule changed rug behaviour. All Δrug figures below should be read the same way; Δwin2x/Δwin5x/Δsub1.5x are the load-bearing numbers.

10B — Each variant, solo, full population vs train vs test

variantn (coverage)win≥2xΔ win2xwin≥5xsub‑1.5xtrain Δtest Δsign-stable?read
(a1) buyers 60‑150 only31 (11.7%)64.5%+28.722.6%29.0%+31.2+20.0yesstrong, robust
(a2) exclude buyers≥300140 (52.8%)40.7%+4.913.6%46.4%+2.8+9.5yesmodest but keeps HALF the volume
(a1) + buyer-count deflated 11.5% (f)40 (15.1%)60.0%+24.220.0%35.0%+21.6+30.0yescorrection widens the bucket (31→40) but doesn't change the verdict
(a2) + deflated (f)159 (60.0%)39.6%+3.813.2%47.8%+3.4+4.7yessame — (f) is low-impact on THIS decision
(b) fusion: buyers 60‑150 + mcap<$50K26 (9.8%)69.2%+33.426.9%23.1%+31.6+36.7yesBEST single lever — high-precision, low-recall
(b) + deflated (f)29 (10.9%)72.4%+36.627.6%20.7%+34.3+41.4yesmarginally better, thin n either way
(c) lead_time_sec <290s83 (31.3%, logging gap)42.2%+6.410.8%45.8%+10.4+5.0yesreal, cheap (already computed, unused today)
(e) mcap ceiling ≤$50K165 (62.3%)37.6%+1.813.3%49.7%+2.6‑0.8NOsign flips — do not ship alone
(e) mcap ceiling ≤$80K231 (87.2%)35.5%‑0.310.4%50.6%‑0.1‑1.0yes (flat/negative)net-zero to slightly negative vs true baseline
(e) mcap ceiling ≤$100K245 (92.5%)35.1%‑0.710.2%50.6%‑1.2+0.1NOno real effect — barely trims the population

Correction to §5/§7-rec-6 of the audit above: that section compared the sub‑$50K band against the $50‑100K band in isolation (real, still true) — but a mcap-ceiling RULE compared against the TRUE full-population baseline (which also contains strong $100‑150K performers) nets out to roughly flat, and the effect isn't even consistently signed across the train/test split. Recommendation downgraded from "ship" to "do not ship as a standalone ceiling change."

10C — (d) top10 concentration penalty — directional only, honest coverage gap

nwin≥2xwin≥5xrug%
top10 ≥60%1957.9%15.8%63.2%
top10 <60%20935.4%8.6%59.3%

Direction matches the earlier audit (§3C) — but n=19 is DIRECTIONAL ONLY, not decision-grade alone. Simulating "un-apply the ‑5pt soft penalty" (the code's real formula: conviction −= risk_score×0.25, risk_score=20 for top10≥60% alone) finds only 3 record-only rows in [40,45) that would newly cross the 45 broadcast floor — too thin to read anything into. Structural limit, stated plainly: the live compound gate HARD-REJECTS at risk_score≥70 (needs top10≥60 STACKED with bundle≥55 or thin-liq, not top10 alone) — those candidates are never logged at all, so full removal of the gate can never be backtested from this dataset; only a live shadow/record-only log (Pete's own prescribed method, §4C above) can close this gap.

10D — (g) Peak vs. hold — the highest-leverage finding in this pass

n (closed only)win≥2xavg mult
peak-based (ground truth used everywhere above)10436.5%7.12x
held to last-tracked price (real trader outcome)1043.8%2.39x

A 32.7-point gap between "what the token did" and "what a trader who didn't sell at the exact top actually captured," on the 104 calls whose tracking window has fully closed (161/265 still tracking, so this will keep filling in — not yet the full picture). This is not a scoring-formula problem — it argues for labelling every future headline number explicitly as "peak-capturable" vs "hold-through," not silently leading with the flattering one. The bigger unlock this points to: a real EXIT signal (when to sell) may matter more to Pete's actual P&L than any entry-side scoring tweak on this list.

10E — Combinations

combon (coverage)win≥2xΔwin≥5xtrain Δtest Δread
fusion (b) + fast lead-time (c)14 (5.3%)57.1%+21.328.6%+17.2+30.0DIRECTIONAL (n<20) — too thin to trust despite consistent sign
sweet-spot (a1) + mcap≤$50K26 (9.8%)69.2%+33.426.9%+31.6+36.7identical set to (b) — same rule, confirms (b) is well-formed
exclude-crowded (a2) + mcap≤$80K134 (50.6%)39.6%+3.813.4%+1.4+9.0best VOLUME-preserving combo — half the calls, modest but real lift, sign-stable
fusion (b) + fast (c) + mcap≤$50K14 (5.3%)57.1%+21.328.6%+17.2+30.0same n=14 set as row 1 — mcap≤50K adds nothing extra here (already implied by fusion)

10F — Recommended combined config (ranked by robustness × impact)

1 Ship a "high-conviction fusion" tag/boost: buyers 60‑150 AND entry mcap<$50K
n=26 (9.8% of population), +33.4pt win2x, +17.1pt win5x, sign-consistent train/test, identical result whether framed as fusion (b) or combo (a1+mcap).
Highest-leverage, most robust single change on this list. High-precision/low-recall by nature — use as an ADDITIVE conviction boost or a distinct "🎯 High Conviction" tag, not a replacement for the broadcast floor (would cut call volume ~90%).
n=26 is real but not huge — re-validate at 2x this sample before leaning on it exclusively.
2 Soft-penalize (don't hard-cut) buyers≥300 as "crowded"
n=140 (53% coverage kept), +4.9pt win2x, sign-consistent train/test — the volume-preserving alternative to the aggressive sweet-spot-only filter.
Best risk/reward for a broad scoring change: modest, real, durable lift while keeping half the call volume. Pair with #1 above rather than choosing one — #1 catches the best of it, #2 improves the rest.
The GT buyer-count correction (f, ~11.5% deflation) does not materially change this call either way — worth fixing for data-quality reasons elsewhere (authenticity score accuracy), but not a blocker for shipping #1/#2.
3 Add lead_time_sec into the composite score
n=83 (31% coverage — logging gap, confirm why 69% of rows lack it before relying on it), +6.4pt win2x, sign-consistent.
Cheapest fix on the list — already computed, currently discarded. Real, modest, durable signal.
Fix the logging gap first so it's populated on close to 100% of rows, not 31%, before weighting it meaningfully.
4 DO NOT ship the mcap-ceiling tightening (e) as evaluated
Every ceiling tested (50K/80K/100K) is flat-to-negative vs the TRUE full-population baseline, and 50K/100K flip sign between train and test.
Corrects an assumption carried over from the earlier band-only comparison (§5) — a real, useful correction, not a null result to hide.
None — this is a "don't do it" finding, zero downside to not shipping it.
5 Hold the top10≥60% penalty-removal (d) for a live shadow test, not a direct ship
n=19, directional-only; the real hard-reject population (top10 stacked with bundle/thin-liq) is invisible to this dataset by construction.
Matches the original audit's own caution (§4C) — grow the sample via record-only logging before touching the live gate.
Acting on n=19 risks overfitting to a handful of tokens.
6 Report BOTH peak-capturable and hold-through win-rate going forward (g)
32.7pt gap (36.5% peak vs 3.8% hold, n=104 closed) — the single largest number in this whole backtest.
Not a scoring change — a reporting-honesty fix, and a strong argument that an EXIT signal may be worth more future effort than further entry-side tuning.
n=104 will keep growing as the other 161 still-tracking calls close — re-check this gap monthly, don't treat it as final.

10G — Overfitting discipline, stated plainly

👁 EyezOn Data — read-only backtest, live engine/DB/deploy untouched · harness: perception_backtest.py (scratchpad) · @sudiyasa_'s desk only