ULTIMATE AI V2 — RESEARCH BRAINSTORM & METRIC REGISTER Deterministic
Blackjack Decision Engine Research Created: 16 September 2026 Status:
FUTURE RESEARCH REGISTER — NOT ACTIVATED Frozen reference release: CLI
V9.2.46 / Historical Replay v15.10.195

====================================================================== 1.
PURPOSE
======================================================================

This file is the standing brainstorm/register for Ultimate AI V2.

Its purpose is to capture testable questions and every useful measurable
input already available, or derivable without inventing evidence, from
the completed Deterministic Blackjack research environment.

It is NOT an instruction to activate Ultimate AI V2. Ultimate AI V1
remains frozen. The released blackjack engine, controller rules,
evidence, cardstreams and Replay remain unchanged.

The register deliberately includes metrics that may ultimately prove
useless. V2 research should be allowed to reject variables as well as
retain them.

Core research question:

Can combinations of information genuinely available at the decision
point provide reproducible explanatory or decision value beyond Ultimate
AI V1, without hindsight, noise, data leakage or false next-card
prediction?

A particularly important future question is:

Why did the small observed/live population reach £200 much more
frequently than the much larger shuffled/counterfactual populations, and
can any measurable difference survive independent counterfactual
validation?

The observed difference is a QUESTION TO INVESTIGATE, not evidence that
a special mechanism exists.

======================================================================
2. NON-NEGOTIABLE RESEARCH FIREWALL
======================================================================

Every candidate metric must be tagged by WHEN it becomes knowable.

PRE-WAGER: Information genuinely available before the wager is
committed.

POST-WAGER / PRE-ACTION: Information available after the opening deal
but before Hit/Stand/Double/Split.

IN-HAND: Information genuinely revealed before the particular decision
being assessed.

POST-HAND: Information available only after settlement. May explain
completed journeys, but must never be fed retrospectively into an
earlier decision.

POST-SESSION: Summary/diagnostic information. Explanatory only unless a
later prospective experiment explicitly predeclares its use.

CORPUS / COUNTERFACTUAL: Population information derived outside the live
decision. Must be clearly separated from contemporaneous evidence and
tested for leakage.

Dealer hole card is unavailable until legitimately revealed. Future
cards, future outcomes and future bankroll paths are forbidden. No
hindsight may be disguised as a predictor.

======================================================================
3. PRIMARY OUTCOMES TO EXPLAIN OR PREDICT
======================================================================

Each can be tested separately. Do not collapse them into one vague
“success” score.

-   Reached £110 / £120 / £130 / £140 / £150 / £160 / £170 / £180 / £190
    / £200.
-   Reached each tier within the 30-hand frame.
-   First hand index reaching each tier.
-   Eventually reached each tier where that definition is applicable.
-   Final bankroll.
-   Net profit/loss.
-   Positive / break-even / negative session.
-   Finished £0.
-   Finished £150 where historically tracked.
-   Finished above/below £100.
-   Maximum bankroll reached.
-   Minimum bankroll reached.
-   Peak hand index.
-   Trough hand index.
-   Maximum drawdown and its timing.
-   Recovery after falling below £100.
-   Number of below-£100 episodes.
-   Number of recoveries back above £100.
-   Time/hands to first below-£100 crossing.
-   Time/hands to recovery.
-   Survival to hand 30 / early table exit.
-   Post-£190 outcome: reached £200 versus subsequent loss.
-   Session W/L/P counts and rates.
-   Longest winning run.
-   Longest losing run.
-   Largest winning bankroll run.
-   Largest losing bankroll run.
-   Risk score / Journey Intensity where eligible under frozen
    equations.
-   Ultimate AI V1 outcome.
-   Future V2 outcome.
-   Difference V2 minus V1 on identical cardstream.

======================================================================
4. CARD / SHOE / COMPOSITION METRICS
======================================================================

RAW CARD STATE - Player opening card 1 rank. - Player opening card 2
rank. - Player opening suits. - Dealer up-card rank. - Dealer up-card
suit. - Exact opening three-card rank+suit state. - Normalised starting
rank state. - 10/J/Q/K normalisation. - Player opening pair unordered
equivalent. - Hard/soft starting total. - Pair status. -
Blackjack/natural status. - Ace presence. - Number of 10-value opening
cards. - Dealer Ace / 10-value up-card. - Player total at every decision
point. - Dealer visible total where legitimate. - Number of player cards
currently visible. - Number of all legitimately visible cards.

H/N/L COMPOSITION Frozen descriptive classification: HIGH = 10/J/Q/K/A
NEUTRAL = 7/8/9 LOW = 2/3/4/5/6 Ace remains HIGH for H/N/L composition;
blackjack arithmetic still uses 1/11.

Candidate metrics: - Current visible HIGH count. - Current visible
NEUTRAL count. - Current visible LOW count. - Current visible H/N/L
percentages. - H-N difference. - H-L difference. - N-L difference. -
HIGH minus expected baseline. - NEUTRAL minus expected baseline. - LOW
minus expected baseline. - H/N/L concentration. - Dominant visible
category. - Tied composition. - Entropy/diversity of visible H/N/L
mix. - Number of visible cards supporting the composition estimate. -
Confidence/reliability weighting based on visible sample size.

RECENT CARD WINDOWS Test separately rather than assuming one optimum
window: - Last 1 visible card. - Last 2. - Last 3. - Last 5. - Last
10. - Last 15. - Last 20. - Cards since latest shuffle. - Previous
completed hand. - Previous 2 hands. - Previous 3 hands. - Previous 5
hands. - Previous 10 hands. - Current hand only. - Current + previous
hand. - Current + previous N hands.

For every window: - H/N/L counts and percentages. - Exact ranks. -
10-value count/rate. - Ace count/rate. - Low-card count/rate. -
neutral-card count/rate. - run length of category. - change from
preceding window. - deviation from shoe-to-date composition.

SHOE POSITION / CONSUMPTION - Sequential card index. - Cards consumed
since shuffle. - Cards remaining where architecture permits. -
Percentage penetration. - Hand number since shuffle. - Number of hands
since shuffle. - Shuffle count. - Cards consumed in previous hand. -
Average cards consumed per hand. - Player additional-card count. -
Dealer additional-card count. - Split-related card consumption. -
Double-related card consumption. - Natural-related consumption. -
penetration band. - pre/post penetration-trigger state. - shoe
first/middle/late segment.

PHYSICAL / SHUFFLE INTEGRITY CONTEXT Research-only: - six-deck
permutation identity where independently auditable. - exact
physical-copy identity in future audit instrumentation. - exact
rank+suit recurrence. - normalised state recurrence. - Method
A/B/C1/C2/C3/C4/D/E. - Frozen/Casual path provenance. - seed/replicate
identity. - cards/reservoir size. - shuffle architecture. - common
deterministic Method-B benchmark identity.

Do not interpret H/N/L concentration or recent card sequences as proof
that the next card is more likely to be HIGH/NEUTRAL/LOW. Any claimed
predictive value must survive prospective/counterfactual testing.

======================================================================
5. HISTORICAL SAME-STATE METRICS
======================================================================

OBSERVED HISTORY For the exact decision-equivalent starting rank
state: - number of previous observed occurrences. - wins. - losses. -
pushes. - W/L/P percentages. - average P/L. - median P/L. - P/L
distribution. - average wager. - wager distribution. - action
frequencies. - Hit frequency. - Stand frequency. - Double frequency. -
Split frequency. - Frozen-followed frequency. - non-Frozen action
frequency where applicable. - previous outcome. - most recent
equivalent-state outcome. - number of sessions containing the state. -
time/hands since previous occurrence.

Strict matching definition: player two opening ranks + dealer up-card;
10/J/Q/K normalised to 10; suits ignored for normalised matcher; player
opening pair/order treated equivalently.

EXACT SUIT+RANK HISTORY — SEPARATE FEATURE - exact player card
ranks+suits + dealer up-card rank+suit. - occurrence count. - session
count. - W/L/P. - average P/L. - compare exact-suit state against
normalised-state result.

Exact suit information must not be treated as strategically meaningful
merely because it is available. It is a candidate control/negative
feature.

======================================================================
6. A–E ROBUSTNESS CORPUS METRICS
======================================================================

For the current session’s own robustness family: A, B, C1, C2, C3, C4,
D, E.

Each method: 5,000 Frozen + 5,000 Casual paths. Total: 80,000
individually searchable policy paths per session.

Frozen and Casual paths qualify independently. Pair identity is
provenance, not an eligibility restriction.

For each matching starting state: - total occurrence count. - Frozen
occurrence count. - Casual occurrence count. - occurrence rate per
resolved hand. - W/L/P count. - W/L/P percentages. - average P/L. -
median P/L if reconstructable. - P/L variance/dispersion. -
positive/negative/tied average. - average wager. - wager distribution. -
action distribution. - Frozen/Casual outcome difference. -
method-specific outcome difference. - method agreement/disagreement. -
number of A–E methods containing the state. - minimum occurrence count
across methods. - maximum occurrence count across methods. - state
prevalence across all 80,000 paths.

H/N/L AFTER OPENING THREE - HIGH count/percentage. - NEUTRAL
count/percentage. - LOW count/percentage. - additional cards consumed. -
HIGH-dominant hand count. - NEUTRAL-dominant hand count. - LOW-dominant
hand count. - TIED-composition count. - average post-opening H/N/L
mix. - distribution by method. - distribution by Frozen/Casual. -
compare observed subsequent composition with corpus distribution.

FULL-HAND COMPOSITION - full-hand HIGH/NEUTRAL/LOW. - full-hand category
percentages. - cards per resolved hand. - composition variance. - method
differences. - policy differences.

CORPUS COVERAGE - whether observed normalised state appears in corpus. -
occurrence count for each observed state. - whether exact rank+suit
opening appears where reconstructable. - coverage percentage. -
rare-state flag. - common-state flag. - corpus sample-size/reliability
flag.

Known validation context should remain documented: observed normalised
opening-state coverage was found to be complete in the post-project
audit; exact-suit/rank coverage requires the retained detail available
for the relevant session/method and must not be invented where
historical provenance is incomplete.

======================================================================
7. WAGER / EXPOSURE METRICS
======================================================================

-   Current wager.
-   Previous wager.
-   Opening wager.
-   Minimum wager.
-   Maximum wager.
-   Mean wager.
-   Median wager.
-   wager standard deviation/dispersion.
-   wager as percentage of current bankroll.
-   wager as percentage of starting bankroll.
-   wager change from previous hand.
-   absolute wager change.
-   percentage wager change.
-   consecutive wager increases.
-   consecutive wager decreases.
-   wager escalation run.
-   wager de-escalation run.
-   survival-hand status.
-   escalation/burst status.
-   rebound status where applicable.
-   wager band.
-   wager tier.
-   £5 denomination compliance.
-   £2.50 residual provenance where applicable.
-   total session exposure.
-   cumulative exposure at current hand.
-   exposure remaining to hand cap (descriptive only).
-   exposure before/after £100 crossing.
-   exposure before/after tier achievement.
-   exposure during winning streaks.
-   exposure during losing streaks.
-   exposure-weighted W/L/P.
-   profit/loss per unit exposure.
-   house-edge/loss-per-exposure diagnostic where appropriate.
-   high-wager win rate.
-   high-wager loss rate.
-   low-wager win/loss rate.
-   wager at natural.
-   wager at double.
-   wager at split.
-   effective exposure after double/split.
-   high exposure coinciding with favourable/unfavourable card outcomes.

======================================================================
8. BANKROLL / JOURNEY METRICS
======================================================================

AT EACH HAND - bankroll before wager. - bankroll after settlement. -
change this hand. - cumulative P/L. - distance from £100. - distance
from £200. - distance from next tier. - current peak. - current
trough. - drawdown from peak. - recovery from trough. - number of
previous below-£100 crossings. - number of previous recoveries. - hands
since peak. - hands since trough. - hands since first below £100. -
hands since last recovery. - current tier reached. - highest tier
previously reached. - tier regression. - tier momentum/change.

JOURNEY SHAPE - maximum drawdown. - maximum run-up. - area below £100. -
area above £100. - time spent below £100. - time spent above £100. -
number of bankroll direction reversals. - volatility of hand-to-hand
bankroll changes. - longest positive bankroll run. - longest negative
bankroll run. - magnitude of positive/negative runs. - recovery slope. -
drawdown slope. - early/middle/late session P/L. - first 5 / 10 / 15 /
20 hand P/L. - remaining hand budget. - £190-to-£200 conversion. - loss
after reaching £190. - £200 reach hand index. - whether £200 reached
after prior drawdown. - whether £200 reached without falling below £100.

Important: future bankroll information is forbidden at a contemporaneous
decision point. These metrics may be explanatory labels for completed
journeys or use only their genuinely known prefix at the decision being
tested.

======================================================================
9. ACTION / DECISION METRICS
======================================================================

-   Frozen/basic decision.
-   actual observed action.
-   Casual action.
-   Ultimate AI V1 action.
-   future V2 action.
-   Hit.
-   Stand.
-   Double.
-   Split.
-   natural resolution.
-   dealer-natural early resolution.
-   action legality.
-   action agreement/disagreement.
-   Frozen followed yes/no.
-   override proposed yes/no.
-   override accepted/rejected.
-   override confidence.
-   evidence layers supporting override.
-   number of supporting independent signals.
-   number of conflicting signals.
-   convergence score.
-   provenance completeness.
-   decision state frequency.
-   rare-decision flag.
-   decision outcome.
-   decision P/L.
-   counterfactual P/L on identical cardstream where valid.
-   action-specific exposure.
-   action-specific corpus outcome.
-   action-specific H/N/L context.

Never judge a decision as correct merely because its realised outcome
was good.

======================================================================
10. OUTCOME / RESULT METRICS
======================================================================

HAND LEVEL - Win / Loss / Push. - natural blackjack. - dealer natural. -
simultaneous natural/push. - player bust. - dealer bust. - double
result. - split child A result. - split child B result. - overall split
result. - hand P/L. - return relative to wager. - return relative to
effective exposure. - cards consumed. - settlement type.

SESSION LEVEL - total wins/losses/pushes. - W/L/P percentages. - natural
counts/rates. - dealer-natural counts/rates. - split count/rate. -
double count/rate. - action counts/rates. - final bankroll. - £200
reach. - all tier reaches. - exposure. - average wager. -
peak/trough/drawdown. - positive/negative result.

======================================================================
11. SEQUENCE / TEMPORAL METRICS
======================================================================

Candidate historical context only where genuinely available: - previous
hand result. - last 2 W/L/P. - last 3. - last 5. - last 10. - current
winning streak length. - current losing streak length. - push streak. -
W/L/P transition frequencies. - win-after-win. - win-after-loss. -
loss-after-win. - loss-after-loss. - action sequence. - wager
sequence. - bankroll-direction sequence. - card-category sequence. -
10-value run length. - no-10-value run length. - HIGH run. - NEUTRAL
run. - LOW run. - changes in H/N/L composition across hands. - result
following wager escalation. - result following wager reduction. - result
following drawdown. - result following recovery.

These are candidate descriptors, not evidence of gambler’s-fallacy
effects. Any predictive claim must beat suitable controls.

======================================================================
12. PLAYER / CONTROLLER CONTEXT
======================================================================

Where evidence exists: - observed/live player. - Frozen. - Casual. -
Stable. - Cautious. - Aggressive bursts. - Max aggressive. -
Rebounder. - Stat-Watching controller. - Ultimate AI V1. - future
Ultimate AI V2. - declared side-controller. - controller state. -
controller rule invoked. - policy agreement. - policy divergence. -
previous controller interventions. - intervention count. - intervention
P/L (diagnostic only). - controller-specific exposure. -
controller-specific £200 reach. - controller-specific drawdown. -
controller-specific recovery.

Do not infer unrecorded human psychology or motivations from outcomes.

======================================================================
13. SESSION CONTEXT / CROSS-SESSION METRICS
======================================================================

-   Session ID.
-   hand index.
-   formal-start status.
-   source completion status.
-   observed/live versus replay/counterfactual.
-   total cards.
-   shuffle count.
-   cards since latest shuffle.
-   correction/provenance status.
-   session evidence completeness.
-   session-specific A–E corpus availability.
-   prior-session state occurrences.
-   cumulative prior observed state occurrences.
-   cross-session state prevalence.
-   cross-session W/L/P for state.
-   cross-session average P/L for state.
-   cross-session wager context.
-   cross-session £200 reach.

Do not let future sessions leak into an earlier historical decision when
testing a contemporaneous AI.

======================================================================
14. SHUFFLE-METHOD / ROBUSTNESS METRICS
======================================================================

-   Method A observed-card permutation.
-   Method B fresh full six-deck deterministic RNG benchmark.
-   C1 50% penetration.
-   C2 65%.
-   C3 75%.
-   C4 80%.
-   D physical-style riffle/strip/cut proxy.
-   E automatic/continuous proxy.
-   method-specific hand count.
-   method-specific £200 reach.
-   method-specific final bankroll.
-   method-specific W/L/P.
-   method-specific exposure.
-   method-specific H/N/L.
-   method-specific state prevalence.
-   method-specific post-opening composition.
-   method-specific divergence from observed.
-   Frozen/Casual within-method difference.
-   across-method variance.
-   across-method agreement.
-   sensitivity to penetration.
-   sensitivity to persistent finite shoe versus remix proxy.

Validation V1 boundary: Method B uses a common deterministic 5,000-seed
benchmark schedule across sessions. Cross-session uses must not be
counted as independent fresh Method-B shoes.

======================================================================
15. NATURAL / SPECIAL-CASE METRICS
======================================================================

-   player natural.
-   dealer natural.
-   simultaneous naturals.
-   natural rate.
-   dealer-natural rate.
-   wager on natural.
-   bankroll effect.
-   natural timing architecture.
-   Clean Peek / historical frozen / no-peek architecture where
    relevant.
-   double eligibility.
-   split eligibility.
-   split pair rank.
-   split child totals.
-   split child outcomes.
-   effective split exposure.
-   double total.
-   double result.
-   residual £2.50 provenance from 3:2 blackjack.

======================================================================
16. RISK / INTENSITY / EXISTING DERIVED METRICS
======================================================================

Where eligible and reproducible from frozen definitions: - Risk score. -
Journey Intensity. - every frozen component input to those scores. -
every frozen component output. - risk band. - intensity band. - exposure
component. - drawdown component. - wager component. - bankroll-path
component. - recovery component where defined. - controller/provenance
component where defined.

Never silently redefine a frozen score to suit V2. If V2 creates a new
score, give it a new name and equation.

======================================================================
17. CANDIDATE INTERACTION FEATURES
======================================================================

Single metrics may be weak while combinations are informative. Candidate
interactions must therefore be tested explicitly.

Examples: - H/N/L × wager size. - H/N/L × bankroll. - H/N/L ×
drawdown. - H/N/L × shoe position. - H/N/L × starting state. - H/N/L ×
action. - H/N/L × exposure. - H/N/L × recent W/L/P. - starting state ×
wager. - starting state × bankroll. - starting state × corpus outcome. -
starting state × method. - wager escalation × subsequent visible-card
context. - high exposure × outcome. - high exposure × double/split. -
drawdown × wager escalation. - recovery state × wager. - tier proximity
× wager. - tier proximity × action. - £190 state × exposure. - £190
state × card composition. - action disagreement × corpus evidence. - V1
confidence × H/N/L evidence. - observed-history evidence ×
robustness-corpus evidence. - recent-card evidence × shoe-to-date
evidence. - exact-state frequency × corpus sample size. - method
agreement × override confidence.

Interactions should be predeclared or corrected for multiple testing
where large-scale exploration is used.

======================================================================
18. WEIGHTING CANDIDATES
======================================================================

V2 may eventually weight evidence, but weights must be learned/validated
rather than chosen because they produce a desirable result.

Possible weighting dimensions: - sample size. - recency. - exactness of
state match. - observed versus synthetic provenance. - number of
independent sessions. - number of A–E methods agreeing. - Frozen/Casual
agreement. - method robustness. - confidence interval width. - outcome
variance. - exposure relevance. - wager similarity. - bankroll
similarity. - shoe-position similarity. - penetration similarity. -
temporal proximity. - decision/action similarity. - H/N/L sample size. -
exact rank+suit match versus normalised match. - prior-only evidence
availability. - provenance quality/completeness.

Candidate weighting families: - equal weighting. - sample-size
weighting. - capped sample-size weighting. - inverse-variance
weighting. - recency weighting. - state-similarity weighting. -
method-consensus weighting. - provenance/reliability weighting. -
Bayesian/shrinkage weighting. - regularised learned weights. -
deliberately simple predeclared scoring rules.

No weight may use the future outcome of the decision it is intended to
inform.

======================================================================
19. BASELINES / CONTROLS V2 MUST BEAT
======================================================================

-   Ultimate AI V1 frozen baseline.
-   Frozen policy.
-   Casual policy.
-   no-HNL ablation.
-   HNL-only experimental layer.
-   shuffled/random HNL labels as negative control.
-   exact-state-only.
-   observed-history-only.
-   robustness-corpus-only.
-   bankroll/wager-only.
-   recent-WLP-only.
-   shoe-position-only.
-   simple equal-weight model.
-   random/noise features.
-   always-stand / always-hit-to-bust controls where scientifically
    relevant.
-   identical cardstream V1 versus V2.
-   identical cardstream feature ablations.

A complicated model that does not materially outperform a simpler
baseline should not be preferred merely because it is more
sophisticated.

======================================================================
20. STATISTICAL / VALIDATION METRICS
======================================================================

Depending on the target: - sample count. - event count. -
percentage/rate. - mean. - median. - standard deviation. - variance. -
interquartile range. - minimum/maximum. - confidence interval. - effect
size. - absolute difference. - relative difference. - odds/risk ratio
where suitable. - correlation. - calibration. - Brier score for
probability outputs. - log loss for probability outputs. - ROC-AUC where
appropriate. - precision. - recall. - specificity. - sensitivity. - F1
where appropriate. - mean absolute error. - root mean squared error. -
rank correlation. - bootstrap interval. - permutation test. -
exact/binomial test where appropriate. - multiple-testing correction. -
false-discovery control. - holdout performance. - cross-session
validation. - leave-one-session-out validation. - method holdout. - seed
holdout. - rare-state performance. - common-state performance. -
stability across A–E. - stability Frozen versus Casual. -
coefficient/weight stability. - ablation effect. - out-of-sample
degradation. - calibration drift.

Statistical significance alone is not sufficient. Practical effect size,
reproducibility and decision usefulness matter.

======================================================================
21. THE £200 REACH INVESTIGATION
======================================================================

Predeclare before analysis.

QUESTION: Why did the observed/live sessions show a much higher
£200-reach proportion than the large shuffled/counterfactual
populations?

FIRST TEST: Determine how unusual the observed count is under the frozen
benchmark populations when the sample size is only 11 sessions.

Do not begin by assuming the observed percentage represents a different
mechanism.

IF THE DIFFERENCE WARRANTS EXPLANATION, compare:

CARDS - opening-state distribution. - exact rank/suit distribution. -
H/N/L visible composition. - post-opening H/N/L composition. -
naturals. - doubles/splits. - card consumption. - shoe position.

TIMING - favourable cards at high exposure. - unfavourable cards at high
exposure. - outcome timing relative to wager escalation. - outcome
timing relative to drawdown/recovery. - early versus late favourable
runs.

WAGERS - average wager. - wager distribution. - escalation. - high-wager
win/loss. - effective double/split exposure. - loss per unit exposure.

JOURNEY - early bankroll growth. - drawdown. - recovery. - time
above/below £100. - £190 conversion. - winning/losing runs.

DECISIONS - Frozen agreement. - action mix. - doubles/splits. - V1
interventions where applicable.

CORPUS CONTEXT - same-state outcomes. - method agreement. - observed
versus synthetic state outcomes. - state prevalence. - post-opening
composition for equivalent states.

A candidate explanation should then be converted into a predeclared
counterfactual hypothesis and tested on evidence not used to discover
it.

======================================================================
22. ANTI-OVERFITTING / ANTI-DATA-MINING RULES
======================================================================

The environment permits hundreds or thousands of retrospective
questions. That is a strength only if controlled.

For every formal V2 investigation record: - hypothesis before result. -
target outcome. - candidate metrics. - temporal eligibility. - discovery
population. - validation population. - exclusion rules. - weighting
method. - success criterion. - rejection criterion. - negative result. -
limitations.

Separate: DISCOVERY from CONFIRMATION.

Do not repeatedly inspect holdout results while tuning weights. Do not
select only favourable sessions. Do not remove inconvenient methods. Do
not redefine a state after seeing its outcome. Do not convert
correlation into causation. Do not claim card counting/prediction merely
from H/N/L imbalance. Do not use future cards. Do not use future
bankroll. Do not use later sessions when emulating an earlier live
decision. Do not treat 80,000 paths/session as 80,000 statistically
independent experiments without accounting for
architecture/pairing/repeated seeds.

======================================================================
23. FEATURE ABLATION PLAN
======================================================================

Any eventual V2 should be decomposable.

Candidate experiments: V1 V1 + HNL V1 + observed same-state V1 +
robustness same-state V1 + bankroll/journey prefix V1 + wager/exposure
V1 + shoe position V1 + recent WLP V1 + HNL + same-state V1 + HNL +
wager V1 + HNL + bankroll V1 + all eligible features V2 minus each
feature family one at a time

This establishes what actually adds value rather than crediting the
whole model for a result generated by one component.

======================================================================
24. EXPLAINABILITY / PROVENANCE OUTPUT
======================================================================

Every V2 decision or research prediction should be reproducible and
explain:

-   session.
-   hand.
-   decision time.
-   evidence available.
-   evidence excluded.
-   V1 recommendation.
-   V2 recommendation.
-   whether V2 differs.
-   candidate metrics used.
-   metric values.
-   weights.
-   weighted contributions.
-   corpus sample sizes.
-   method provenance.
-   confidence/uncertainty.
-   reason for override/non-override.
-   counterfactual result after settlement, clearly labelled hindsight.
-   model/version.
-   experiment/version.
-   deterministic seed where relevant.

The end user should be able to ask: “Why did V2 say this?” and receive
an evidence-based answer rather than an opaque score.

======================================================================
25. STOP / REJECTION CONDITIONS
======================================================================

V2/HNL should be rejected, simplified or retained only as descriptive
context if:

-   effect disappears on holdout.
-   effect depends on hindsight.
-   effect depends on one session.
-   effect depends on one shuffle method.
-   weights are unstable.
-   random/noise features perform similarly.
-   H/N/L adds no value beyond starting state.
-   apparent gain disappears after sample-size correction.
-   improvement is statistically detectable but practically trivial.
-   performance requires excessive complexity.
-   provenance cannot be reconstructed.
-   V2 increases override frequency without improving outcomes.
-   result cannot be reproduced on identical streams.

A negative result is a successful research outcome.

======================================================================
26. FUTURE DATA / EXTERNAL VALIDATION
======================================================================

If genuine external casino/blackjack data becomes available: -
reconstruct cards in exact chronological/deal order. - verify
reconstruction fidelity first. - do not invent ambiguous card order. -
replay through frozen mechanism. - compare external fit with old model
and casino-realism architecture. - classify fit as closer to either,
between, or neither. - preserve source provenance. - test V2 only after
its architecture/weights are frozen. - do not retrain on the external
validation set.

======================================================================
27. STANDING BRAINSTORM QUESTIONS
======================================================================

Examples to retain for future prioritisation:

-   Why was observed £200 reach unusually high?
-   Was favourable outcome timing more important than overall card
    composition?
-   Did high wagers coincide unusually often with favourable
    states/outcomes?
-   Were observed drawdowns less damaging because favourable hands
    arrived at particular exposure points?
-   Did doubles/splits contribute disproportionately to successful
    journeys?
-   Does H/N/L add anything once exact starting state is known?
-   Does recent H/N/L add anything once shoe-to-date composition is
    known?
-   Does same-state historical W/L/P add anything beyond the frozen
    decision?
-   Does robustness-method agreement identify stable versus noisy
    states?
-   Are observed successful sessions distinguishable before they
    succeed?
-   Can a model trained without the observed 11 sessions explain them?
-   Do the three non-£200 observed sessions contradict candidate
    explanations derived from the eight successful ones?
-   Are exact-suit matches useful only as a negative/control feature?
-   Does shoe penetration alter the usefulness of any H/N/L metric?
-   Is apparent V2 improvement merely increased exposure/risk?
-   Does V2 improve final bankroll while worsening drawdown?
-   Does V2 improve £200 reach while increasing ruin?
-   Are gains robust across A–E rather than concentrated in one proxy?
-   Does V2 retain value under Frozen and Casual paths?
-   Can the explanation survive leave-one-session-out testing?
-   Does the explanation survive an independent synthetic benchmark
    corpus?
-   Would an SME/croupier predeclared alternative controller outperform
    the original on the identical journey?

======================================================================
28. PRIORITISATION RULE
======================================================================

Classify proposed work:

A. INTEGRITY / VALIDATION Does the apparatus do what the research
claims?

B. SCIENTIFIC EXTENSION Does a predeclared new hypothesis add
reproducible knowledge?

C. EXPLORATORY CURIOSITY Interesting, but not automatically worthy of
formal testing.

Normally execute A and carefully selected B. Archive C unless it becomes
a clearly justified hypothesis.

======================================================================
29. LAYMAN SUMMARY
======================================================================

Ultimate AI V2 should not be built simply by giving an AI every
statistic and asking it to find a way to win.

We now have enough data to find patterns by accident.

The useful question is whether information that was genuinely known at
the time — cards already seen, the current hand, bankroll, wager,
previous results, shoe position and large same-state robustness
populations — contains a repeatable signal that still works when tested
on data that was not used to discover it.

Every candidate metric can be measured and weighted, but it does not
have to survive. A strong V2 may ultimately use only a small subset.

The £200 difference is an excellent research question because it is
large enough to deserve explanation, but the observed sample is only 11
sessions. The first job is therefore to establish whether 8 successful
sessions out of 11 is genuinely unusual relative to the frozen
benchmark. Only then should AI look for measurable reasons.

If a proposed explanation works only after the outcome is already known,
it has failed.

If H/N/L proves descriptive but not predictive, that is a valid finding.

If V2 cannot reproducibly improve on frozen V1 under identical
cardstreams, V1 remains the stronger model.

======================================================================
36. SESSION-PLATE METRIC INTEGRATION — REQUIRED SME WORKSTREAM
======================================================================

Ultimate AI V2 must not become a separate analysis silo. Candidate V2 signals
must be compared with the established metrics already reported/verified for
the session-evidence plates.

Compare, where eligible:
- final bankroll, net P/L and £200/tier reach.
- bankroll path, peak/trough and hand indexes.
- maximum drawdown and recovery behaviour.
- exposure, average wager and wager/action alignment.
- W/L/P counts, rates and sequences.
- doubles, splits and naturals.
- Risk and Journey Intensity with frozen component inputs/outputs.
- A–E shuffle robustness.
- Frozen versus Casual counterfactuals.
- Stat-Watching/other declared side-controller results where present.
- Ultimate AI V1 results.
- source-completion/provenance and card/shuffle/correction status.

The existing session plates are the established evidence framework.
V2 must determine whether new evidence explains, complements or materially
improves upon those existing measures.

Research flow:
PREPARED EVIDENCE -> V2 CANDIDATE EVIDENCE -> SESSION-PLATE METRICS ->
WEIGHT/INTERACTION ANALYSIS -> COUNTERFACTUAL + CROSS-SOURCE VALIDATION ->
SME FINDINGS.


======================================================================
37. PHYSICAL-CARD OPENING ALIGNMENT ACROSS CORPORA
======================================================================

Across every corpus where immutable physical-card provenance exists, or can be
legitimately reconstructed, investigate physical identity alignment among:

PLAYER OPENING CARD 1 + PLAYER OPENING CARD 2 + DEALER UP-CARD.

Record separately:
- 3/3 physical-card alignment.
- 2/3 physical-card alignment.
- 1/3 physical-card alignment.
- 0/3 physical-card alignment.
- which two positions align: P1+P2, P1+Dealer, P2+Dealer.
- exact rank+suit agreement.
- normalised starting-state agreement.
- physical deck-copy agreement.

Compare with shoe pointer, penetration, cards since shuffle, shuffle method
A/B/C1/C2/C3/C4/D/E, corpus/source, £200 reach, bankroll, exposure, drawdown,
wager/action, H/N/L and established session-plate metrics.

High-value interaction:
PHYSICAL-CARD ALIGNMENT × SHOE POSITION × SHUFFLE ARCHITECTURE.

Question:
When similar openings recur, are they represented by different physical cards
as expected, or is there repeated physical-card alignment; and does its
frequency vary with penetration, shuffle architecture or journey outcome?

PROVENANCE BOUNDARY:
Matching rank+suit does NOT establish matching physical deck copy.
The frozen Card representation did not retain immutable Deck-1...Deck-6 copy
identity. Where a corpus lacks physical-copy provenance, report:
NOT DETERMINABLE FROM RETAINED PROVENANCE.
Never manufacture identity from card appearance.


======================================================================
38. VALIDATION V1 — PHYSICAL SIX-DECK / SHUFFLE CONTEXT
======================================================================

Completed post-project Validation V1 provides an integrity reference.
Frozen CLI V9.2.46 and Replay v15.10.195 remained unchanged.

Independent audit-only Deck-1...Deck-6 identities were used for Method B.

Results:
- 5,000 deterministic Method-B seeds audited.
- every starting shoe contained exactly 312 physical identities.
- every shuffled shoe contained all 312 identities exactly once.
- failed permutations: 0/5,000.
- no physical identity duplicated, lost or replaced.
- all 5,000 physical three-card opening identity tuples were unique.
- repeated normalised states were represented by different physical tuples.
- exact rank+suit patterns could repeat using different physical deck copies.

Method-B boundary:
the same deterministic benchmark seed schedule is reused across sessions.
Do not multiply it across sessions and claim independent fresh Method-B shoes.

For A/C/D/E and external corpora, physical-copy conclusions are allowed only
where the source preserves sufficient provenance.


======================================================================
39. AUTOMATIC SHUFFLER / CONTINUOUS-SHUFFLE RESEARCH BOUNDARY
======================================================================

SMEs should distinguish:
1. Automatic shuffler — shuffles a pack/shoe, then finite-shoe dealing occurs.
2. Continuous shuffling machine — played/discarded cards progressively return
   toward the available population during ongoing play.

Method E remains an AUTOMATIC/CONTINUOUS-SHUFFLING PROXY:
completed hand -> consumed cards returned -> captured reservoir remixed ->
next hand.

It is not an exact engineering simulation of a particular commercial CSM.
Real machines may differ in reinsertion mechanics, latency and procedure.

V2 questions:
- Does a finite-shoe depletion signal weaken under Method E?
- Does recent H/N/L retain value when completed cards are returned/remixed?
- Does shoe-pointer value collapse when persistent depletion is removed?
- Do effects strengthen with C1/C2/C3/C4 penetration?
- Are effects consistent under A/B/D persistent/rearranged shoe conditions?
- If a claimed depletion effect remains under E, is another mechanism, leakage
  or noise a better explanation?

Method E therefore serves as a potential MECHANISM-BREAKING / NEGATIVE CONTROL,
not merely another shuffle method.


======================================================================
40. SHOE POINTER + PHYSICAL ALIGNMENT + COMPOSITION
======================================================================

Candidate interactions for SME testing:
- physical opening alignment × shoe pointer/penetration.
- physical opening alignment × H/N/L.
- physical opening alignment × exact/normalised state.
- physical opening alignment × wager/exposure.
- physical opening alignment × £200 reach.
- shoe pointer × H/N/L × wager.
- shoe pointer × H/N/L × bankroll/drawdown.
- shoe pointer × same-state corpus evidence.
- penetration × post-opening composition.
- shuffle method × shoe pointer × H/N/L.
- persistent-shoe versus Method-E comparison.

No interaction is assumed useful merely because it can be calculated.


======================================================================
41. OBSERVED OPENING-STATE COVERAGE — VALIDATION CONTEXT
======================================================================

Completed post-project check:
- 273/273 observed historical opening hands had their NORMALISED starting state
  represented in the corresponding session-specific A–E corpus.
- coverage = 100.0%; missing = 0.

This is normalised decision-equivalent rank-state coverage, not physical-card
identity coverage.

Exact rank+suit follow-up:
- Sessions 2–11: 243/243 observed openings positively represented at least once
  in their own A–E robustness corpus.
- Session 1 retains older provenance limitations.
- one Session-1 dealer up-card suit is historically uncertain.
- four further Session-1 exact-suit openings cannot be conclusively resolved
  across the retained full A/C/D/E detail.
- these are NOT DETERMINABLE, not proven absent.


======================================================================
42. SME HANDOVER — WHAT IS COMPLETE / WHAT REMAINS
======================================================================

PREPARED:
deterministic architecture; frozen CLI/Replay; exercised live environment;
GUI/CLI UAT; observed journeys; internal PRIMARY/HOLDOUT; BlackjackBench;
Casino Benchmark; corrected 50M/parquet; A–E robustness; H/N/L same-state
corpora; historical same-state evidence; session plates/metrics; Ultimate AI
V1; provenance/integrity work; physical six-deck Validation V1; V2 metric
brainstorm.

REMAINS FOR SME:
prioritise hypotheses; determine useful metrics/interactions; establish
weighting; compare V2 with session-plate metrics; investigate £200 disparity;
test shoe pointer/penetration; persistent-shoe versus Method-E mechanisms;
physical-card opening alignment where provenance permits; cross-source
transfer; ablation/negative controls; independent validation; and determine
whether a future V2 implementation is scientifically justified.

Ultimate AI V2 is an SME RESEARCH OPPORTUNITY, not an unfinished promised
feature.


======================================================================
43. FINAL SME PRINCIPLE
======================================================================

The environment is prepared to allow deeper questions to be asked.

The SME must determine which relationships are genuine, descriptive only,
destroyed by controls, transferable across independent populations, robust on
identical cardstreams, and useful beyond frozen V1.

The strongest result may be a smaller V2 than initially imagined, or a finding
that V1 should remain unchanged. Negative results remain valid scientific
results.


======================================================================
44. PREAMBLE / ENTRY-CONTEXT EFFECT — SME INVESTIGATION
======================================================================

Preamble must be treated as a distinct candidate research dimension rather
than silently merged into the formal session. Existing evidence architecture
already preserves the preamble as entry/shoe context and keeps it separate
from formal-session hand count, bankroll path, exposure, W/L/P, drawdown and
decision totals.

The observed experience that the two preamble sessions felt unusually
difficult and ended with heavily depleted bankrolls is a research observation
/trigger for investigation, NOT evidence that preamble caused the depletion
or produced harder cards. The small observed sample must not be used to infer
a mechanism.

SME questions include:
- Does entering the formal journey after a preamble materially change the
  distribution of subsequent cards, opening states or outcomes?
- Is any apparent effect explained simply by the shoe pointer / penetration
  reached before formal play begins?
- Does preamble alter the remaining-shoe H/N/L composition or the distribution
  of exact/normalised starting states subsequently encountered?
- Are difficult-looking hands, high-exposure losses, doubles/splits, naturals,
  drawdowns, recovery behaviour or £200 reach different after preamble?
- Does the effect remain after controlling for starting bankroll, wager,
  exposure, shoe position, penetration, shuffle architecture and policy?
- Compare observed preamble journeys with Frozen-from-preamble-start and
  Casual-from-preamble-start counterfactuals where source-supported.
- Compare preamble versus no-preamble populations on otherwise equivalent
  conditions where a valid comparison can be constructed.
- Test whether any apparent preamble effect survives A/B/C1/C2/C3/C4/D/E,
  particularly persistent-shoe methods versus Method E.
- Examine PREAMBLE × SHOE POINTER × H/N/L × WAGER/EXPOSURE interactions.
- Compare preamble findings with established session-plate metrics rather than
  creating a disconnected V2 measure.

Candidate metrics:
- preamble hand count and cards consumed.
- shoe pointer / cards-since-shuffle at formal entry.
- penetration at formal entry.
- preamble W/L/P, P/L, exposure and bankroll movement where source-supported.
- formal-entry bankroll context, kept distinct from formal-session metrics.
- subsequent opening-state frequency and exact-state difficulty descriptors.
- post-entry H/N/L composition.
- first-N-formal-hands W/L/P and P/L.
- early drawdown, trough timing and recovery.
- wager/exposure timing after formal entry.
- doubles/splits/naturals and their P/L contribution.
- £110–£200 tier reach / £200 conversion where comparable.
- Frozen/Casual preamble-start outcomes.
- method-specific and cross-source replication.

CRITICAL BOUNDARIES:
PREAMBLE is not automatically a causal feature and must not be assigned a V2
weight because two observed examples were poor. Discovery and confirmation
must be separated. Formal-session metrics remain formally separate from the
preamble unless an explicitly declared analysis studies their relationship.
No future information may cross the temporal firewall. A null result — that
preamble adds no information once shoe position/penetration and other entry
conditions are controlled — is a valid and important result.

======================================================================
45. £200 REACH — ANCHOR OUTCOME, NOT THE SOLE SUCCESS MEASURE
======================================================================

£200 reach should remain prominent because it is a common and interpretable
cross-population endpoint, but Ultimate AI V2 must not treat it as the sole
definition of journey success.

SME use:
- use £200 reached/not reached as an anchor comparator across compatible
  observed, synthetic, benchmark, parquet and robustness populations.
- test candidate V2 features for association with £200 reach.
- then test the same features against richer continuous/session outcomes:
  maximum bankroll, final bankroll, tier/time-to-reach, maximum drawdown,
  exposure, recovery, W/L/P, Risk/Journey Intensity and wager/action behaviour.
- distinguish a broad journey relationship from a threshold-crossing artefact.
- a journey reaching £199 is not analytically equivalent to one remaining near
  £100 merely because both are coded "did not reach £200".
- a journey briefly touching £200 and subsequently collapsing should not be
  treated as identical to one that reaches and sustains a high bankroll.

The observed 8/11 £200 reach remains an important research trigger, but the SME
must first establish whether that result is unusual for n=11 under the frozen
reference populations before seeking a mechanism.


======================================================================
46. SAME-STATE CORPUS ECONOMIC OUTCOME — FROZEN AND CASUAL AVG P/L
======================================================================

The historical observed same-state panel already reports:
AVG P/L PER MATCHED HAND
for previous observed occurrences.

Ultimate AI V2 / SME analysis should also use the equivalent economic outcome
from the much larger same-state robustness corpus, kept separate by policy:

- FROZEN — average £ P/L per matched hand.
- CASUAL — average £ P/L per matched hand.

Do not require a Frozen-vs-Casual delta as a primary displayed metric.

Where evidence permits, SMEs may additionally examine:
- average P/L per unit wager/exposure, to normalise differing stakes.
- A/B/C1/C2/C3/C4/D/E method-specific average P/L.
- average P/L conditional on shoe pointer/penetration.
- average P/L conditional on H/N/L and exact/normalised starting state.
- average P/L alongside bankroll, drawdown, wager/action and £200 reach.
- observed-history average P/L versus robustness Frozen/Casual average P/L,
  while respecting the very different sample sizes and provenance.

Average P/L is an outcome descriptor, not automatically a recommendation.
A favourable historical or corpus average must not by itself trigger a V2
decision. It must pass temporal eligibility, weighting, counterfactual,
cross-source and independent-validation requirements.

This metric directly links same-starting-state frequency/composition evidence
to financial outcome and should be compared with established session-plate
metrics rather than analysed as an isolated signal.


======================================================================
FINAL DOCUMENT STATUS
======================================================================

Publication title:
ULTIMATE AI V2 — SME RESEARCH FRAMEWORK & INVESTIGATION REGISTER

Companion:
Ultimate AI V2 — Research Design & Handover Plate (for SME)

This register is the detailed SME companion to the concise handover plate.
It records prepared evidence, candidate metrics, research questions,
interactions, validation controls, provenance boundaries and future
investigations. It does not claim that the candidate relationships are
predictive or that a V2 implementation is required.

======================================================================
END OF FINAL SME RESEARCH FRAMEWORK & INVESTIGATION REGISTER
======================================================================
