📄 What's Inside
This report walks through the computer opponent's decision engine end-to-end. This is the same code that runs in your copy of Just Cribbage. No simplifications. Every section quotes the exact GDScript function responsible for each behavior.
🏗 1 System Architecture
The computer opponent is built from a clean pipeline of pure, stateless modules. No single function reaches across into another's state; decisions flow through an explicit data contract.
The AI is a pure function: given the same observation and the same RNG seed, it always produces the same move. This is what makes the game verifiable, not just a promise in the docs, but a property you can test by swapping the opponent's cards and confirming the output doesn't change.
🔒 2 The Data Contract: AlleyObservation
This is the most important structural decision in the whole AI. Before the brain can
make any decision, it receives an AlleyObservation, and that type
is defined to be physically incapable of carrying the player's hidden cards.
var own_hand: Array = [] # the Alley's OWN cards (its 6 to discard, or its play hand) var played: Array = [] # the public pegging pile, in order var running_count: int = 0 # public running count in the current pegging sub-round var opp_card_count: int = 0 # how many cards the opponent HOLDS (a COUNT, never the cards) var starter = null # the cut/starter card -- public once cut (null before) var crib_is_alley: bool = false var dealer: String = "" var scores: Dictionary = {} # public {"player": int, "alley": int}
Notice what is not here. There is no opp_hand field. Not hidden, not
encrypted. The field does not exist. The brain cannot access it because the type
system enforces it. This is the difference between a promise and a
proof.
The one thing that might look like cheating: opp_card_count. But count
of cards in hand is public information at any real cribbage table. You can count the cards fanned in your opponent's hand. The brain uses it only to
weight the probability that the opponent holds a particular rank during threat modeling.
Ranks and suits are never here.
The discard observation is built by AlleyObservation.for_discard(own_six, crib_is_alley)
and the pegging observation by AlleyObservation.for_play(own_hand, played, count, opp.size(), starter).
In both cases player_hand is never passed to the constructor, only its .size(). The test suite in tests/test_no_peek.gd swaps the opponent's
hidden cards for arbitrary values and confirms the brain's output is byte-identical.
The two factory constructors enforce consistent construction everywhere:
static func for_discard(own_six: Array, crib_is_alley_in: bool, ...) -> AlleyObservation: var o := AlleyObservation.new() o.own_hand = own_six o.crib_is_alley = crib_is_alley_in # ...dealer, scores -- but NEVER the opponent's cards return o static func for_play(own_hand_in: Array, played_in: Array, running_count_in: int, opp_card_count_in: int, starter_in = null, ...) -> AlleyObservation: var o := AlleyObservation.new() o.own_hand = own_hand_in o.played = played_in o.running_count = running_count_in o.opp_card_count = opp_card_count_in # a COUNT (int), not the actual cards o.starter = starter_in return o
🃏 3 Discard: Expected Value Analysis
The AI's discard decision is the most intellectually interesting part of the system. It evaluates every possible 2-card throw from its 6-card hand by simulating what happens across all 46 cards that could be cut as the starter.
The Top-Level Entry Point
func choose_discard(obs: AlleyObservation, skill: float, decision_seed: int) -> Array: _o = obs var hand: Array = obs.own_hand if hand.size() <= 2: return hand.duplicate() var ranked := _eval_discards(hand, obs.crib_is_alley) # rank all 15 possible throws if ranked.is_empty(): return hand.slice(0, 2) var t := _temperature(skill) # 0 = expert (exact argmax), >0 = weighted random if t <= 0.0: return (ranked[0]["thrown"] as Array).duplicate() # expert: always the best EV throw var roll := RandomNumberGenerator.new() roll.seed = decision_seed # build a probability distribution over ALL ranked options, weighted by EV var vals: Array = [] for r in ranked: vals.append(float(r["score"])) var pick := _softmax_pick(vals, t, roll) return (ranked[pick]["thrown"] as Array).duplicate()
Evaluating Every Possible Throw
With 6 cards and 2 to throw, there are exactly C(6,2) = 15 possible throws. The AI evaluates all of them. For each candidate throw it keeps 4 cards and calculates the expected score across all 46 possible starter cards.
func _eval_discards(hand: Array, my_crib: bool, ranked_order := false) -> Array: var results: Array = [] var cuts := _unseen_cuts(hand) # the 46 cards not in the AI's hand var inv := 1.0 / float(maxi(1, cuts.size())) for i in range(n): for j in range(i + 1, n): var thrown: Array = [hand[i], hand[j]] var kept: Array = [] # the 4 cards the AI would keep for k in range(n): if k != i and k != j: kept.append(hand[k]) var avg_hand := _avg_hand(kept, cuts) * inv # E[hand points | cut] var own_only := _avg_own_crib(thrown, cuts) * inv # E[these 2 cards in crib | cut] var avg_crib := StatsScript.crib_value(thrown[0].rank, thrown[1].rank, my_crib) var expected := avg_hand + avg_crib var net := avg_hand - avg_crib # dealer wants to maximise; pone wants to minimise crib var safety := avg_hand + (own_only if my_crib else -own_only) results.append({ "thrown": thrown, "avg_hand": avg_hand, "avg_crib": avg_crib, "expected": expected, "net": net, "safety": safety, "score": expected if my_crib else net, })
The scoring formula is crib-seat aware:
Pone score = avg_hand − avg_crib
When you're the dealer, the crib is yours, so high-crib throws are good. When you're the pone, the crib goes to your opponent, so the AI penalises dangerous throws (5s, pairs into the opponent's crib) even if they'd score well on their own.
The 46-Card Starter Pool
func _unseen_cuts(hand: Array) -> Array: var seen := {} for c in hand: seen[c.rank * 4 + c.suit] = true var cuts: Array = [] for su in range(4): for rk in range(1, 14): if not seen.has(rk * 4 + su): cuts.append(Card.new(rk, su)) return cuts
The AI averages over all 46 cards it can't see: its own 6 minus the 52-card deck. It does not exclude the player's hand from the pool, because it doesn't know what's in the player's hand. This is exactly what a fair human player would do. If it excluded the player's cards, that would be a form of peeking.
The Tiebreaker: Safety
func _cmp_disc(a: Dictionary, b: Dictionary) -> bool: if absf(float(a["score"]) - float(b["score"])) > 0.001: return float(a["score"]) > float(b["score"]) return float(a["safety"]) > float(b["safety"]) # within 0.001 EV: prefer safer throw
When two throws are within 0.001 points of each other in expected value, the AI prefers the safer throw. "Safety" here measures the expected crib value contributed by your own thrown cards. High safety means the thrown cards don't help a future crib much either way.
🎶 4 Difficulty Tiers
Four named difficulty levels, each mapping to a skill value between 0 and 1. These are not arbitrary; they're calibrated against explicit win-rate targets validated by the benchmark runner.
| Tier | Skill Value | Temperature (T) | Expert Win Rate Target | Character |
|---|---|---|---|---|
| Easy | 0.12 | 1.2 | 95–98% | Near-misses and real blunders; plays like a beginner |
| Normal | 0.35 | 0.5 | 75–85% | Often plays well, makes exploitable errors |
| Hard | 0.85 | 0.15 | 55–65% | Strong play, occasional suboptimal choices |
| Expert | 1.0 | 0.0 | – | Always picks the highest EV option. No randomness. |
static func _temperature(skill: float) -> float: var s := clampf(skill, 0.0, 1.0) var anchors := [ [0.0, 2.0], # skill 0.0 -> T 2.0 (very erratic) [0.12, 1.2], # skill 0.12 -> T 1.2 (Easy) [0.35, 0.5], # skill 0.35 -> T 0.5 (Normal) [0.85, 0.15], # skill 0.85 -> T 0.15 (Hard) [1.0, 0.0] # skill 1.0 -> T 0 (Expert: exact argmax) ] for i in range(1, anchors.size()): if s <= float(anchors[i][0]): var a: Array = anchors[i - 1] var b: Array = anchors[i] var f := (s - float(a[0])) / maxf(0.000001, float(b[0]) - float(a[0])) return lerpf(float(a[1]), float(b[1]), f) return 0.0
The function linearly interpolates T between adjacent anchors. At Expert (skill 1.0), T = 0 exactly, and the code short-circuits to pure argmax, no randomness at all. At lower skills, T > 0 feeds into the softmax model in the next section.
The previous difficulty model had a "worst-of-15 throw" branch where Easy would pick the worst possible discard roughly 77% of the time, tossing 5-5 into your own crib while keeping garbage, which reads as broken rather than weak. The current model instead uses softmax: every tier picks from the same EV ranking with a calibrated temperature, producing human-plausible near-misses and the occasional real blunder rather than deliberately self-destructive play.
📈 5 The Softmax Mistake Model
This is how the game generates "realistic" mistakes at lower difficulty levels without simply playing randomly. It's a probabilistic selection over EV-ranked candidates, where the probability of picking a suboptimal option decays exponentially with how much EV it gives up.
When T is large (Easy), the probability mass spreads out widely, even options that give up 2+ points of EV get picked sometimes. When T approaches 0 (Expert), the mass concentrates entirely on the best option.
static func _softmax_pick(vals: Array, t: float, roll: RandomNumberGenerator) -> int: var best := -1.0e20 for v in vals: best = maxf(best, float(v)) var weights: Array = [] var total := 0.0 for v in vals: var w := exp(-(best - float(v)) / t) # best option always gets weight 1.0 weights.append(w) # worse options get fractional weight total += w var x := roll.randf() * total # uniform draw scaled to total weight for i in range(weights.size()): x -= float(weights[i]) if x <= 0.0: return i # this is the chosen option return 0
The key property: options more than ~8T below the best are almost never picked. At Easy (T=1.2), anything more than ~9.6 points worse than the best is essentially off the table. This means the AI never throws the single worst possible discard just because it's Easy. It generates near-misses and the occasional medium blunder, which is how real beginners lose points.
In ranked PvE sessions the engine RNG is replaced by FairDeal.hash_roll(), a SHA-256-derived uniform draw that any language can reproduce. This means the server
can verify every AI decision without trusting the client. The softmax logic is
identical; only the source of randomness changes. See choose_discard_ranked()
and choose_play_rolled().
♥ 6 Pegging Decisions
Pegging is harder than discarding: decisions happen in sequence, depend on the public pile, and have to balance immediate points against what the reply might score. The brain handles this with a card-counted net-value model.
Entry Point: choose_play()
func choose_play(obs: AlleyObservation, skill: float, rng: RandomNumberGenerator): _o = obs var legal: Array = _legal_plays(obs.own_hand, obs.running_count) if legal.is_empty(): return null var s: float = clampf(skill, 0.0, 1.0) var t := _temperature(s) if t <= 0.0: return _best_play(1.0) # expert: full threat model, exact argmax var vals: Array = [] for card in legal: vals.append(_net_value(card, s)) # sophistication level scales the threat weight return legal[_softmax_pick(vals, t, rng)]
Legal plays are cards that won't push the running count over 31.
Each legal card is evaluated by _net_value(), and the results are fed
through the same softmax used for discards.
The Net Value Function: What Makes a Peg Card Good
func _net_value(card, sophistication: float) -> float: var value: int = card.get_cribbage_value() var new_count: int = _o.running_count + value var temp: Array = _o.played.duplicate() temp.append(card) var score: float = 0.0 if new_count == 15: score += 2.0 # fifteen-for-two if new_count == 31: score += 2.0 # thirty-one score += float(_score_pegging_pairs(temp)) # pairs, pairs royal, double pairs royal score += float(_score_pegging_run(temp)) # runs of 3, 4, 5... # Subtract: expected opponent reply (card-counted, weighted by sophistication) if sophistication > 0.0 and new_count < 31: var threat_w: float = sophistication # Endgame-aware: if opponent is close to winning, replies hurt 1.5x more if sophistication >= 0.85 and int(_o.scores.get("player", 0)) >= 115: threat_w *= 1.5 score -= threat_w * _opponent_threat(new_count, temp) # Heuristic penalties if new_count == 5 or new_count == 21: score -= 1.0 # leaves easy 15 or 31 score += float(new_count) * 0.02 # prefer playing higher counts if _o.running_count == 0 and card.rank == 5: score -= 1.2 # don't lead a 5 if sophistication > 0.0 and new_count >= 22 and new_count < 31: score += sophistication * 0.10 * float(new_count - 21) # like being close to 31 return score
This function combines several independent signals:
| Signal | Points | Explanation |
|---|---|---|
| Fifteen-for-two | +2 | Immediate score |
| Thirty-one | +2 | Immediate score |
| Pair / run | +2 to +12 | Immediate score from pegging combos |
| Opponent threat | −0..−N | Expected reply score, scaled by sophistication |
| Leaving count 5 or 21 | −1 | Hands opponent easy 15 or 31 |
| Leading a 5 | −1.2 | Classic beginner trap |
| Count closeness to 31 | +0..+0.9 | Hard/Expert prefer being close to 31 |
| Raw count | +0..+0.62 | Tie-break: slightly prefer higher counts |
The sophistication parameter (which equals the skill value during pegging)
directly scales the threat weight. At Easy (0.12), the threat term contributes only
12% of its full strength. The AI barely thinks about what the opponent will
play back. At Expert (1.0), the full threat is applied plus the endgame multiplier
when the player is near 121.
⚠ 7 Opponent Threat Model
At higher skill levels, the AI doesn't just count its own immediate points. It estimates what the opponent is likely to score in reply. This is the "card counting" part of the AI.
The Unseen Rank Pool
The AI builds a probability table of how many copies of each rank are still unaccounted for, not in its hand, not in the public pile, not the starter card.
func _unseen_rank_counts() -> Dictionary: if _rc_for == _o and _o != null: return _rc_cached # memoized per observation -- computed once per decision var counts: Dictionary = {} for r in range(1, 14): counts[r] = 4 # start: 4 of every rank for c in _o.own_hand: counts[c.rank] = int(counts[c.rank]) - 1 # subtract own cards for c in _o.played: counts[c.rank] = int(counts[c.rank]) - 1 # subtract public pile if _o.starter != null: counts[_o.starter.rank] = maxi(0, int(counts[_o.starter.rank]) - 1) # subtract starter _rc_for = _o _rc_cached = counts return counts
The starter subtraction (added in AI-8) is an important fairness detail: the starter card is public knowledge at a real table, and including it in the card count gives the AI a small but legitimate advantage over naive play. Not subtracting it would be leaving information on the table a skilled human would use.
Threat Calculation
func _opponent_threat(new_count: int, seq: Array) -> float: var rank_remaining: Dictionary = _unseen_rank_counts() var unseen_total: int = 0 for r in rank_remaining: unseen_total += int(rank_remaining[r]) var opp_cards: int = _o.opp_card_count if opp_cards <= 0 or unseen_total <= 0: return 0.0 var threat: float = 0.0 for r in rank_remaining.keys(): var copies: int = int(rank_remaining[r]) if copies <= 0: continue var pts: int = _reply_points(int(r), new_count, seq) # points if opp plays this rank if pts <= 0: continue var weighted: float = float(pts) * _prob_opponent_holds(copies, unseen_total, opp_cards) if weighted > threat: threat = weighted # worst-case reply, not expected reply return threat
Notice that the function returns the maximum weighted threat across all ranks, not the sum. This is a deliberate design choice: the AI plays defensively against the worst plausible reply, which is more conservative than averaging. It makes the AI avoid "Russian roulette" pegs that hand over a rare but crushing reply.
Probability of Holding a Rank
func _prob_opponent_holds(copies: int, pool: int, hand_size: int) -> float: if copies <= 0 or hand_size <= 0: return 0.0 if copies >= pool: return 1.0 var p_none: float = 1.0 for k in range(hand_size): p_none *= float(pool - copies - k) / float(pool - k) if p_none <= 0.0: return 1.0 return 1.0 - p_none # P(holds at least one copy) = 1 - P(holds zero copies)
This is the hypergeometric probability: given a pool of pool unseen cards,
copies of which are the target rank, and the opponent holds hand_size
of them. What's the probability they have at least one? The complement approach
(1 minus P(holding zero)) is exact and efficient.
Endgame Defense (AI-10)
One special case: when the player is within 6 holes of winning (score ≥ 115), Hard and Expert ramp up their defensive weight by 1.5×. The line:
if sophistication >= 0.85 and int(_o.scores.get("player", 0)) >= 115: threat_w *= 1.5
The AI tightens up in the stretch run. And critically, scores is public
information. It's on the board for everyone to see.
🔢 8 ScoreEngine: The Math
All scoring lives in core/ScoreEngine.gd: pure static functions,
no state, no UI. There are two versions of each scorer: a full event-producing
version for the Show UI, and a fast allocation-free version for EV inner loops.
Finding Fifteens
static func _find_fifteens(cards: Array) -> Array: var events: Array = [] var n: int = cards.size() for mask in range(1, 1 << n): # enumerate all 2^n - 1 non-empty subsets var sum: int = 0 var combo: Array = [] for i in range(n): if mask & (1 << i): # bit i set = card i is in this subset sum += cards[i].get_cribbage_value() combo.append(cards[i]) if sum == 15 and combo.size() >= 2: events.append({"type": "fifteen", "points": 2, ...}) return events
For 5 cards, this checks 31 subsets (2&sup5; − 1). Face cards count as 10, aces as 1. Every combination that totals 15 with at least 2 cards scores 2 points.
Fast Scorer (EV Inner Loop)
The fast version avoids allocating arrays for every subset by using a PackedInt32Array and scanning runs directly:
# Runs: consecutive-rank streaks; each streak of length L>=3 scores # L * (product of counts) -- identical to enumerating the cartesian combos. var streak_len: int = 0 var streak_mult: int = 1 for r in range(1, 15): if r <= 13 and rc[r] > 0: streak_len += 1 streak_mult *= rc[r] # multiply by copies of this rank else: if streak_len >= 3: pts += streak_len * streak_mult # double runs, triple runs etc. arise naturally streak_len = 0 streak_mult = 1
The run scoring formula is elegant: streak_length * product_of_rank_counts
correctly handles double runs (a run of 4 where one rank appears twice) and triple
runs without enumerating all combinations explicitly.
The Perfect 29
static func is_perfect_29(hand: Array, starter) -> bool: if starter.rank != 5: return false var has_jack := false var five_count := 0 var jack_suit := -1 for c in hand: if c.rank == 11: # Jack has_jack = true jack_suit = c.suit elif c.rank == 5: five_count += 1 return has_jack and five_count == 3 and jack_suit == starter.suit
The rarest hand in cribbage: J♦ + 5♠5♣5♥ + 5♦ starter (Jack matching the cut's suit). Exactly one combination out of the ~12 billion possible cribbage deals.
📈 9 Crib Odds Tables
Rather than simulating the full crib completion at runtime for every discard candidate, the AI reads from pre-computed tables generated offline by exact enumeration.
# Computed by tools/gen_crib_table.py via exact enumeration of all 58,800 crib completions const AVG_CRIB := [ # A 2 3 4 5 6 7 8 9 10 J Q K [5.53, 4.45, 4.57, 5.47, 5.74, 4.26, 4.09, 4.13, 4.04, 3.96, 4.20, 3.86, 3.75], # A ... [5.74, 5.77, 6.43, 7.00, 8.99, 7.10, 6.42, 5.76, 5.74, 7.03, 7.26, 6.93, 6.82], # 5 ]
The 5-5 entry (row 5, col 5): 8.99 average crib points. That's why throwing a pair of 5s into your own crib is so powerful, and why the AI is reluctant to let a pair of 5s go into the opponent's crib as pone.
The basic AVG_CRIB table assumes a uniform opponent completion. But real
cribbage has seat effects: the dealer completes their crib greedily, while the
pone defensively throws their worst crib cards into the dealer's crib. The AI
uses two additional tables, AVG_CRIB_DEALER and AVG_CRIB_PONE
, generated by policy-conditioned self-play (n=3,000 per rank pair) to capture this
asymmetry. The 5-5 dealer table entry is 8.85; the pone table entry is 9.10, higher because the dealer completes aggressively.
The lookup is a single array access at decision time: O(1), no math, no simulation.
static func crib_value(rank_a: int, rank_b: int, my_crib: bool) -> float: var a := clampi(rank_a, 1, 13) - 1 var b := clampi(rank_b, 1, 13) - 1 if my_crib: return AVG_CRIB_DEALER[a][b] # seat-conditioned: opponent completes defensively else: return AVG_CRIB_PONE[a][b] # seat-conditioned: opponent completes greedily
🎓 10 Coach & Hints (CoachEval)
The in-game hint system and post-game coaching grade the player's decisions using the exact same model the AI uses for itself. One source of truth; the game never holds the player to a higher standard than it holds the computer.
Live Grade: Blunder Detection
func grade_play(own_hand: Array, pile: Array, count: int, opp_count: int, chosen, starter = null) -> Dictionary: var obs := AlleyObservation.for_play(own_hand, pile, count, opp_count, starter) var best_val: float = _brain.peg_best_value(obs) var chosen_val: float = _brain.peg_play_value(obs, chosen) var ev_lost: float = maxf(0.0, best_val - chosen_val) # Named trap detection var led_five: bool = (count == 0 and chosen.rank == 5) var left_danger: bool = new_count in [5, 10, 11, 21, 22] return { "best": ev_lost <= PEG_BEST, # within 0.05 of optimal "ok": ev_lost < PEG_SLACK, # within 0.35 -- no nagging "ev_lost": ev_lost, "best_card": best_card, "reason": ..., "tags": tags, }
The slack constants matter. PEG_SLACK = 0.35 means if you give up less than
0.35 expected peg points, the coach says nothing. Pegging deltas are small and
noisy and the coach is designed not to nag over coin-flip decisions.
Named Traps: Positive Coaching
The coach also recognises specific tactical formations in the player's kept hand and calls them out positively, teaching technique, not just scoring blunders.
| Trap Name | Formation | How it works |
|---|---|---|
| Five-card trap | 4-6-6 | Lead a 6; they can't safely answer with a 5 (6-5-4 run threat); you force 31 |
| Pair-royal bait | Low pair (A-4) | Lead one; if they pair it, drop the third for pairs royal (6 points) |
| Magic eleven | Two cards summing to 11 | Any ten-card in the reply lets you seize 31-for-2 |
| Run bait | Three consecutive ranks | Offer one end; hold the extension to re-trap into a longer run |
Deep Analysis: Post-Game Expectimax
The opt-in deep analysis after the game runs a full expectimax simulation over the opponent's possible hands, averaging across sampled completions to find leads that score nothing immediately but create a structural peg advantage: traps the greedy grader can't see.
func analyze_lead(player_cards: Array, unseen_pool: Array, opp_count: int, chosen, max_samples: int = 120) -> Dictionary: var samples: Array = _sample_hands(unseen_pool, opp_count, max_samples) for oc in samples: for c in player_cards: totals[c] += _play_recurse(player_cards, oc, [], 0, P, c) # player maximises # opponent inside _sim() plays fair greedy -- AlleyBrain.peg_best_card(obs)
The simulation samples up to 120 opponent hands from the unseen pool (or fewer if C(pool, k) ≤ 120, in which case it enumerates exactly). For each sample it computes the net peg margin from best play by both sides through the remainder of the phase. A "trap" is flagged when the best lead scores nothing immediately yet nets ≥1.5 margin over the sample, and the chosen lead underperformed by ≥1 point.
🔐 11 FairDeal: Cryptographic Fairness
FairDeal is the anti-cheat backbone: it proves that neither the game nor the AI could have "re-rolled" the deal after seeing the cards, and that every hand was predetermined before a single card was shown.
Commit-Reveal Protocol
func begin_match(master_seed: int, client_seed_in: String = "") -> String: client_seed = client_seed_in var ctx := HashingContext.new() ctx.start(HashingContext.HASH_SHA256) ctx.update(_seed_bytes(master_seed)) ctx.update(("|%s|server" % client_seed).to_utf8_buffer()) server_secret = ctx.finish() commit_hex = _sha256_hex(server_secret) # this goes to the player BEFORE dealing return commit_hex func hand_seed(hand_index: int) -> int: var ctx := HashingContext.new() ctx.start(HashingContext.HASH_SHA256) ctx.update(server_secret) ctx.update(("|%s|%d" % [client_seed, hand_index]).to_utf8_buffer()) var d: PackedByteArray = ctx.finish() var s: int = 0 for i in range(8): s = (s << 8) | int(d[i]) return s & 0x7FFFFFFFFFFFFFFF # 63-bit positive seed
At match start you see commit = SHA256(secret). SHA-256 is a one-way function:
nobody can find the secret from the commit. So the entire deal tree, every hand,
every shuffle, every AI decision in ranked mode, was locked in before any card was
dealt. At match end the secret is revealed and you (or any verifier) can replay every
hand from scratch using verify_hand_seed() and confirm the cards match.
Ranked Mode: Hash Rolls for AI Decisions
static func hash_roll(seed_val: int, tag: String, n: int) -> float: # roll(seed, "aid", 0) for discard; roll(seed, "aip", n) for nth pegging play # Returns a uniform float in [0, 1) derived from SHA-256 var ctx := HashingContext.new() ctx.start(HashingContext.HASH_SHA256) ctx.update(b) # seed as big-endian 8 bytes ctx.update(("|%s|%d" % [tag, n]).to_utf8_buffer()) var v: int = 0 for i in range(7): v = (v << 8) | int(d[i]) v = v >> 3 # 53 bits return float(v) / 9007199254740992.0 # / 2^53
In ranked PvE sessions, the AI's softmax draws come from this function instead of
the engine RNG. The server's replay verifier calls the same formula
(mirrored in functions/src/ai.ts) and gets byte-identical results, proving the AI played at the claimed skill level and didn't "get lucky" with
unusual RNG draws.
📊 12 Strength Benchmark
The AI's strength is not asserted. It is measured. The benchmark runner
in tests/bench_ai.gd plays the shipped AlleyBrain against itself at
different skill levels over thousands of matches and validates the results against
hard win-rate targets.
# AI-5 spacing targets -- the weaker tier's win rate vs Expert should land near: # Easy 2-5% (expert wins 95-98%) # Normal 15-25% (expert wins 75-85%) # Hard 35-45% (expert wins 55-65%) # Outside those bands => retune the _temperature() anchors in core/AlleyBrain.gd. const N_MAIN := 10000 # expert vs random -- headline sanity check const N_LADDER := 3000 # expert vs each weaker tier
The benchmark alternates which seat each policy sits in and which player deals first, so neither seat-advantage nor dealer-advantage biases the result. It runs 10,000 matches for the headline number and 3,000 for each ladder rung.
func _wilson(k: int, n: int) -> Array: var z := 1.96 # z* for 95% CI var p := float(k) / float(n) var denom := 1.0 + z * z / float(n) var centre := (p + z * z / (2.0 * float(n))) / denom var margin := (z * sqrt(p * (1.0 - p) / float(n) + z * z / (4.0 * float(n) * float(n)))) / denom return [(centre - margin) * 100.0, (centre + margin) * 100.0]
The results write to data/ai_benchmark.json which the in-game "AI Strength"
panel reads, so the numbers displayed in-game are always from the latest actual
benchmark run, not hardcoded estimates.
🔥 13 Can You Trap the AI?
Yes, and the code explains exactly when it works and when it backfires. The AI's pegging model is a one-ply lookahead with a card-counted threat term. That means it catches obvious immediate replies but has no explicit multi-card forward planning of its own. Humans who play multi-play deceptions can exploit this gap, if the skill level is low enough.
AlleyBrain evaluates each legal card with _net_value(), which looks one
opponent reply ahead. It has no explicit lookahead for the second or
third play. So traps that take two or three plays to pay off can work even at Hard, though the endgame defense and threat weighting make them
harder to land than on Easy.
Trap-by-Trap Breakdown
_prob_opponent_holds() gives non-zero probability, and the 6-point pairs royal reply suppresses the AI's pairing instinct on high-skill tiers.kept_traps() specifically calls out.score += sophistication * 0.10 * float(new_count - 21). This can be exploited: if you know you hold the card that hits 31, invite the AI to push the count into the mid-20s by playing your safe cards first. The AI obliges (it wants that range), and then you close it out.running_count = 0 and a fresh pile. It recalculates from scratch, including a -1.2 penalty for leading a 5. If you held your trap cards back across the go, you can re-deploy them into the new sub-round with the AI having no memory of the previous sequence's structure.What the AI Defends Against
The code explicitly penalises several of the most common player attacks:
| Attack | Penalty in Code | Effect |
|---|---|---|
| Leading into count 5 or 21 | score -= 1.0 for leaving danger count |
AI avoids handing you an easy 15 or 31 |
| You lead a 5 | if count == 0 and card.rank == 5: score -= 1.2 |
AI actively avoids leading 5, which blocks mirror pressure |
| Pair bait (you hold the third) | _prob_opponent_holds() * pairs-royal reply pts |
At Hard/Expert, the threat weight includes your probable holding of the third rank |
| You're close to pegging out | threat_w *= 1.5 at scores ≥ 115 |
Hard/Expert plays much tighter defense in the endgame |
| Immediate reply scoring | Full _opponent_threat() at full sophistication |
Expert almost never hands you a straightforward 15, pair, or run |
If you attempt a trap that the AI should have fallen into but didn't, the
post-game deep analysis (analyze_lead()) will notice. It simulates
the same situation with expectimax and reports whether your lead was objectively strong.
If it was the best line but the AI dodged it by luck, the analysis confirms
that your lead was correct and the AI got lucky, not that you blundered.
★ 14 Master Strategies & How the Code Responds
These are the techniques advanced cribbage players use. For each one, here's what the code actually does when you employ it, or when the AI does.
Discard Strategy (The Show)
| Technique | What Masters Do | How the Code Handles It |
|---|---|---|
| Crib awareness as pone | Avoid throwing 5s, pairs, or suited connectors into the dealer's crib | The AI uses AVG_CRIB_PONE which is the seat-conditioned table where the
dealer completes the crib greedily. The AI's score = net as pone
directly penalises dangerous throws. The 5-5 pone entry (9.10) makes a pair of
5s extremely costly to toss away. |
| Feed your own crib | As dealer, throw cards that combine well with any completion | score = expected = avg_hand + avg_crib as dealer. The seat-conditioned
AVG_CRIB_DEALER reflects a defensive opponent completion, so the AI
doesn't over-trust paired throws into its own crib. It uses realistic averages. |
| Safety discard near-ties | When two keeps are nearly equal in EV, throw the less dangerous pair | _cmp_disc() uses a 0.001 epsilon tiebreak on safety. The AI already does this. |
| Keep connected cards | Prefer keeping runs and pairs that maximise cut potential | The 46-starter EV average across _avg_hand() automatically favours connected
cards because more of the 46 starters complete runs or pairs with them. |
| Sacrifice hand for crib | As dealer, throw a scoring pair into your crib even if it weakens the hand | The EV formula sums hand + crib explicitly. The AI will sacrifice 2 hand points to gain 3 expected crib points. The math tells it to. As pone, it would never do the same. |
Pegging Mastery
| Technique | Master Play | Code Behaviour |
|---|---|---|
| Card counting | Track which ranks are gone; adjust probability of opponent replies | The AI does this exactly via _unseen_rank_counts(). It subtracts
its own hand, the public pile, and the starter. At Expert this is used at full weight. |
| Endgame desperation pegging | When close to winning, peg aggressively even if handing back points | Scores are public (in AlleyObservation.scores). Expert at score ≥115
increases threat weight by 1.5x, meaning the AI pegs more conservatively
against you, not more aggressively. The player who's behind should be the
one taking risks. |
| Do not lead a 5 | Classic rule: leading a 5 hands the opponent 15-for-2 with any 10-value card | if _o.running_count == 0 and card.rank == 5: score -= 1.2. The penalty is 1.2 points, enough to push it below virtually any alternative lead.
Expert almost never leads a 5. |
| Prefer safe leads | Lead low (A-4) to limit the opponent's reply options | The AI doesn't have an explicit "lead low" heuristic, but the threat model achieves the same result: low cards leave fewer dangerous reply ranges, so their net value is higher at high skill. Leading a 4 scores nothing but leaves fewer immediate replies that score points. |
| Count management | Keep the count in ranges where you hold the winning card | The AI's score += float(new_count) * 0.02 tiny bonus for higher counts is
a weak proxy for this. True count management requires knowledge of your own hand
. The AI uses the card-count pool for it. Humans who know exactly which counts
their hand can reach have an edge the AI's probability model only approximates. |
| Sacrifice to reset | Say "go" deliberately to reset the count when holding a powerful combo | The AI cannot deliberately say go. It plays a card if any legal play exists, even a low-value one, because passing is only permitted when no card plays under 31. This means a human who engineers a count where the AI is forced to go, then re-leads into a trap, is playing a depth the AI's one-ply model doesn't explicitly plan around. |
| Flush your dangerous cards early | Play your high-risk cards (5s, cards that form 15s) first to avoid being stuck | The AI penalises leaving 5 and 21 on the count because those hands the next player something. Humans who think one step further ("if I play this now, what counts do I leave on my next turn?") are playing a deeper lookahead than the AI explicitly models. |
The Trap the AI Is Most Blind To
The biggest structural gap in the AI is multi-play sequence planning.
AlleyBrain values each card play independently via _net_value() and looks
only one opponent reply deep. A master human player thinks in sequences:
If you hold 7-8-9 and lead the 7: the AI checks whether replying with something that hits 15 or pairs is good. Say it plays a 6 (count=13). You drop the 8 (count=21). Now the AI faces count 21 and it knows -1.0 applies to leaving 22, so it avoids playing a card that brings it to 22. But if you hold the 2 (total=23, no score) you've neutralised its defence and your 9 later scores the 32 reset. The AI's one-ply window cannot anticipate the whole line. At Expert it resists better than at Easy, but the structural limit remains.
What the AI Does Better Than Most Humans
It's not all in the human's favour. The AI at Hard/Expert is better than the average player at several things:
| Skill | Why the AI Is Better |
|---|---|
| Exact EV discard calculation | Humans estimate crib odds intuitively; the AI has 58,800-completion exact tables. It never forgets that 5-J (same suit) is better than 5-Q. |
| Seat-conditioned crib defence | Many players throw suboptimal cards into the opponent's crib. The AI uses the
AVG_CRIB_PONE table and always knows the true expected cost. |
| Consistent threat weighting | Humans get tired, miss pair threats, forget how many of a rank have been played. The AI never does. It maintains exact card counts every play. |
| No tilt | The AI doesn't play differently after losing five hands in a row. Its temperature is fixed by difficulty tier and never changes based on game history. |
⚡ Key Takeaways
This is not a policy. It is a type-system enforcement. AlleyObservation
has no field for the opponent's hidden cards. AlleyBrain is a pure function
of an AlleyObservation. The test suite confirms this structurally.
Expected-value discard over 46 starter cards. Card-counted pegging with threat modeling. Seat-conditioned crib tables. Endgame defense when you're close to 121. Leading-5 penalty. Magic-eleven awareness. These are the same concepts a strong human player uses.
The softmax model produces human-plausible errors at lower difficulties: near-misses and occasional real blunders weighted by how much EV they give up. No tier ever plays the single worst option unless the hand is genuinely close. The difficulty targets are validated by 10,000-match benchmarks.
Every deal is committed before any card is shown via SHA-256 commit-reveal. You can verify every hand at match end. In ranked mode, every AI decision is a deterministic SHA-256 draw the server can replay independently.
| File | Role | Lines |
|---|---|---|
| core/AlleyBrain.gd | Complete decision engine (discard + peg) | ~490 |
| core/AlleyObservation.gd | Data contract (no-peek enforcement) | 42 |
| core/ScoreEngine.gd | Pure scoring, fast + full-event variants | ~240 |
| core/PeggingController.gd | Pegging state & AlleyBrain bridge | ~155 |
| core/CoachEval.gd | Hint/grading system & deep analysis | ~366 |
| core/FairDeal.gd | Commit-reveal fairness & ranked hash-rolls | ~160 |
| scenes/stats.gd | Crib tables, par benchmarks, stat tracking | ~400 |
| tests/bench_ai.gd | 10k-match strength benchmark | ~166 |
Generated from live source code for Just Cribbage (Real Cribbage), July 2026. All code quotations are verbatim from the shipped GDScript files.