Newest first — every version bump below reflects an actual change to how SAM predicts, logs, or displays picks.
WTA 12-FACTOR IMPORTANCE WEIGHTS. The women's model now runs on a 12-factor weighted composite instead of the hold-%-is-the-engine stack — twelve individual factors, each weighted on its own: Return Points Won % 21, Break Points Converted % 17, Service Points Won % 10, Return Games Won % 9, Break Points Saved % 8, 1st Serve Points Won % 8, 2nd Serve Points Won % 8, Surface Win % 8, Service Games Won % 5, Double-Fault Rate 2%, H2H 2%, Ranking Differential 2% — 100% total. MECHANICS: wtaWeightedCompositeEdge computes one composite edge E per matchup (each factor a head-to-head differential normalized by 2x the 2026 tour standard deviation, clamped to +/-1, double faults inverted); E feeds the Monte Carlo loop directly — every simulated service game and tiebreak draws against edge-adjusted holds (+/- E x 0.05 per side), so the 12 factor weights bank against fresh noise game by game instead of pre-shifting the hold inputs — a max edge moves each hold 5pp in opposite directions, typical matchups 1-2pp per side. DATA: the WTA's own free stats API (api.wtatennis.com, no key) — one paginated sweep builds the full 2026 tour table, cached 6h in the Cache API, name-matched per player, minimum 5 matches to count. MISSING-DATA RULE: any missing input goes neutral (0), never a fabricated default — if either player's WTA row is missing, the composite stays off and the standard hold path runs untouched. NO-DATA GATE (all tennis): when neither player has verifiable data (no rank, no season W/L on either side), the matchup is declined as unable to process instead of simulating on pure estimates — the model either runs on real stats or it doesn't run. WTA and ITF: the composite runs for both tours (active only when both players have WTA tour rows); ATP never enters that path — its hold-% engine, surface logic, and writeups are unchanged. The WTA writeup now leads with the return-game framing. Untouched: 0.40 shrinkage, down-only band calibration, 50/50 rule + caps, best-of-3 for WTA, model 1.6.1. Covered by tests/wta-importance-weights.test.js and tests/tennis-no-data-gate.test.js. Full suite green.
CONSERVATISM SHRINKAGE SET TO 0.40 (0.55 → 0.40). The daily shrinkage watch's drift note showed the empirical optimum moved to 0.40 (90% noise band 0.35-0.47), with the 0.55 setting sitting outside the data's noise band — the sim's edges were being over-flattened into coin flips. Easing shrinkage keeps more of the sim's lean (a raw 70% edge reports ~58% instead of ~61%, a raw 80% edge ~62% instead of ~67%). All deflationary mechanisms stay in place: down-only band calibration, 50/50 rule + caps, grading, branding, model 1.6.1.
MANUAL ANNOUNCEMENT BLASTS. New "announcement" alert type plus an admin /admin/send-announcement endpoint (admin-key authed, title/body/url params): Drew can now push any message he wants to every push-subscribed user on demand. Every blast still flows through the single dispatchAlert choke point — per-user prefs (new "Announcements from ParlayKings" toggle in notification settings, default on), quiet hours, the 12/day cap, and 24h dedupe on identical messages (re-sending the same text within a day is a no-op, not a double blast). Announcements respect quiet hours — only game-day score alerts bypass them, by design. Covered by tests/alerts.test.js. Untouched: 0.55 shrinkage, down-only band calibration, 50/50 rule + caps, grading, model 1.6.1.
FOOTBALL 15-FACTOR IMPORTANCE WEIGHTS. Drew's exact importance weights are now the football model — 15 factors, 100% total: EPA/efficiency 17%, QB factor 15%, Passing EPA/dropback 12%, Success rate 10%, Pressure/protection 9%, Turnover differential 8%, Explosive plays 7%, Points per drive 5%, Red-zone TD 4%, Early-down EPA 3%, 3rd/4th down 3%, Rush EPA 2%, Sack rate 2%, Special teams 1.5%, Penalty efficiency 1.5%. MECHANICS: footballWeightedCompositeEdge computes one composite edge E per matchup (each factor a head-to-head expected-points edge, clamped to +/-6 before weighting, normalized vs league average); bridgeToTeamSim applies it additively as +/- E/2 on the two Poisson lambdas before the realistic-range clamp. The old tiered-hierarchy multiplicative stack is bypassed for NFL/CFB — when the composite is on, the QB, RB, WR, O-line, pass-rush, kicker, efficiency, and turnover-margin multiplier slots all go neutral (exactly 1.0) inside calculateLambdaMultiplier so nothing is double-counted; rest, travel, form, weather, injuries, and SOS are untouched, and other sports never set the flag. GONE from the sim path: the standalone RB/WR multiplicative factors (not on Drew's list — team rushing lives in Rush EPA at 2%). DATA: NFL EPA table (nflverse play-by-play, all 32 teams) is embedded in the bundle and refreshed weekly by nfl_epa_refresh.py — EPA, success rate, passing/dropback EPA, rush EPA, early-down EPA, sack rates, special-teams EPA, points per drive, all shrunk by gp/8 early in the season. The QB factor (individual QB vs league average: CPOE/accuracy, decisions, sack avoidance, pressure play, recent form) and Passing EPA (the whole dropback ecosystem: QB, receivers, protection, scheme, opposing pass D) stay separate inputs at Drew's combined 27% — overlap acknowledged, not increased without backtesting. CFB runs the same 15-factor structure with the advanced EPA factors neutral until a real CFB EPA feed lands — the kicker factor is its special-teams proxy; nothing is faked. MISSING-DATA RULE: any missing input goes neutral (0), never a fabricated default — one-sided missing data produces no phantom edge. Untouched: 0.55 shrinkage, down-only band calibration, football always-picks-a-winner, neutral-site HFA behavior, 50/50 rule + caps for other sports, model 1.6.1. Covered by tests/football-importance-weights.test.js (102 checks: exact weights and 100% sum, 32-team EPA table sanity, neutral missing-data behavior, reversal symmetry, QB/passing separation, exact 4x early-season shrinkage, per-factor clamp, composite-on neutralization with legacy path intact). Full suite green.
UFC 15-FACTOR IMPORTANCE WEIGHTS. Drew's exact importance weights are now wired into the Monte Carlo loop — 15 factors, 100% total: Significant Strike Differential 16%, Opponent-Adjusted Performance 15%, Takedown Defense 10%, Recent Performance / Trajectory 10%, Significant Strike Defense 8%, Significant Strikes Absorbed/min 7%, Strength of Schedule 7%, Knockdown Rate 6%, Takedowns/15 5%, Takedown Accuracy 4%, Round-by-Round Performance 4%, Significant Strike Accuracy 3%, Submission Attempts/15 2%, Control-Time Differential 2%, Finish Rate 1%. Drew's five emphasized inputs carry 59%. MECHANICS: ufcWeightedCompositeEdge computes one composite edge E per matchup from the 15 normalized factors; each simulated round banks E * 0.75 (UFC_ROUND_EDGE_SCALE) against noise. GONE: the old strike-differential kicker (0.22 edge, x2.4 into the decision), the 1.15 takedown control credit, and the 0.98-1.02 form volume multiplier. KEPT: the KO/sub/TD event rolls at their old math (takedown prob multiplier 2.2, sub-finish coefficient 0.08, KO chance on landed strikes x0.035), the last-five exact-tie tiebreak (|raw-50| < 0.5), and the CANNOT_VERIFY gate. NEW DEEP STATS from ufcstats.com fight-detail pages (last 3 fights): knockdown rate per 15, control-time differential per round, round win % on significant strikes, and a trajectory metric (newest fight vs older average) blended 70/30 into the Recent Performance factor. FightMatrix factors (opponent-adjusted 15%, strength of schedule 7%) are neutral 0 until the FightMatrix feed lands — nothing is faked. MISSING-DATA RULE: no fabricated defaults — when SApM is missing on either side, both the SApM factor and the strike-differential factor go neutral instead of inventing a 95%-of-SLpM phantom. Untouched: 0.55 shrinkage, down-only band calibration, 50/50 rule + caps, the tale-of-the-tape prose, model 1.6.1. Covered by tests/ufc-importance-weights.test.js (28 checks: all 15 keys wired, weights sum 100%, top-5 = 59%, FightMatrix neutral, SApM phantom fix, single-factor isolation, deep-parser fixtures, 3/5-round sims), tests/ufc-form-demoted.test.js (form as the real 10% factor), tests/ufc-tiered-hierarchy.test.js (composite structure, event rolls, physicals capped at +/-0.10), plus updated ufc-tiebreak-and-rounds and ufc-tale-of-the-tape. Full suite: 842 passed, 0 failed.
AUTO-GRADE DATE-MATCH HOLE CLOSED. findFinalResult's exact Event Date check previously only applied when multiple ESPN events matched a pick's matchup — a single candidate was graded without any date verification. This morning's run-all proved the hole: ten 9/24 MLB picks were graded against 9/23's finals (same matchups, series continuing) before any 9/24 game was played. Now a single candidate must also match the pick's Event Date exactly, otherwise the pick stays Pending — no guessing, single or multiple. The ten bad grades were wiped back to Pending in Airtable and the replica sheet. Untouched: 0.55 shrinkage, down-only band calibration, 50/50 rule + caps, model 1.6.1.
BASEBALL TIERED HIERARCHY. Final baseball importance order, enforced by per-tier clamps with strictly decreasing max swing: TIER 1 = the starting pitcher as the single dominant weight — the QB of baseball — clamp widened 0.86-1.16 to 0.70-1.40 so an elite ace moves the number like the football QB; ERA regression toward 4.2, the sample gate (w<0.25 neutral), and starter-lock timing logic are untouched. TIER 2 = the bullpen with real weight: (a) FIXED the confirmed quirk where bullpen ERA was shrunk toward the 4.2 starter anchor so every elite bullpen flattened to 0.86 — relievers now regress toward their own 3.90 anchor with a tier-2 clamp of 0.80-1.25; (b) the full-lock pitching blend rebalanced 62/38 to 58/42 starter/bullpen (staff proxy 78/22 to 75/25), so the pen carries genuine tier-2 weight while the starter keeps the majority. TIER 3 = the bats — team runs scored stays the offensive engine, unchanged. TIER 4 = infielders sampled INDIVIDUALLY for the first time (the model previously knew zero batters — no lineups at all): each team's starting SS, 1B, and top 2B/3B (pooled, one combined factor) identified from the live roster by at-bats (min 100, regular-season only, OUT players skipped) and factored by OPS in Drew's positional order — SS 0.88-1.12 > 1B 0.90-1.10 > 2B/3B 0.92-1.08; Drew's caveat honored (bat over glove, OPS-driven). TIER 5 = home field (+0.15), rest, weather, travel, recent form — unchanged and smallest. Each tier rides its own lambda multiplier slot (starter+pen blend stays in the pitcherFactor slot suppressing the opponent; new ssFactor/firstBaseFactor/secondThirdFactor slots boost a team's own expected scoring). Untouched: 0.55 shrinkage, down-only band calibration, bivariate Poisson, NO 50/50 for MLB (Drew's standing call — baseball always picks a winner), neutral-site HFA behavior, model label SAM 1.6.1. Still open (flagged, not approved): neutral-site HFA gap, changelog-vs-code doc mismatches (BDL 50/50 vs 65% prose, rho 0.02 vs 0.28 note, recent-form exclusion note), park factors, platoon splits, umpires, bullpen workload. Covered by tests/baseball-tiered-hierarchy.test.js. Model 1.6.1.
FOOTBALL TIERED HIERARCHY. Final football importance order, enforced by per-tier clamps with strictly decreasing max swing: TIER 1 = the QB as the single dominant named-player weight, now clamped 0.70-1.40 (was 0.80-1.25 — like the pitcher in baseball, the QB gets the highest weight in the game); TIER 1 also = the lead back by rushing YPC (0.88-1.12) and the WR1 by receiving YPR (0.90-1.10), sampled individually from each roster by actual usage; TIER 2 = the offensive line via sacks allowed, rush YPC, and yards per play (0.92-1.08); TIER 3 = the pass rush via the OPPONENT's sacks created, TFL, and turnovers forced (0.93-1.06); TIER 4 = the kicker as a standalone meaningful input via the actual kicker's season FG% (0.95-1.04 — a good kicker can swing a coin-flip game); TIER 5 = the remaining team-efficiency bits (3rd/4th down, red zone, INT%, TOP, pace, punts, penalties), demoted from a 0.90-1.10 clamp to 0.96-1.03; TIER 6 = home field, rest, weather, travel, recent form — unchanged and smallest. QB CONSOLIDATION: the QB was triple-counted (team PPG, pass YPA/pass YPG inside the generic efficiency factor, standalone QB factor) — pass YPA, pass YPG, and yards per game are OUT of the efficiency factor, sacks allowed moved to the O-line factor, opponent sacks/TFL moved to the pass-rush factor, FG/XP bits moved to the kicker factor. Team PPG stays the scoring engine. Each tier rides its own lambda multiplier slot (qbFactor stays in the pitcherFactor slot; new rbFactor/wrFactor/olineFactor/passRushFactor/kickerFactor params on calculateLambdaMultiplier). The skill factors sample real athletes: QB by passing attempts (min 50, existing path), lead back by rushing yards (min 40 attempts), WR1 by receiving yards (min 20 receptions), kicker by kicking points; OUT players are skipped and never boost a team; missing data is neutral. Untouched: 0.55 shrinkage, down-only band calibration, football always-picks-a-winner, neutral-site HFA behavior, kickRetAvg sign quirk, CFB sosDiff (still 0), model label SAM 1.6.1. Covered by tests/football-tiered-hierarchy.test.js (50 checks: QB clamp widening, RB/WR/O-line/pass-rush/kicker factors, efficiency demotion cap, strict tier ladder, elite QB beats good form, elite WR beats max rest, O-line swing beats home field, kicker flips a seeded 50/50 sim, dead-even games still resolve). Tonight's UFC tiered hierarchy, tennis hierarchy, basketball fixes, wiring audits, NASCAR removal, and NHL restoration all verified still present. Model 1.6.1.
UFC TIERED HIERARCHY. Final UFC importance order: TIER 1 = striking trio (striking output SLpM > strike accuracy > strikes absorbed) AND grappling (takedowns/subs) at full strength — control score 1.15 per landed TD, takedown probability multiplier 2.2, sub-finish coefficient 0.08. TIER 2 = physicals: reach, height, stance, age. TIER 3 = everything else: recent form (+/-2% band), finish mix (+/-10%), home-country boost (0.015), weight-class finish biases (all demoted earlier tonight and still demoted). The earlier grappling demotion was REVERTED — grappling sits alongside the striking trio in tier 1, not below it. The striking trio's own math is untouched (SLpM x accuracy strike volume, x0.035 KO scaling, 0.22 strike-differential edge, x2.4 into the decision), as are the CANNOT_VERIFY gate, 0.55 shrinkage, down-only band calibration, UFC always-picks-a-winner behavior, and the prose. Covered by tests/ufc-tiered-hierarchy.test.js (12 checks: tier-1 fighter beats physicals-only fighter 73-27, even tier-1 with physical edge decides 72-28, tier-3 inputs with perfect form and maxed mix only nudge 53-47). Also folded into this entry: (1) LAST-FIVE EXACT-TIE TIEBREAK: last-five form is the very last thing — only when the raw sim split is within |raw-50| < 0.5 (inside Monte Carlo noise) does the better last-5 winPct get the pick; a raw 50.9-49.1 is a decision, not a tiebreak; last-5 tied or missing -> raw higher side stands; applied before shrinkage so confidence shrinks normally (lands 51-49); the prose explicitly states when last-five broke a dead-even sim. (2) EXPLICIT PRELIM ROUNDS: prelim fights are 3 rounds, period — the card-position rule is the authority (title fights and main events go 5; anything not identifiable as title/main-event is 3), ESPN's format.regulation.periods is the explicit scheduled-rounds field, the description '/N Rnd/' parse is only a cross-check now. (3) ROUNDS VERIFIED END TO END: traced fetchUFCFightRounds -> runMatchupSimulationCore UFC branch (roundsToUse = rounds === 5 ? 5 : 3) -> runUFCMultiStatSimulation (safeRounds = rounds === 5 ? 5 : 3, loop 'for (rnd = 0; rnd < safeRounds && !finished; rnd++)' iterates exactly rounds times, no hardcoded 3/5 in the loop path, no off-by-one; fatigue fade and per-round finish rolls scale with the actual round count). Covered by tests/ufc-tiebreak-and-rounds.test.js (21 checks: tiebreak fires on seeded dead-even sim with 5-0 vs 2-3 and picks the 5-0 fighter 51-49, 50.9-49.1 does not invoke form, prelim resolves to 3 even with missing description, title/main-event resolve to 5, 5-round sim shows materially higher finish rate than 3-round). Tonight's basketball fixes, wiring-audit fixes, NASCAR removal, tennis importance weights, UFC form demotion, and UFC tale-of-the-tape all verified still present. Model 1.6.1.
UFC TALE OF THE TAPE ON TOP. The physical tape and striking/grappling stat differentials (reach, height, age, stance, SLpM/accuracy/absorbed, takedown avg/accuracy, submission avg, strike and takedown defense) are now the top tier of the UFC model by construction: everything below them was demoted, nothing above them was boosted. DEMOTED: (1) ufcFinishMixScale compressed from clamp(0.72 + share*0.9, 0.78, 1.32) to clamp(0.90 + share*0.35, 0.90, 1.10) — a fighter with 100% KO history now gets at most a +10% finish-chance modifier instead of +32%. (2) The weight-class finish-bias table compressed halfway toward 1.0 across all classes (e.g. heavyweight KO 1.18 -> 1.09, flyweight KO 0.82 -> 0.91). (3) The home-country strike boost halved from +0.03 to +0.015. UNCHANGED AND FULL STRENGTH: the striking/grappling core math, frame edges (combined clamp +/-0.08), stance edges, the age speed edge and 35+ cliff, the 0.98-1.02 form band (form stays the bottom input), the CANNOT_VERIFY gate, 0.55 shrinkage, down-only band calibration, UFC always-picks-a-winner behavior, and the prose (the tale of the tape keeps printing in full — it is information AND the engine). Covered by tests/ufc-tale-of-the-tape.test.js (extracts the real bundle functions). Tonight's basketball fixes, wiring-audit fixes, NASCAR removal, tennis importance weights, and UFC form demotion all verified still present. Model 1.6.1.
UFC RECENT FORM DEMOTED. Recent form (raw last-5 W/L, no opponent-quality adjustment) is now the BOTTOM of the UFC importance hierarchy — the least important input in the sim. The form multiplier in runUFCMultiStatSimulation was compressed from 0.92 + 0.16 * (winPct/100) (a 0.92–1.08 swing) to 0.98 + 0.04 * (winPct/100) (a 0.98–1.02 swing): a 5-0 fighter now gains at most +2% strike volume and an 0-5 fighter loses at most 2%. Form can nudge a genuinely close matchup but can no longer flip a fight that the striking/grappling differentials decide. Prose is unchanged — the last-5 is still printed in the tale of the tape as information, just demoted in the math. Covered by tests/ufc-form-demoted.test.js (extracts the real bundle function). Untouched: all UFC sim math (striking, grappling, defenses, frame, stance, age, weight-class scale, finish mix, rounds, home-country boost), the CANNOT_VERIFY gate, 0.55 shrinkage, down-only band calibration, UFC always-picks-a-winner behavior, and tonight's basketball fixes, wiring-audit fixes, NASCAR removal, and tennis importance weights — all verified still present. Model 1.6.1.
TENNIS IMPORTANCE WEIGHTS. "Hold serve %, court surface, number of sets" are the three most important inputs — everything else subordinate. The tennis model now enforces that hierarchy in tennisServeHoldProb and the sim writeup. (1) HOLD % is the engine: real serve tape is used directly as the hold probability, untouched and dominant. (2) SURFACE is the co-primary driver, rules unchanged: grass +0.04 / clay -0.04 / hard 0 on hold probability plus the player's own surface W-L edge at 0.15 weight (5+ surface matches, delta capped at +/-25pp), total surface contribution clamped at +/-0.06, and NO modifier on top when real surface-specific hold tape already captures the surface effect (no double-counting). (3) SET FORMAT compounds structurally: verified that runTennisMatchSimulation already implements best-of-N as real set-by-set probability math (first to 2 sets in BO3, first to 3 in BO5) with no fudge factor — the stronger server's edge genuinely grows in BO5 (a 0.85-vs-0.78 hold edge gains ~5pp of match win probability going BO3 to BO5). ATP Slams = best-of-5, everything else = best-of-3, unchanged. DEMOTED to tiebreak-level nudges only: recent form's coefficient halved (0.015 -> 0.0075, so a perfect recent record moves hold by 0.00375 — it cannot overrule the serve/surface signal), and the no-tape fallback estimate's rank swing compressed (0.82 - log10(rank)*0.09 -> 0.84 - log10(rank)*0.045, so rank 1-vs-100 moves the estimate ~5pp instead of ~10pp) with the season W-L band halved — rank and generic W-L still differentiate when no serve tape exists, but they can no longer manufacture a favorite out of a ranking. H2H edge (max +/-0.008) was already tiebreak-level and stays. The sim writeup now states the hierarchy: "Hold % is the engine, then surface, then the set format. Rank, recent form and season W/L are tiebreak-level nudges only." Covered by tests/tennis-importance-weights.test.js (25 checks against the real bundle functions). Untouched: 0.55 shrinkage, tennis 0.72 both-taped carve-out, down-only band calibration, 50/50 rule + caps, per-league bias (removed, stays removed), tennis weather wiring, tonight's basketball fixes, wiring-audit fixes, and NASCAR removal — all verified still present. Model 1.6.1.
NASCAR REMOVED FROM THE DAILY SLATE. The daily board-run loop (buildDailySlate, "run all") does not touch NASCAR — it is not in SLATE_LEAGUES, so racing/nascar-premier is never fetched for the daily slate. Reason: there is no way to do NASCAR effectively — no real statistical data source exists, so predictions were running on model-estimated 0-100 power ratings with no live per-driver stats. The single-prediction path is unchanged: an explicit NASCAR matchup request still verifies via get_game_schedule and estimates driver ratings as before. Covered by tests/nascar-slate-removal.test.js. Untouched: 0.55 shrinkage, down-only band calibration, 50/50 rule + caps, per-league bias (removed, stays removed), tonight's basketball fixes and wiring-audit fixes, model 1.6.1.
BASKETBALL FIXES. Five repairs across NBA/WNBA/CBB, all conservative and tested: (1) OUT STAR NO LONGER BOOSTS: fetchTeamRosterRaw never attached injury data, so isPlayerOut() was a silent no-op — an OUT leading scorer still boosted his team's lambda up to 1.25x while the injury haircut took at most ~6% (the 1.5.7 changelog claimed OUT suppression worked, but the branch never fired). Rosters now attach live injury statuses from ESPN's league injuries endpoint (same source getInjuryReport uses; a cached GET, one cheap call per league, fail-open to old behavior if down) matched by athlete id with display-name fallback — the existing suppression branch now actually fires, so an OUT star's factor reverts to neutral. This also repairs the same dead check for NFL/CFB/NHL key players. The general injury haircut is untouched. (2) NEUTRAL-SITE HOME COURT: a neutral venue previously fell through as "teamB is home" and handed +3/+4 to an arbitrary side — systematic in CBB conference tournaments and March Madness. New isNeutralVenueEvent() reads ESPN's neutralSite flag and homeAway "neutral" markers; neutral games now award 0 home-court points to either side and set both travel tiers to medium (neither side is "home"). Non-basketball leagues' behavior is unchanged. (3) CBB TRAVEL: basketball/mens-college-basketball added to SPORT_PATH_TO_TRAVEL_LEAGUE plus a new 72-program CBB arena-coordinate table (major conferences + Gonzaga/Memphis/SDSU/Dayton; newest arenas verified — Baylor's Foster Pavilion Waco 31.5567,-97.1216, Texas' Moody Center Austin). Same coarse tiers as every league (<200 mi short, >1500 mi long, else medium); unlisted programs fail open to the old medium default instead of the old flat 0.99x-for-everyone. (4) KEY-PLAYER SELECTION: basketball now picks the pool winner by max PPG, not max gamesPlayed — ESPN roster order is not depth-chart order, so a 12-PPG ironman was beating a 28-PPG star. NFL/CFB still sort by usage (passing attempts = the real starter, per the Dak Prescott verification) and NHL by games. (5) PER-LEAGUE SCORER SCALES: the single (18, 40) NBA-shaped scale compressed WNBA/CBB stars into a near-constant band — a 15-PPG CBB scorer even graded BELOW average (0.925). New keyPlayerScaleForLeague() mirrors the qbScaleForLeague precedent: NBA (18, 40) unchanged, WNBA (13, 29), CBB (11, 25), derived proportionally from each league's team-scoring environment (sim base rates ~115/83/72 PPG). Now an elite scorer in any league lands near the 1.25 clamp and an average starter sits at 1.00, exactly like the NBA. Covered by tests/basketball-fixes.test.js (51 checks against the real bundle functions). Deliberately NOT implemented (proposal-only): SOS builds, draw-sigma changes, 3PT-variance modeling, early-season-blend re-weighting, and the other 30 re-weight proposals from the loop audit. Untouched: 0.55 shrinkage, down-only band calibration, 50/50 rule + caps, per-league bias (removed, stays removed), model 1.6.1.
WIRING-AUDIT CLEAN TARGETS. Five small fixes, no re-weighting. (1) SOCCER STAR RESOLUTION: getSoccerStarFactors returned after the first competition where EITHER team had a scorer, so when the two teams played in different competitions the second team's star silently dropped (no adjustment at all). New pure resolveSoccerStarsByCompetition() resolves each team's star independently in ITS OWN competition, matching how the base-rate fetch already works — e.g. Arsenal's scorer from the Premier League and Real Madrid's from La Liga in the same pick. Early fetch exit preserved (stops once both teams resolve). (2) SOCCER VENUE-SPLIT TRACE: the writeup built the 50/50 TOTAL-vs-HOME/AWAY venue-split note with += then immediately overwrote it with = — math was fine, the prose lied. The note now actually prints. (3) NHL GOALIE SAMPLE GATE: the probable-goalie path fetched games played but never used it — a 2-game .960 call-up got 1-(0.960-0.905)/0.1 = max 0.8 goal suppression from 2 games of noise. New nhlGoalieEffSavePct(): under 3 games the SV% is NOT used (goalie stays identified, factor 1, no fallback to a different goalie's numbers), 3-9 games shrink toward the .905 league average by games/10, 10+ games full weight. Exact magnitudes: min games 3, full sample 10, league-avg SV% .905; existing 0.8-1.25 factor clamp untouched. (4) UNUSED FOOTBALL STATS WIRED: the ESPN football fetch already pulled yardsPerGame/passYpg/rushYpg/fourthDownPct but footballProfileToFactor never read them — data fetched and discarded. Now wired conservatively: yardsPerGame 0.0002 coef clamped 0.98-1.02 and ONLY when passYpa is missing (no double-count with per-attempt efficiency), passYpg 0.00015 clamped 0.99-1.01, rushYpg 0.00015 clamped 0.99-1.01, fourthDownPct 0.0006 clamped 0.99-1.01. League anchors: NFL 335/218/117/52, CFB 390/245/145/48 (total YPG / pass YPG / rush YPG / 4th-down %). Existing early-season shrink and final factor clamp unchanged. (5) DEAD-CODE CLEANUP: removed unused tennis fromTape, computeTennisPowerRating and its ratingA/ratingB adjustment chain, the unused forecast from getTennisMatchupPowerRatings, and the unconsumed rawA from runTennisMatchSimulation; parsed.ufcData marked as an explicit dead field (parse shape preserved). computeUFCPowerRatings was flagged as orphaned but is live through getUFCMatchupPowerRatings — verified, kept. Covered by tests/wiring-audit-3.test.js (28 checks, all green). Untouched: 0.55 shrinkage, tennis 0.72 both-taped carve-out, down-only band calibration, 50/50 rule + caps, model 1.6.1.
CONTINUED AUDIT: CLEAN DATA WIRINGS ONLY. Empirical graded-pick review (30 days through 2026-09-23): UFC weakest at 51.4% against 63.9% published confidence, NHL weak on a tiny sample, NFL slightly overconfident; WTA/tennis/WNBA/NCAAF/soccer all working — nothing about their wiring was touched. Six activations, all pure connections of data the worker already fetches, zero new model factors: (1) UFC fallback merge now keeps real SApM/strikeDefense/takedownDefense/finishMix/country from the ufcstats.com chain instead of dropping them and forcing the sim onto league-average constants (mergeUfcFallbackStats). (2) UFC unknown stance is now neutral (null) instead of defaulting to orthodox — missing data can no longer manufacture a phantom +/-0.025 stance edge vs a known southpaw. (3) NFL/CFB cold weather: Open-Meteo tempF already arrived but never reached the sim — now <=32F = 0.99x and <=20F = 0.985x on both offenses, symmetric, conservative, no heat drag for football. (4) Soccer venue splits: the football-data.org standings response already contains HOME/AWAY tables in the same payload (zero new calls) — each side's home/away-table scoring rates now blend 50/50 with its TOTAL season rate (TOTAL stays the anchor); neutral-venue tournaments (WC/EC) and venue samples under 3 games fail open to TOTAL, and fallback paths (BALLDONTLIE/ESPN) pass through untouched. (5) Golf now prefers the current-season scoring-average split over career/all-time when ESPN labels splits; unlabeled responses fall back to prior behavior. (6) Bowling blends all available PBA season rows 60/30/10 (most recent first) instead of reading only the first 20xx row — identical behavior with a single season. Also a trace honesty fix: darts no longer claims "current-season" averages since the parser can't verify season context. Deliberately NOT implemented (proposal-only, flagged for Drew): NHL PP/PK (the deep-dive's claim was half-wrong — the data exists on NHL's own API as powerPlayPct/penaltyKillPct but NOT in any already-fetched response, so it's a new fetch, not a free wiring; hockey literature also rates PP/PK weak vs 5v5 goal differential), UFC layoff/ring-rust (fight dates not proven present in the already-fetched page), NBA 4-in-5 density drag, widening the basketball key-player candidate pool, tennis thin-tape rank shrinkage, and golf/bowling/darts modeling tweaks. Covered by tests/wiring-audit-2.test.js (49 checks against the real bundle functions). Untouched: 0.55 shrinkage, down-only band calibration, 50/50 rule + caps, NHL run-all exclusion, per-league bias (removed, stays removed), model 1.6.1.
QB KEY-PLAYER FACTOR: PER-SOURCE SCALES. The deep-dive audit flagged that NFL/CFB both ran their QB factor through one scale (avg 90, divisor 150) — but ESPN's athlete-overview "QBRating" field is a DIFFERENT stat per league, verified live against ESPN's real API 2026-09-23. NFL = classic passer rating: Mahomes 103.8, Allen 125.8, Fields career 84.7, top-25 mean ~104.8 (0-158.3 scale, starter average ~90) — the old params were CORRECT here, kept as-is. CFB = NCAA passer efficiency: Sayin 182.4, leaders mean ~134.6 / median 132, FBS average ~135 (scale runs ~100-200+, uncapped) — the old params were WRONG: an average 135-efficiency college QB computed 1 + (135-90)/150 = 1.30 and maxed the 1.25 clamp, and nearly every qualifying college QB sat at or near the ceiling, so the factor couldn't differentiate. The NCAA.com "Pass Eff" fallback is the same NCAA-efficiency metric, so one CFB scale fixes both the ESPN path and the fallback. Fix: new qbScaleForLeague() gives each league its own (avg, divisor) — NFL (90, 150), CFB (135, 225) — and new qbFactorFromValue() applies the shared 0.8-1.25 clamp; the factor now always measures "how far above/below average for THIS stat". The CFB divisor mirrors the NFL curve in standard-deviation terms: elite (182) -> ~1.21, average (135) -> 1.00 neutral, bad (95) -> ~0.82. The trace label is honest too: CFB now reads "Pass Eff" instead of "QB Rating". Backup-QB substitution behavior untouched (OUT QBs still revert to neutral); shrinkage, bands, 50/50 rule, caps untouched. Covered by tests/qb-factor-scale.test.js (18 checks: elite/avg/bad/terrible per scale, monotonicity, no floor-drag for good QBs, no max-out for average QBs, plus a regression check showing the old params maxed a 135-efficiency QB at 1.25). Untouched: everything else, model 1.6.1.
TENNIS SURFACE IS A FIRST-CLASS INPUT. Drew played college tennis at Eastern Illinois, and Two things were wrong. (1) DETECTION: the surface came from fuzzy venue-city matching against TennisMyLife tournament names (e.g. ESPN's "All England Club" or "London" never matching "Wimbledon"), and the tournament name ESPN already returns was ignored entirely — so surface detection silently failed and everything downstream of it (surface hold tape, surface win%, modifiers) went neutral. Now the tournament name is checked first: the Slams have FIXED surfaces (Australian Open = Hard, Roland-Garros/French Open = Clay, Wimbledon = Grass, US Open = Hard), with venue fuzzy matching only as a last resort. ESPN's tennis scoreboard was checked live and exposes no surface field, so there is nothing to read from ESPN. Unknown = neutral, never a guessed surface. (2) WEIGHT: surface was a tiny nudge (clay -0.03 / grass +0.025, and only when there was no surface hold tape at all). Now surface sits alongside hold-serve % as the two dominant inputs. Exact magnitudes: when real surface-specific hold tape exists (12+ service games on that surface from TennisMyLife), the tape IS the surface effect — no modifier on top, no double-counting. Otherwise the surface moves hold probability directly: grass +0.04, clay -0.04, hard 0 (roughly the real ATP/WTA surface differential off the hard baseline, symmetric), PLUS the player's own surface W-L edge at 0.15 weight (needs 5+ surface matches from TennisMyLife, edge delta capped at +/-25pp — real observed data, never fabricated). Total surface contribution clamped to +/-0.06 so it modulates the pick instead of single-handedly deciding it; overall hold stays clamped to [0.50, 0.94]. The trace now names how the surface was detected (tournament name vs venue match). Covered by tests/tennis-surface.test.js (24 checks). Untouched: 0.55 shrinkage (0.72 when both players have hold tape), down-only band calibration, 50/50 rule + caps, Grand Slam best-of-5 logic, tennis weather wiring, model 1.6.1.
MLB WEATHER IN THE SIM. The Stadium Intelligence weather machinery was gated to football only — MLB got the display line but zero run-scoring effect (the audit flagged Wrigley wind as the example). Now MLB games feed real Open-Meteo temp/wind/precip into the simulation through a new baseballFactor lambda input: heat 90F+ = 1.02x expected runs both sides (thinner air, more carry), 80s = 1.01x, cold 40F or below = 0.99x, 25mph+ wind = 1.015x, 15mph+ = 1.005x, 70%+ rain = 0.99x — symmetric on both teams, clamped to 0.97-1.03 total so most games move 0-1%. Deliberately NOT the football wind math (a passing/kicking suppression would be wrong for baseball): windMph/precipPct are never set for MLB. No park-orientation data exists in the stadium table, so wind stays direction-neutral (temperature does the heavy lifting) rather than faked. Indoor venues and missing forecasts skip gracefully to neutral. Tennis already had weather wired at the hold-probability level (wind/heat/humidity move the hold number) — verified, left untouched. Untouched: 0.55 shrinkage, down-only band calibration, 50/50 rule + caps, model 1.6.1.
NHL CONFIRMED STARTING GOALIE. The goalie is the key player in hockey, but the sim was crediting the season usage-leader's save% — if the backup starts, the wrong goalie was modeled; if the usage-leader was out, neither was. ESPN publishes probableStartingGoalie (with a Confirmed status) on NHL scoreboard competitors — the same probables shape MLB uses for probable pitchers, verified live. New getNHLProbableGoalies() identifies the game matchup-first (same selection logic as the MLB pitcher fix: eventDate is only a tiebreaker) and the confirmed starter's real save% now drives the goalie factor ahead of the usage-leader. Same opponent-direction math and 0.8-1.25 clamp as before — only the selection got smarter, not bigger. When ESPN lists no probable goalie (or no A-vs-B game is found), the usage-leader path still runs but the trace says so explicitly instead of silently modeling the wrong goalie. Untouched: 0.55 shrinkage, down-only band calibration, 50/50 rule + caps, model 1.6.1.
TENNIS GRAND SLAM FORMAT FIX. Men's Slams were simulating best-of-3: the slamHint detector was built from the surface string ("Hard"/"Clay"/"Grass"), which can never match a Slam name, so the best-of-5 branch could never fire — systematically overstating underdogs at the four biggest events of the year. Slam detection now reads the REAL tournament name from ESPN's scoreboard (Australian Open, Roland-Garros/French Open, Wimbledon, US Open, variants handled): ATP at those events simulates best-of-5, ATP everywhere else best-of-3, WTA best-of-3 everywhere (women's Slams are best-of-3), ITF best-of-3. Unknown tournament defaults to best-of-3, never a guessed best-of-5. The trace now prints the format line with the tournament name. Untouched: 0.55 shrinkage, down-only band calibration, 50/50 rule + caps, tennis hold/weather wiring, model 1.6.1.
SAM 1.6.1. Model version bumped from SAM 1.6 to SAM 1.6.1: the label written to new Airtable pick logs (Model Version), the chat display label, and the test ping all read 1.6.1; historical changelog entries and graded rows keep their 1.6 labels. Rolls up tonight's data-integrity fixes: home-field advantage now enters MLB/NHL/soccer sims (+0.15 runs / +0.25 goals / +0.3 goals in expected-goals math), MLB probable-pitcher lookup identifies the game by matchup instead of trusting the model's date (wrong dates no longer pull the wrong starter; explicit uncertain result with 0% starter weight and a model-facing WARNING when no A-vs-B game is found), real great-circle travel distance tiers (159-venue table; long 0.97x / short 1.0x branches now fire), soccer star-power factor from football-data.org top scorers, and removal of the dead injurySeverity parameter (severity was already encoded per-player in the injury haircut — wiring it would have double-counted). Known exception: strength of schedule (sosDiff) is still neutral everywhere — honest SOS needs full-season schedules plus every opponent's win%, new fetching plumbing that is flagged for a follow-up build, not hacked in. Untouched: 0.55 shrinkage, down-only band calibration, 50/50 rule + caps, NHL run-all exclusion, grading, branding.
SOCCER STAR-POWER FACTOR. Soccer was the only team league with no key-player input at all. New getSoccerStarFactors() pulls each competition's top scorers from football-data.org (same source already used for soccer scoring baselines, limit 50 so every club's top scorer is covered), takes each side's top scorer goals/game (minimum 5 appearances), and converts it to a "self" lambda multiplier analogous to the NBA leading-scorer PPG factor: league-average 0.35 goals/game is neutral, an elite ~0.75 goals/game striker lands near +20%, clamped to 0.8-1.25 like every other key-player factor. Applied per side through the same pitcherFactor slot the other leagues' key-player factors use. The model-facing trace names the scorer with goals and appearances. Fail-open: no scorer data means no adjustment, never a penalty. Also removed the dead if (key === "SOCCER") line inside getKeyPlayerFactor (unreachable since soccer never entered that function). Untouched: 0.55 shrinkage, down-only band calibration, 50/50 rule + caps, model 1.6.
REAL TRAVEL DISTANCE TIERS. The long (0.97x) and short (1.0x) travel branches in calculateLambdaMultiplier could never fire: getHomeAwayContext only ever returned "home" or "medium", so every road team got the same 0.99x whether it bussed across town or flew coast-to-coast. Now the away side's tier comes from the real great-circle distance between the two teams' home stadiums: under 200 miles = short (1.0x, e.g. Yankees at Mets, Jets at Giants sharing a stadium = 0 miles), over 1500 miles = long (0.97x, true coast-to-coast), everything else stays medium (0.99x). Built a verified 159-venue coordinate table covering NFL (32), NBA (30), WNBA (15 incl. 2026 expansion Portland Fire + Toronto Tempo), MLB (30, Athletics at Sutter Health Park Sacramento), NHL (32, Utah Mammoth), and EPL (20, current 2026-27 lineup) — 18 venues spot-checked against a second geocode source, all matched. Lookup reads the ESPN event's own team abbreviations with alias + tolerant matching (GSW/GS, CONN/CON, ATH/SAC) so variant spellings still resolve; anything unresolved falls back to the old home/medium behavior, never an error. CFB/CBB have no table yet (hundreds of venues — needs a bigger verified build or runtime ESPN team-venue lookups) and keep the previous behavior; flagged as a follow-up. Untouched: 0.55 shrinkage, down-only band calibration, 50/50 rule + caps, model 1.6.
DEAD MULTIPLIER INPUTS RESOLVED. (1) injurySeverity REMOVED, not wired. calculateLambdaMultiplier accepted an injurySeverity input (OUT = 0.94x, Questionable = 0.97x) that no caller ever set. Checked for double-counting before wiring it: scoreInjuryReport already encodes severity per player (OUT = 1.0x, Doubtful = 0.7x, Questionable/day-to-day = 0.4x of position weight, capped at 0.82 floor). A flat most-severe-status multiplier on top would have counted the same signal twice, so the dead parameter and its branches were deleted instead of wired — the richer per-player encoding is the severity mechanism, now documented in a code comment. No live pick changes: the parameter was never populated. (2) sosDiff LEFT UNWIRED, flagged for follow-up. Strength of schedule is still 0 (neutral) everywhere: honest SOS needs full-season schedules plus every opponent's win%, which is new fetching plumbing the pipeline doesn't have (the 17-day rest/form window only covers recent games — a "recent SOS" from two weeks of data would be fake precision, so it was not hacked in). The sosDiff input stays in calculateLambdaMultiplier ready for that build. Untouched: 0.55 shrinkage, down-only band calibration, 50/50 rule + caps, model 1.6.
MLB PROBABLE-PITCHER MATCHUP-FIRST LOOKUP. getMLBProbablePitchers picked its game by matching the model-passed eventDate against an 11-day scoreboard window — a wrong or missing date silently pulled the wrong game's starter. Now the game is identified by the MATCHUP: candidates are already filtered to games where the two teams play each other, and the pick is the date-matched game when eventDate hits (doubleheader-safe: soonest upcoming game that day), otherwise the nearest upcoming A-vs-B game in the window. A wrong model date no longer changes the answer. The uncertain fallback now fires only when no game between the two teams exists in the window at all — in that case the function returns an explicit uncertain result (instead of null) so the model gets a WARNING line and starter weight goes to 0 rather than silently running pitcher-neutral. Untouched: 0.55 shrinkage, down-only band calibration, 50/50 rule + caps, model 1.6.
NOTIFICATION TAP DOUBLE-SEND FIXED (CROSS-TAB). Tapping "prediction ready" could render the reply twice: the original chat tab was still polling and the notification opened a second view, and the render guards were per-tab memory only. Seen/started flags are now also in localStorage so the second tab skips a job the first tab already rendered.
DUPE CHECK NOW KV-ONLY. The "is this game already logged" lookup was falling through to a slow Airtable call. Now it checks the KV replica only — instant. The replica is kept in sync on every log and grade.
PREDICTION WRITEUPS UPGRADED TO PROFESSIONAL ANALYST GRADE. Every prediction now reads like a professional analyst of that specific sport broke the game down: matchup setup, key numbers with named players, a substantive professional breakdown walking through the real edges (matchups, injuries, rest, form, weather, pitching, home/away), then the win-percentage split and definitive winner. No more short blurbs that hide the analysis.
HOME-FIELD + PITCHER DATA FIXES. (1) HOME-FIELD NOW ENTERS MLB/NHL/SOCCER SIMS. The sim told the model home advantage was "factored directly into this simulation" for these leagues, but no home-field number ever reached the trials — only NFL/CFB/NBA/WNBA/CBB got one. Now the home team gets a real edge in the expected-goals/runs math: MLB +0.15 runs, NHL +0.25 goals, soccer +0.3 goals (conservative, grounded in measured per-game home scoring edges). (2) MLB PROBABLE-PITCHER WRONG-GAME FALLBACK. When ESPN listed several games in a series and none matched the requested date, the sim silently used the first game's starter and the model got no warning. Now the pitcher data is flagged uncertain: a WARNING line goes to the model, and starter lock weight is cut in half so a possibly-wrong pitcher can't drive the pick — the blend shifts toward bullpen/staff instead. Untouched: 0.55 shrinkage, down-only band calibration, 50/50 rule + caps, NHL run-all exclusion, grading, branding, model 1.6.
LEAGUE BIAS REMOVED. The per-league calibration bias is gone entirely: no more Supabase bias fetch/adjust on the sim text the model sees (runMatchupSimulation), the grader no longer trains it (updateCalibrationBias call removed), and the /admin/backfill-calibration endpoint is removed. Deleted dead code: applyCalibrationToResultText, getCalibrationBias, applyCalibrationBias, updateCalibrationBias, backfillCalibrationFromAirtable, CALIBRATION_LEARNING_RATE/MAX_BIAS, and tests/calibration-picked-side.test.js. The sim's numbers now reach the model unadjusted by league history. Untouched: 0.55 shrinkage, down-only band calibration, 50/50 rule + caps, NHL run-all exclusion, grading, branding, model 1.6.
CONSERVATISM SHRINKAGE SET TO 0.55 (0.50 → 0.55). The sim's raw edges were being flattened into coin flips: post-shrinkage sims average ~63% picked-side with only 10% landing 48-52, but published confidence lands at median 49% with 53% inside 48-52 — the band calibration fits on its own deflated outputs and compounds the pull-down. Easing shrinkage keeps more of the sim's lean (a raw 70% edge reports ~61% instead of 60%, a raw 80% edge ~67% instead of 65%) while all deflationary mechanisms stay in place: down-only band calibration, 50/50 rule + caps, NHL run-all exclusion, grading, branding, model 1.6.
DATA PIPELINE GAPS CLOSED. Deep audit found the model was missing data it should have had: (1) Rest days + recent form now factor into MLB, NHL, and Soccer sims — previously only NFL/NBA/WNBA/CFB/CBB got them, leaving NHL back-to-backs and MLB hot/cold streaks unadjusted. (2) CBB injury weights fixed — was silently using NFL position weights for college basketball; now uses proper basketball weights. (3) MLS/EPL/Premier League queries now normalize to SOCCER instead of silently degrading with NFL injury weights and no rest/form data.
SLATE TIMEOUT RAISED TO 5 MINUTES. The Run Today slate runner was aborting prediction requests at 120 seconds, causing "signal timed out" failures on slow games. Bumped to 300 seconds so slow signals can finish instead of failing.
NHL BACK IN RUN ALL. NHL is restored to the daily slate leagues after the 2026-09-20 removal. Single NHL predictions were never affected.
MARKET COMPARISON RESTORED, STAKING STAYS OUT. The model-vs-market commentary is back on every prediction: "SAM sees +X% edge over the market" / "SAM is X% below the market implied probability" / "SAM and the market are roughly aligned" now render on the Market Line block for new predictions and logged-pick displays. The edge-gate verdicts (BET/HOLD/PASS) and all stake sizing stay deleted — this is market-agreement numbers only, no imaginary-money recommendations.
DEFLATIONARY CALIBRATION RESTORED. The 2026-09-21 calibration redesign is reverted: (1) fitBandCalibration buckets on the LOGGED band-calibrated "Confidence %" again instead of the raw sim picked-side % — bands learn on their own past outputs; (2) the symmetric +/-10pt clamp is replaced by the down-only clamp (published confidence may only shrink toward the band's realized hit rate, never inflate); (3) the online league-bias learner trains on the logged confidence again instead of the raw sim split. Buckets are back to 50-59/60-69/70-79/80-89/90%+. Untouched: 0.5 shrinkage, 50/50 rule + caps, NHL run-all exclusion, the picked-side bias fix, grading, branding, model 1.6.
REFUSED PREDICTIONS NO LONGER EAT A FREE PREDICTION. A live test showed a guest asking for a matchup that isn't on the board ("Tigers vs Mariners" — the model correctly refused instead of guessing) still burned one of the 3 daily free predictions: the counter dropped from 3 to 2 with no pick delivered. The guest cookie was being incremented on every /api/chat reply, even when no pick came back. Now the incremented cookie is only set — and the remaining count only drops — when a real pick is produced. Logged-in users were already protected (their deduction only writes when a pick exists); guests are now covered too. Same fix applied to the tennis coin-flip decline path, which had the identical leak, and /api/chat now fails closed if the billing backend key ever goes missing instead of serving unmetered predictions.
MODEL SWITCHER REMOVED ENTIRELY. The "ParlayKings 1.6" tag by the chat input is gone completely — no button, no fold-out, no label. One model only (SAM 1.6 on Muse Spark 1.3), so there was nothing to switch or explain. The picker markup, all its styles, and its toggle script are deleted; the chat input now sits alone. Nothing else touched.
MODEL TAG FOLDS OUT AGAIN. Tapping the "ParlayKings 1.6" tag by the chat input now opens the info card again, showing the model name and its explanation line ("Sports Analysis Machine (SAM) 1.6 original model still most accurate."). Still one model only — SAM 1.6 on Muse Spark 1.3, nothing to switch between; the card is informational. Tapping anywhere else closes it.
MODEL LABEL NO LONGER LOOKS LIKE A BUTTON. The chat input's "ParlayKings 1.6" tag was still styled like the old dropdown (border, pill shape) so it looked tappable but did nothing when tapped. It's now plain text. Single-model setup is untouched: SAM 1.6 on Muse Spark 1.3, no picker, nothing to choose.
GAME-CARD FOLLOW-UPS REMOVED. The post-simulation gameCard feature from earlier tonight is fully removed: the five card functions, the /api/chat card wiring (request parsing, cached-response property, prompt grounding, main response property), the background-job card property, and all client-side card state (hold/send/receive/reset). Chat behavior is back to exactly what it was before the feature existed. Nothing else touched.
EDGE GATE REMOVED FROM OUTPUT. The edge-gate verdict block is gone from every prediction reply: no more BET/HOLD/PASS-by-edge lines, no edge-points numbers, no model "fair price", no stake sizing, and none of the "SAM sees +X% edge / X% below the market / roughly aligned" commentary. The Market Line block stays as plain market data (book, both sides' prices, implied %). The 50/50 note is unchanged. The gate's code (edgeGate/edgeGateVerdict, Kelly sizing, odds converters, thresholds) had zero remaining consumers after the rendering was cut, so it was deleted outright; the postmortem faded_market tag still works because it computes from the logged confidence vs the market-line implied % numerically. Regression test replaced: tests/edge-gate-removed.test.js (was edge-gate-demotion.test.js). Untouched: prediction flow, grading, calibration, 50/50 rule, 0.5 shrinkage, NHL run-all exclusion, alerts, recap, branding, model 1.6, /api/speak.
GAME-PRO FOLLOW-UP WORDING FIX. The follow-up grounding rule told the model to say the canned line "my sim notes don't cover that" when a question goes beyond the game card. Drew flagged it as robotic — replaced with an instruction to say so in its own words (e.g. that wasn't part of this simulation). No behavior change: still answers from the card, still never invents sim details. Nothing else touched.
POST-SIM "GAME PRO" FOLLOW-UPS (Drew's request). After a prediction, /api/chat now returns a compact gameCard built from the real sim outputs (matchup, league, both teams' sim percentages, picked winner, calibrated confidence, the sim's key-factors line, and the Market Line block — under 1.5KB). The client holds it per-conversation and sends it back with the next message; a new chat or opened thread starts clean so cards never leak across conversations. If the next message is a follow-up about that game, the card is injected as structured grounding with an anti-freelancing rule: answer FROM the card, say so in their own words when the card doesn't cover the question, never invent sim details — and if the message names a different matchup, run a fresh simulation for it. Follow-up vs new-game detection is deterministic: a message naming teams that match neither card team drops the card; everything else keeps it. Client-supplied cards are strictly sanitized (allowlisted fields, length caps, newlines collapsed) before they can reach the prompt. Background-job predictions also arm the card. Additive only: no card present means byte-identical behavior; prediction, grading, calibration, 50/50 rule, edge gate, shrinkage, staking, NHL exclusion, and all prompt rules untouched. Regression tests: tests/game-card.test.js (37/37).
BET VERDICTS DEMOTED TO HOLD-ONLY. The 10,000-run Monte Carlo staking backtest found the edge gate inverted: picks it classified BET (edge of 3+ pts) won only 39.1%, and half-Kelly sizing on them ruined the bankroll in 99% of simulated worlds — the gate bets exactly where SAM is overconfident vs the market. Per Drew's "only if it wins" rule no staking parameter was changed; instead the published verdict no longer prints a BET signal or any stake sizing. A BET-classified pick now reads as HOLD with the edge info preserved ("+X pts vs market … no stake recommended while the gate is under review"). Internal gate classification is kept for the eventual rework; HOLD/PASS verdicts unchanged. Regression test added (tests/edge-gate-demotion.test.js). Untouched: 0.5 shrinkage, 50/50 rule, NHL run-all exclusion, staking math itself, branding, model 1.6.
CALIBRATION REDESIGN: BANDS NOW FIT ON RAW SIM VALUES + SYMMETRIC CLAMP (Drew: "do whatever you think best"). The 2026-09-21 audit found the band calibration deflationary: 62% of picks logged below 50% confidence while that group hit 55%+, Brier 0.2434 vs 0.2444 for a constant base-rate predictor. Root causes fixed: (1) fitBandCalibration now buckets on the RAW band-stage input — the bias-adjusted sim picked-side % (max of logged "Team A %" / "Team B %") — instead of the already band-calibrated logged "Confidence %", which had created a feedback loop (bands fit on deflated values, then pulled new picks further down). New bands match the sim's honest structure: 50-60 / 60-70 / 70-80 / 80%+. (2) The down-only clamp is replaced by a symmetric +/-10pt clamp (BAND_CALIBRATION_MAX_SHIFT_PTS) inside applyBandCalibration, so legitimate pull-ups are possible again; exact-50 is never touched (the 50/50 rule owns it). (3) The online league-bias learner (updateCalibrationBias) now trains on the same raw picked-side % via new biasLearnerPredictedPct(), instead of the deflated logged confidence — it was learning on one scale and correcting another. Backtest (386 decided SAM 1.6 picks, chronological walk-forward, 21 re-graded excluded): Brier 0.2434 -> 0.2393, share logged below 50%: 62% -> 8%, new logged calibration is monotonic (50-55%: 52% hit, 70-75%: 79% hit), and the 70-80% raw sim segment is preserved (85 picks, mean logged 74.6%, 75.3% hit). Supabase sam_calibration values left to re-learn (not reset): the 21 re-graded picks' total contamination was only -10.5pts spread across 6 leagues (worst -4.5 NFL/WTA), the historical values are unreachable from here (worker-side secret), and the online updater self-corrects within a few graded picks per league now that inputs are honest. STAKING BACKTEST (Drew: "if it wins more bets... i am happy" — data-driven only): the edge-gate (BET at SAM-minus-market edge >= 3) is INVERTED on 243 decided picks with DraftKings lines — BET picks went 39.1% (69 BETs at 36.2% on old conf), and 10,000 bootstrap Monte Carlo runs show half-Kelly on gate BETs ruins the bankroll in 99.2% of worlds (median terminal 0.6u from 100u); quarter-Kelly, flat stakes, and a tighter +5 gate all lose too; even the diagnostic inverted gate (bet edge<0) only reaches median 97.9u. NO staking change shipped — nothing wins robustly, so per Drew's condition the staking layer is untouched. FLAG FOR DREW: the published "BET ... half-Kelly" verdicts are -EV as designed; recommend deciding whether to keep, demote, or rework the edge gate. Untouched by design: 0.5 shrinkage, 50/50 rule + caps, NHL run-all exclusion, staking params, branding, model number.
NHL MADE A FIRST-CLASS LEAGUE IN THE DATA PIPELINE. Airtable already carried "NHL" as a League select option and the worker was already wired for it (ESPN paths, sim distributions, grading margins, alert prefs) — the only gap was the system prompt's [PICK] rule-6 league list, which never told the model NHL was a valid League value, so NHL predictions would have been logged as "Other" and fragmented the history. Added NHL to that list. The Bets-replica chain (Airtable → KV → bets-replica.json → Google Sheet "Bets" tab) is league-agnostic with no per-league tabs or validation, so NHL rows flow cleanly; verified 19 league values already coexist there. The 2026-09-20 run-all exclusion is untouched: the slate loop still skips NHL unless Drew says otherwise.
CALIBRATION BIAS NOW APPLIES TO THE PICKED SIDE + 21 HISTORICAL 50/50 PICKS RE-GRADED TO PUSH. (1) Logic-bug fix in applyCalibrationToResultText: the Supabase league bias is computed from picked-side outcomes (stated confidence of the pick vs. actual), but it was being added to teamA's percentage unconditionally — teamA/teamB are just tool-arg order. When the model picked teamB in an overconfident league (NFL bias ≈ −12.6), e.g. teamA 35%/teamB 65% printed as teamA 22%/teamB 78% — the printed confidence got inflated instead of corrected, and feedback from inflated picks drove the bias further negative. The bias is now applied to the presumptive pick's percentage (higher sim side; teamA on an exact tie, matching the [PICK] rule) with the other side mirrored as 100−adjusted. ±15 clamp, 1–99 range, and the calibration note are unchanged. No Supabase calibration values were modified — this is a presentation/prompt-text correction, not a calibration value change. (2) Drew approved re-grading the 21 historical SAM 1.6 picks logged at exactly 50% confidence that the auto-grader had scored W/L (7–14) instead of Push, violating the standing 50/50 rule — each was verified live in Airtable before its Result was set to Push; nothing else on those rows was touched. Flag: those 21 picks previously fed the calibration updater as W/L, so their historical influence on the running Supabase bias values can't be cleanly unwound — the bias values keep that history by design. Untouched by design: 0.5 shrinkage, calibration math/values, confidence/staking/thresholds, risk settings, the SAM XAI 2.1 history classifiers, and /api/speak.
GRADER 50/50-RULE ENFORCEMENT + TENNIS/WTA LABEL STANDARDIZATION. (1) The auto-grader now enforces Drew's 50/50 rule on the confidence number itself: any pick logged at exactly 50% confidence grades as Push, even when the model named a side (previously only "50/50"/"no pick" winner labels were caught, so 21 historical conf-50 picks were auto-graded W/L). New isConfidenceFiftyFifty() helper; the graded Push is also excluded from calibration updates and tagged in Notes. Going-forward only — the 21 historical rows were left untouched per Drew. (2) Women's tennis is now a single label everywhere going forward: the [PICK] tag format and both tool league enums now state Tennis = men's ATP only and WTA = ALL women's matches (the model knew — e.g. calling Potapova/Andreeva "a WTA tour-level matchup" — but logged League=Tennis, fragmenting per-league calibration). Historical Airtable rows were not rewritten. Untouched by design: 0.5 shrinkage, all calibration math/values, confidence/staking/thresholds, risk settings, the SAM XAI 2.1 history classifiers, and /api/speak.
SPARK 1.3 SINGLE-MODEL CLEANUP. Drew's call after the Grok removal: with SAM 1.6 (Muse Spark 1.3) as the only model, strip everything that only existed for the old multi-model days. (1) Onboarding copy: the Workflow explainer still said “ParlayKings 1.6 or ParlayKings XAI 2.1, whichever you picked” — now just ParlayKings 1.6. (2) Deleted the entire dead Gemini LLM stack: callGemini/requestGemini, GEMINI_MODEL, TOOLS_GEMINI, geminiApiKeys (zero callers each, verified) plus dead provider===“gemini” branches in extractReplyText, fetchCacheCandidateRecords, fetchPendingSlateRecords, and fetchPredictionHistory; runMatchupSimulation's default provider is now “meta”. (3) buildDailySlate no longer fetches today's logged records twice — one KV-replica read; the alreadyLoggedGemini admin stat now aliases the meta count. (4) isPublicReplicaProvider() is now a constant true; the chat page's model-picker dropdown (one option) became a static “ParlayKings 1.6” label and the picker JS was removed. (5) fetchPredictionHistory's up-to-10-page Airtable scan is now cached 10 minutes in the SAM_BETS KV namespace (provider-scoped key); all downstream computation and postmortem side effects still run per message. (6) New single-model test asserts one provider target, no picker options, and no dead model stacks. Untouched by design: 0.5 shrinkage, calibration math, confidence/staking, the 50/50 rule, the NHL slate exclusion, the SAM XAI 2.1 history classifiers, and /api/speak (xAI orion voice).
SAM XAI 2.1 (GROK) REMOVED — SAM 1.6 IS NOW THE ONLY MODEL. Drew's call: the Grok picker option was barely used (one logged pick in the week before removal) and confused people more than it helped. Removed end to end: the "ParlayKings XAI 2.1" model-picker option, Grok from the auto-grade loop (gradeAllModelBets now only grades SAM 1.6 rows), the Meta/Grok cross-provider failover in /api/chat, the "get a second opinion" button (it only existed to ask the other model), and all xAI chat-model code (callGrok/requestGrok, GROK_MODEL, GROK_URL, the Grok key rotation + quota-skip machinery, allGrokKeysDown). The xAI text-to-speech endpoint (/api/speak, voice "orion") is untouched — it still reads GROK_API_KEY for voice only. Historical Airtable rows labeled "SAM XAI 2.1" are left alone; the Meta label matcher still excludes them so they never leak into SAM 1.6 track records or calibration.
NHL OUT OF RUN ALL. The Run All board run no longer fetches or runs NHL games — hockey is off the slate until Drew says otherwise. Single predictions can still pick NHL explicitly; this only stops the automatic whole-board run from touching it.
BOARD-RUN ALERTS GO TO EVERYONE, SCATTERED THROUGH THE DAY. When the admin runs the whole board, high-edge pick alerts now fan out to every push-subscribed user — each person only gets the picks matching their own leagues and confidence floor, with per-user dedupe so nobody's delivery blocks anyone else's. Previously the alert only went to whoever ran the prediction. To avoid one big burst, picks for games happening today go out immediately while future-game picks drip out at most one every 30 minutes per user (quiet hours still respected, same 12-a-day cap). Covered by tests/he-fanout.test.js.
GAME-DAY TEAM ALERTS. ParlayKings now pings you for your teams: Michigan football, Michigan State football, the Lions, and the Tigers — game start, halftime (football), and final score, straight from ESPN's scoreboard. Pick your teams in Profile under "My teams" (all off by default); team alerts go out live even during quiet hours, with 24-hour dedupe and the same 12-a-day cap as everything else. The worker checks on its every-minute cron but only calls ESPN when somebody actually has a team on, and only on days that team plays. Covered by tests/game-alerts.test.js.
RECAP STOPS COUNTING EVERYTHING TWICE. Both provider targets resolve to the same Airtable base, so the nightly recap fetched the same table twice and doubled every W-L line (MLB showed 14-4 for a real 7-2). fetchGradedRecords now dedupes by Airtable record ID — counts are identical whether one target or many resolve to the base. No other alert logic, thresholds, or timing touched. Covered by tests/alerts.test.js (59/59).
NAV DRAWER TRIMMED, GRIMALDI.TV UNDER TRIVIA. The hamburger menu carried a full podcast menu (live broadcast, photo gallery, blog, merch, clip vault) on top of every sports link. The podcast stack is back to HOME + EPISODES only, in both drawer copies (shared builder and main chat page), and the GRIMALDI.TV link sits under TRIVIA on every host instead of only parlaykings.net. Covered by tests/nav-links.test.js (29/29).
NOTIFICATIONS SAY WHAT THEY ARE. Every push ticket woke the phone with a fresh look at the latest finished job, so an alert ticket arriving later re-announced "Your prediction is ready" for a job long done — burying the real alert banners. The service worker now only announces jobs finished within the last 15 minutes, and when there is no fresh job and no pending alerts it shows nothing instead of a stale generic. Covered by tests/alerts.test.js wiring checks (53/53).
TAP A NOTIFICATION ONCE, SEE THE ANSWER ONCE. Tapping "prediction ready" could type the reply twice a few seconds later: the startup catch-up and the tap path raced each other (and a tap can arrive twice — service-worker message plus ?job= URL), and nothing marked a render as started until it finished, so an in-flight render never blocked a second one. Every render path is now single-flight (startedBgJobs is marked synchronously the moment a render begins; the startup sweep skips the tapped job; repeat taps are no-ops while failed taps stay retryable). Tapping while the run is still going also adopts the same 10-minute watch window as a fresh background send instead of risking eternal dots. Covered by tests/bg-resume.test.js (18/18) plus a bg-poll.test.js harness sync.
PICK ALERTS, NIGHTLY RECAPS, LINE-MOVE INTEL — ONE DISPATCHER. Every alert now flows through a single choke point (dispatchAlert): per-user prefs, quiet hours, and rate caps are enforced before any push goes out. Three alert types: high-edge pick alerts fire the moment SAM logs a pick above your minimum confidence; a grading recap lands nightly at 9 PM ET; line-movement intel fires when a logged pick's line moves 25+ cents (or 15+ cents within an hour), with a 6-hour cooldown per pick so it can't spam. Alerts are off until you enable them in settings — per-league toggles, 50-100 confidence floor, and quiet hours included. Capped at 12 alerts a day with 24-hour dedupe. Covered by tests/alerts.test.js (51/51).
PROFILE PHOTO SAVES WITH SAVE. Picking a photo used to strand you on a second Upload button nobody could find: preview, then Save does everything — photo goes up first, settings after, footer avatar and profile picture refresh together. Uploads also survive the storage lag that used to say "saved" while showing nothing (background retries plus self-reloading images), and iPhone HEIC shots get a plain-English nudge instead of a dead error. Covered by tests/profile.test.js (81/81).
SPORTSBOOK TAB TAPS FIRST TIME, EVERY TIME. The edge tab slid sideways on hover, so on touch screens the button moved out from under your finger and the tap landed behind it — flash, then nothing. The tab never moves now; it also carries an invisible wider tap halo and sits above every page layer. A well-meant 350ms close-blocker that ate quick follow-up taps is gone too. Covered by tests/sportsbooks.test.js (42/42).
NOTIFICATION TAPS LAND ON THE ANSWER, NOT THE DOTS. Tapping "prediction ready" could just foreground the frozen page you left — still showing the old thinking dots, with nothing re-running — until you clicked out and back in. The tap now wakes the open page and loads that prediction inline (adopts its chat, types the reply, clears the dots); only browsers that can't steer the open window fall back to a fresh load. Covered by tests/bg-resume.test.js (14/14).
USER PROFILES: NOTIFY PREFS, SETTINGS, AVATARS. The KV profile record (built for pins) is now a full user profile: notification prefs (master toggle, per-league toggles, 50-100 minimum confidence, quiet hours) plus settings with display name and avatar. New profile/settings UI in the chat drawer: prefs form plus picture upload with client-side 512px downscale. Avatars target R2 (bucket env.AVATARS) with MIME magic-byte check, 2 MB cap, and server-side dimension parsing for JPEG/PNG/GIF/WebP - and the whole flow degrades gracefully until R2 is enabled (upload control hidden, everything else works). To enable: create bucket parlaykings-avatars and add [[r2_buckets]] binding="AVATARS" to wrangler.jsonc. Covered by tests/profile.test.js (55/55). No notification sending logic - prefs only.
PIN CHATS AND PICKS. Chat history now supports pinning entire chats and individual games/picks, freely combined. Pins are per-user and persistent in a KV profile record (profile:
TENNIS COIN FLIPS STAND, CAPPED AT 3 A DAY. Genuinely even tennis (model calls 50/50, or sim split within 2 points) now stands as an honest 50/50 — logged, grades Push — instead of being forced onto a side. Daily KV counter (ET day, 48h TTL) caps stands at TENNIS_FIFTYFIFTY_DAILY_CAP; past the cap the request is declined as too-close-to-call with nothing logged and no credit deducted. Other leagues untouched; counter failures fail open to standing. Covered by tests/tennis-fiftyfifty-cap.test.js.
BAYESIAN BAND CALIBRATION BACK ON. The per-band fit/apply helpers stayed in the bundle computing params, but the call-site block had been replaced with void finalCalibrationParams, so the correction was computed and thrown away. Restored after sim lock in the /api/chat handler (primary and failover paths share it): fitted bands may pull the sim-locked number DOWN toward the band's historical hit rate, never above the sim; the visible reply text is rewritten to match the logged value. Thin bands stay near raw via shrinkage (prior weight 10); the fit needs 20+ graded picks or it no-ops. Covered by tests/band-calibration.test.js (12/12: overconfident 60-69 band pulls 64 to 53, thin 80-89 band stays near raw); the sim-lock structural guard now allows the sanctioned down-only block.
SAVED CHATS KEEP THEIR BUTTONS. Opening a conversation from history rendered plain bubbles: no thumbs, no share, no second opinion. The restore path now runs every revived reply through the same action row as live messages (copy/up/down/share/speak/second-opinion; provider falls back when the turn predates it). One-line fix in openConvo. Covered by tests/saved-chat-actions.test.js (9/9).
SELF-SERVE PASSWORD RESET. Forgot-password link on the login page; emailed single-use reset links (1-hour expiry) via Email Sending from noreply@parlaykings.net; new-password page with min-8 + confirm. Same message whether or not the account exists (no enumeration), 3-per-hour throttle per identifier, tokens burned on use. Reuses PBKDF2 hashing + Supabase PATCH path. Covered by tests/password-reset.test.js (9/9).
POSTMORTEMS SAY WHY, NOT JUST "OOPS". The grader now reads the final score: league-aware margin bands (NFL/CFB 8/17, MLB 2/6, NBA/CBB 6/15, NHL 1, soccer 1/3) tag blowout losses rout (sim misevaluated strength) and within-one-score losses one_score (the short end of variance, process intact). Fading the market 5%+ and losing now also tags market_called_it (the market was right to be colder). coinflip_miss/favorite_miss demoted to fallback-only when no causal tag fires; high_confidence_miss untouched so the conservatism loop keeps its input. Autopsy lines state final margin plus a plain verdict. Covered by tests/postmortem-upgrade.test.js (54/54).
POSTMORTEMS CHECK THEIR OWN HOMEWORK. Two upgrades. (1) CALIBRATION OVERLAP GUARD: the sim-side conservatism pull (0.9 toward 50 after repeated high-confidence misses) and the per-band calibration correction could both fire on the same phenomenon and double-count. fetchPredictionHistory now measures both side by side: when global 70%+ bands are already correcting down 5+ points, the sim pull softens to 0.95 and says so in its note instead of stacking the full pull on top. (2) EFFICACY TRACKING WITH AUTO-RETIRE: every adjustment activation is recorded per league (KV-backed) with its tag baseline; after 8+ further losses the tag rate is re-measured — halved means working (baseline refreshed), flat means retired as ineffective, and high bands over-hitting while the pull is on means retired as overcorrecting. Retirements hold 30 days with the reason surfaced in the sim note, then the pattern can re-arm. Pattern faded means the record drops. Covered by tests/postmortem-upgrade.test.js (29/29).
TAPPING A NOTIFICATION NO LONGER LANDS ON ETERNAL DOTS. Opening ?job= on a fresh page meant a new chat ID, so the finished job never matched, and a stale storage read showing "pending" meant the page never started watching. The page now adopts the tapped job's chat, renders a finished job straight into it, and keeps polling when the first read is stale; every completed background render also clears the thinking dots. Covered by tests/bg-resume.test.js (11/11).
NOTIFY JOBS START ON SEND. Armed-bell sends used to sit pending until the minute cron noticed them (up to 60 seconds of dead wait): the enqueue handler swept the whole queue instead of running the job in hand, and storage-list lag could hide the fresh job until the next tick. The handler now runs that job object immediately via ctx.waitUntil(runBgJob(...)) and the push still fires the second it finishes. The minute cron keeps grading plus the replica dump and retries only stuck runs (a "running" job older than 6 minutes, down from 15). Bell, push, SW, jobs routes, and client polling untouched.
REASONING BUBBLE NO LONGER VANISHES FROM SAVED CHATS. Opening a chat from History showed only the user/answer bubbles -- the collapsible "ParlayKings Reasoning" bubble was gone every time. Root cause, three links: (1) the client saveConvoTurn posted only {convoId, message, reply}, never the thinking text; (2) /api/convos/save stored only {user, reply, at} per turn; (3) openConvo mapped thinking to '' and rendered only user/ai bubbles. Fix: saveConvoTurn now sends thinking (capped 8KB like replies), the server stores it on the turn, the background-job internal save threads job.result.thinking through too, and openConvo restores the bubble collapsed-by-default in its original spot plus keeps thinking in conversationHistory for follow-ups. Old saved chats without thinking render exactly as before. The bubble itself is unchanged behaviorally (collapsed by default, click the label to expand, chevron toggles) but renderReasoning is now built on DOM APIs instead of an innerHTML template, and whitespace tidying moved into a tiny tidyReasoningText helper with identical display rules -- both covered by tests/reasoning-history.test.js (26/26: collapsed default, toggle, literal-text safety, history restore, save round-trip). No notification, push, jobs, cron, or billing code touched.
NOTIFY-WHEN-DONE BACKGROUND PREDICTIONS. Bell button next to send (gold when armed, choice persisted per browser); armed sends go to POST /api/jobs (sign-in required, 3 pending cap) instead of the blocking /api/chat, and the every-minute cron runs them through the real pipeline via internal replay (same sim, same [PICK], same credit deduction), saves the turn server-side, and fires a payload-less Web Push ticket. New routes /sw.js, /api/push-public-key, /api/push-subscribe, /api/jobs, /api/jobs/mine, /api/job; new secrets VAPID_PUBLIC_KEY, VAPID_PRIVATE_KEY, VAPID_SUBJECT. Open page polls every 15s and renders finished jobs; tapped notifications deep-link /?job=. Follow-ups, same release train: jobs/push routes resolve their own session (fixed 500s on phones), live traffic rolled back to the message-queue build after the first notify builds broke core chat on mobile, bell moved to the History drawer footer with silent background runs, bell-off runs zero notification code (startup catch-up only for browsers that ever used background), and the thinking dots keep watching until an answer actually renders instead of clearing early. Verified live end-to-end (enqueue 200 through cron done in ~70s with reply saved to convo history).
TOOL-CALL TUNING + EDGE GATE + UFC FALLBACK. (1) LEAGUE ENUMS: every TOOLS_OPENAI league field now carries a JSON Schema enum (was prose-only lists the model had to parse), and the schedule/sim/injury lists gained the missing NCAAF/NCAAB values — all enum values verified against normalizeLeague. Legacy Gemini functionDeclarations left untouched. (2) SPARK SAMPLER: Muse Spark requests now send temperature 0.2 for more reliable tool arguments; reasoning_effort stays high. (3) PARALLEL HINT: system prompt rule 9 now says independent calls (schedule plus both sides' injury reports) may go out in one parallel block; schedule -> sim stays serial. (4) ARG-ERROR RETRIES: malformed tool arguments now return a retryable ERROR to the model instead of silently executing with empty {} (both Grok and Meta loops). run_matchup_simulation description intentionally NOT trimmed. (5) EDGE GATE: new deterministic bet/hold/pass layer (edgeGate/edgeGateVerdict + odds converters, pure functions). BET at +3 pts edge or better with a half-Kelly stake; HOLD within \u00b13 with a target price; PASS when the market sits 3+ above SAM or on coin flips. Verdict appended inside the Market Line block on fresh and cached paths so it persists in Notes with no Airtable schema change, worded to never false-fire the faded_market postmortem parser. (6) UFCSTATS VIA ANTIBOT IN FALLBACK: fetchUFCFallbackFighter now tries ufcstats.com first through the existing ZenRows-then-ScrapingAnt chain (was fightingdata/ultimatefightingstats only, env never threaded through), then fightingdata, then ultimatefightingstats — returns null fast when keys are absent so nothing else changes. (7) UFC CRASH FIX: the ESPN fallback path reassigned a const espnResult (TypeError whenever thin tape met a found fallback) — now let.
SIDE NAV IS DARK PURPLE WITH GOLD TYPE. The hamburger drawer background is a dark Parlay Kings purple; MENU, section labels, and links are gold. On parlaykings.net GRIMALDI.TV sits under Trivia instead of at the top of the podcast stack. ai.grimaldi.tv still has HOME in the podcast list.
UFC SIM GETS THE PREDICTIVE TAPE WE ALREADY FETCH. Significant strike differential (SLpM minus SApM) now swings decisions. Career finish mix (KO/sub/decision shares from the ufcstats fight log on the same page) scales KO vs submission chances. Home-country vs event country is a small cage boost when ESPN has both. Weight-class finish bias: heavyweights KO more, fly/straw less. Broadcast records (most wins, fastest KO) stay out of the 25k.
BASKETBALL FOUR FACTORS IN THE SCORE MODEL. NBA/WNBA/CBB now pull the ESPN team sheet (same one-call-per-side as football): eFG%, TOV% from estimated possessions, OREB vs the opponent's DREB, FT rate and FT%, 3P%, pace, AST/TO. Home court is +3 NBA/WNBA and +4 CBB. No PER/BPM/usage, no possession-by-possession rewrite.
FOOTBALL SCORE MODEL GETS THE ESPN SHEET ITEMS WE CAN ACTUALLY USE. Rush YPC and pass YPA (means — ESPN has no play-level stdev), sack % per dropback, INT % per pass, fumbles lost per play (still luck-regressed), FG% by <40 / 40-49 / 50+, XP%, punt net and inside-20, kickoff touchback % and return avg, seconds per play, 4th-down attempts/game. Home NFL side gets +2 expected points (+2.5 CFB). No drive sim, no EPA/DVOA, no player-level.
UFCSTATS.COM IS BACK AS A TAPE ENRICHMENT, VIA ZENROWS THEN SCRAPINGANT. ESPN is still the default UFC path. If SApM, strike defense, or takedown defense are missing, the Worker loads ufcstats.com with ZenRows js_render+antibot, then ScrapingAnt browser=true — not a raw fetch. Those three numbers now move the 25k (chin/defense, stuffing takedowns, KO chance). Challenge pages are rejected and cached fighter pages are reused.
FOOTBALL SIM USES THE ESPN TEAM-STAT SHEET, NOT JUST PPG. Same team-statistics payload now feeds yards/play, pass YPA, rush YPC, third/fourth down, red zone, time of possession, sacks taken vs created, tackles for loss, penalties, FG%, and explosive plays. Early-season samples are shrunk toward league average. Wind from the venue forecast now sets windMph on both sides (the old hook never fired). Combined football factor is capped 0.90-1.10 so scores stay in the real range.
FOOTBALL SCORES STAY IN A REAL RANGE + TURNOVERS IN THE SIM. CFB ESPN standings were feeding season totals (69-45 after two games) as if they were points per game, so the 25k sim printed 99-102 type scores. Football now converts totals to per-game, clamps NFL ~13-38 and CFB ~14-48 expected points, and caps a trial at 55 (NFL) / 70 (CFB). Takeaways and giveaways from ESPN team statistics now move expected scoring (~4.5 points per turnover, luck-regressed).
XAI 2.1 PICKER DESCRIPTION. ParlayKings XAI 2.1 subtitle is now "Sports Analysis Machine (SAM) XAI 2.1 Grok Version".
1.5.9 PICKER DESCRIPTION. ParlayKings 1.5.9 subtitle is now "Sports Analysis Machine (SAM) 1.5.9 original model still most accurate." XAI 2.1 subtitle unchanged.
PASSWORD IS ALWAYS REQUIRED TO SIGN IN. /login and /api/verify-login no longer let a sam_bettors username or email through if password_hash is empty. Blank password is rejected. Login form password field is required. Stripe can still create a row to attach a subscription; that row cannot sign in until a password is set.
HAMBURGER FITS ONE SCREEN. Nav letters are 13px with tighter row padding so the stack stays on one viewport. Lib News is out of the stack. No other menu items added or restored.
CRON IS EVERY MINUTE. Grade Pending bets and re-dump Airtable into the replica on * * * * *. No hourly trigger.
USER-FACING NAME IS PARLAYKINGS. Chat bubble, reasoning toggle, model picker, workflow, Run today status, failover notice, and connecting-error reply say ParlayKings / ParlayKings 1.5.9 / ParlayKings XAI 2.1. Airtable Model Version still writes SAM 1.5.9 / SAM XAI 2.1 so replica and auto-grade keep matching.
SPARK REASONING SET TO LEVEL 4 (high). Meta Muse Spark 1.3 reasoning_effort is now high (1 minimal, 2 low, 3 medium, 4 high, 5 xhigh). Was level 2 (low). Grok / SAM XAI 2.1 stays at medium.
REPLICA RE-SEEDS FROM AIRTABLE EVERY HOUR. Deletes in Airtable were leaving ghost tennis rows in KV (Michelsen vs Tiafoe still Pending in the replica after the Airtable row was gone, so tennis.grimaldi.tv showed a duplicate). Cron now grades then dumps the live table into the replica so the sheet ping and sport sites match Airtable. Google Sheet public CSV is still unpublished (401); the Drive file itself is updated via Apps Script from bets-replica.json.
MARKET LINE BACK ON THE REPLY. ESPN's scoreboard JSON no longer includes moneylines; they live on summary pickcenter (DraftKings homeTeamOdds.moneyLine). Flattening UFC/tennis cards was also dropping event id, so the summary fetch never ran. Odds lookup now keeps the event id, reads pickcenter first, then The Odds API. Cached replies keep the stored Market Line when stripping Auto-graded notes.
RUN ALL IS EVERY SPORT TODAY, NOT DWCS. Slate still pulls MLB, NFL, NCAAF, NBA, NCAAB, NHL, WNBA, UFC, Tennis, WTA, EPL from ESPN scoreboard?dates=YYYYMMDD. Dana White's Contender Series / DWCS is not a numbered UFC card and is skipped. Already-logged games (any Result, not just Pending) are leftovers-skipped so a second click doesn't look like "UFC only" after MLB already wrote. In-progress UFC fights are skipped.
CLOSE GAMES KEEP THE FULL WRITEUP. 50/50 is still gone — tennis/UFC always name a winner. The model was treating that as permission to skip stats and dump a short coin-flip blurb. Tool + prompt now require the four-part writeup even on a 51-49: game summary, cited hold%/tape/rank/W/L, why the higher side wins, then the 25k split that sums to 100.
DELETED AIRTABLE ROWS CAN BE RE-RUN. Cache/dupe checks were treating the KV replica as proof the pick still existed, so deleting Michelsen vs Tiafoe (or any row) and asking again served the old writeup and skipped the new log. Cache hits now GET the Airtable record; 404 drops it from the replica and logs fresh. Also strip Auto-graded / POSTMORTEM lines from cached user replies — those belong in Notes, not the chat.
/stats READS AIRTABLE LIVE. The public stats page and Recent Picks were reading the KV replica, so grades (including today's 50/50 Push rows) could lag the Ai Grimaldi Bets table. Both now paginate Airtable appnbWauDhdcM7mS9 / tbl9X3RoO9uY4jOhL. Sport sites still use the replica.
PARLAYKINGS.NET NAV TRIM. On parlaykings.net / parlaykings.app the hamburger drops Live Broadcast, Episodes, Photo Gallery, Podcast Merchandise, Clip Vault, Workflow, Blog, and the nav Run Today link. Launch Sports Analyzer is labeled ParlayKings.net. HOME is labeled GRIMALDI.TV. drewgrimaldi still gets the header Run today / Run All button on every host. ai.grimaldi.tv keeps the full Grimaldi network menu.
UFC WORKFLOW SHOWS TALE OF THE TAPE. /workflow now has a UFC path: ESPN age/height/weight/reach/stance/SLpM/TD/sub/last-5, those numbers run every round of the 25k sim, and the answer prints TALE OF THE TAPE plus WHY THE SIMULATION PICKED. If the model omits the tape, the Worker appends the sim's tape. ESPN physicals are kept when fight stats fall back to fightingdata.
CLOSE GAMES STILL PICK A SIDE. Tennis/WTA/ITF and UFC no longer lock Picked Winner to 50/50 when the sim is even or data is thin. Yesterday about a quarter of the slate logged 50/50 and graded Push. The 25k sim still prints the real split (51-49 stays 51-49). [PICK] must name a winner — higher sim side, first-listed on an exact 50-50. Already-logged 50/50 rows still grade as Push.
LOGIN CHECKS PASSWORD WHERE ONE IS SAVED. /login and /api/verify-login accepted any existing sam_bettors username or email with no password check. Both now verify the password against password_hash in Supabase sam_bettors with PBKDF2-SHA256 on any account that has one saved, before creating a session or returning ok:true. Accounts with no password saved keep the old behavior until one is set.
FOOTBALL SIM PLUMBING: OPPONENT ADJUSTMENT + TURNOVER REGRESSION + WIND. calculateLambdaMultiplier accepts three new optional inputs, all neutral by default so live picks are unchanged: sosDiff, turnoverMarginPerGame (regressed 70% to the mean at ~4.5pts per turnover), and windMph (0.985 at 15+mph, 0.97 at 20+mph, both sides). New tunables: FOOTBALL_WIND_15_FACTOR, FOOTBALL_WIND_20_FACTOR, SOS_POINTS_TO_FACTOR, TURNOVER_LUCK_DAMPEN, POINTS_PER_TURNOVER.
UFC TALE OF THE TAPE DEEP INTEGRATION. (1) SIMULATION ENGINE: The 25,000-trial UFC Monte Carlo now deeply factors the full Tale of the Tape into every round — physical frame leverage (height & reach combined distance control, center of gravity on takedowns), age dynamics (youth & athletic prime advantage, 35+ chin durability degradation and late-round cardio fatigue), and stance clash (Southpaw/Switch open-stance angles). (2) EXPLANATION & WRITEUP: Tale of the Tape advantages are explicitly calculated, reported, and analyzed in the prediction breakdown alongside striking and grappling metrics.
SAM 1.5.9 (META AI MUSE SPARK 1.3) IS THE MAIN FRONTIER MODEL. (1) FRONTIER SWITCH: SAM 1.5.9 (Meta AI Muse Spark 1.3 via META_API_KEY) is the default frontier model in the picker and for unspecified traffic on ai.grimaldi.tv and parlaykings.net. (2) REPLICA FEED: SAM 1.5.9 is the public replica provider. (3) SECOND OPINION: SAM XAI 2.1 (Grok 4.6) is the second opinion option in the picker. Cached lookups and auto-grade still skip the LLM.
SAM META 1.0 AUTOMATIC FAILOVER ON SAM XAI CREDIT RUNOUT. If SAM XAI 2.1 (Grok 4.6) encounters credit exhaustion, billing issues, or quota limits (HTTP 402/403/429), SAM META 1.0 (Meta Muse Spark 1.2) automatically jumps in first to finish the predictions and write them to the Ai Grimaldi Bets Airtable base and public replica.
SAM XAI 2.1 IS THE MAIN FRONTIER MODEL — ALL PREDICTIONS PULL OVER TO AI GRIMALDI BETS. (1) FRONTIER SWITCH: SAM XAI 2.1 (Grok 4.6) is the default frontier model in the picker and for unspecified traffic on ai.grimaldi.tv and parlaykings.net. (2) SINGLE AIRTABLE BASE: All models log into the single official "Ai Grimaldi Bets" base (appnbWauDhdcM7mS9), distinguished by Model Version. Other bases were deleted. (3) REPLICA FEED: SAM XAI 2.1 is the public replica provider, feeding sport sites (ufc, tennis, football, mlb, wnba, basketball). (4) SLATE: Run today runs SAM XAI 2.1 first to feed the public book and sport sites. (5) NAV CLEANUP: Removed Press & Upcoming Guests from all navigation menus.
RUN TODAY SIMS ALL THREE MODELS. Gemini (SAM 1.5.8) still goes first and is the only write that feeds sport sites (ufc / tennis / football / mlb / wnba / basketball.grimaldi.tv) via the Gemini replica. Then SAM XAI 2.1 and SAM META 1.0 each log the same slate into their own Airtable bases. Click Run today again for leftovers per model. Credits stay untouched. Cached lookups and auto-grade still skip the LLM.
THREE AIRTABLE BOOKS, ONE PER MODEL. SAM 1.5.8 (Gemini) writes AI Grimaldi Bets - Gemini — that stays the public replica / sport-site / Run today feed. SAM XAI 2.1 writes AI Grimaldi Bets - Grok. SAM META 1.0 writes AI Grimaldi Bets - Meta Ai. Asking Grok about a Gemini-logged game runs a fresh sim instead of serving the Gemini cache. Auto-grade walks all three bases. SAM 1.5.8 remains the frontier default and Run today always uses Gemini.
SAM 1.5.8 NOW RUNS gemini-3.8-flash (was gemini-3.7-flash) AND IS THE DEFAULT. Picker default stays SAM 1.5.8, labeled Gemini - Original Model. Thinking stays MEDIUM. Same intro rate as 3.7 ($0.75/$3.75 per 1M through end of 2026). SAM XAI 2.1 stays Grok, SAM META 1.0 stays Meta AI. Run today still uses 1.5.8. Cached lookups and auto-grade still skip the LLM.
PICKER SUBTITLES NAME THE ENGINES. SAM 1.5.8 now reads Gemini, SAM XAI 2.1 reads Grok, SAM META 1.0 already reads Meta AI.
THREE MODELS IN THE PICKER. SAM 1.5.8 is Gemini again (gemini-3.7-flash, native tool loop). SAM XAI 2.1 stays Grok 4.6. New third option is SAM META 1.0 (Meta Muse Spark 1.2 via META_API_KEY). Default and Run today stay SAM 1.5.8. Same picker on parlaykings.net and ai.grimaldi.tv (one Worker). Cached lookups and auto-grade still skip the LLM.
SAM 1.5.8 IS THE MAIN FRONTIER MODEL FOR ONE MORE DAY. Picker default is SAM 1.5.8 (Meta Muse Spark 1.2), labeled Frontier SAM Model. Unspecified traffic and Run today / Run All use 1.5.8. SAM XAI 2.1 (Grok 4.6) stays in the picker as the second-opinion option so Drew can compare another day of data. Cached lookups and auto-grade still skip the LLM.
SAM XAI 2.1 IS THE MAIN FRONTIER MODEL. Picker default is SAM XAI 2.1 (Grok 4.6), labeled Frontier SAM Model. Unspecified traffic and Run today / Run All use XAI 2.1. SAM 1.5.8 (Meta Muse Spark 1.2) stays in the picker as the second-opinion option. Cached lookups and auto-grade still skip the LLM.
PARLAYKINGS.NET SEO. Title, description, keywords, canonical, Open Graph, JSON-LD, robots.txt, and sitemap.xml so Google can find Parlay Kings for sports, sports betting, gambling, DraftKings, FanDuel, and MGM searches. Homepage copy: best sports gambling analysis on the web.
PREMIUM SPORTS ACCESS EXPIRES WITH THE MONTH. premium_until already gates SAM. Sport-site cookies now expire on that date, verify-login returns premiumUntil, and /api/sports-access lets each sport site revoke leftover localStorage the day after the month ends. Pack buyers still have no sports access. Lifetime / drewgrimaldi / Grimillionaire stay.
$5 / 5-PACK ON STRIPE. Clicking Predictions Remaining now offers 5 Predictions for $5 alongside the $10 / 10-pack and $25 Premium. Same Checkout flow as the 10-pack: paid session credits +5 to sam_bettors.credits_remaining, packs never expire.
PARLAYKINGS.NET GETS THE FULL GRIMALDI.TV HAMBURGER. Same podcast + sports list as the rest of the network (HOME, Live, Episodes, Gallery, Blog, Merch, Clip Vault, Press, sports sites). At the bottom: Parlay Kings links — Stats, Workflow, About, Version Notes, Sign In/Out, and Run Today for Drew.
UFC 50/50 LOCK (same rule as tennis). After the 3-round vs 5-round fix, the Aug 22 UFC card still went 5-5 because several fights were 53-54% coin flips that still got a crowned winner. UFC now matches tennis: if the sim is 54-46 or closer (or UFC_DATA: coin), log Picked Winner = 50/50, confidence 50, do not auto-grade Win/Loss (Push once final). Real edges still pick a side. 5-round title/main events keep using the ESPN scheduled round count.
SAM 1.5.8 SPARK REASONING STAYS AT LEVEL 2 (low). Drew tried level 3 (medium) then kept level 2. Meta Muse Spark 1.2 reasoning_effort is low (1 minimal, 2 low, 3 medium, 4 high, 5 xhigh).
WRITEUP OPENS WITH A GAME SUMMARY, THEN STATS, THEN WHY THE PICK. Every prediction now leads with a short 2-4 sentence summary of this game, then the stats, then a detailed explanation of why that team or player wins (the actual edges from the tools), then the win-percentage split. Run today / Run All always uses SAM 1.5.8 (Spark), even if the picker is on Grok.
SAM 1.5.8 SPARK REASONING SET TO LEVEL 2 (low). Meta Muse Spark 1.2 reasoning_effort dropped from medium to low (Meta's scale: 1 minimal, 2 low, 3 medium, 4 high, 5 xhigh). Cheaper and faster on the main model. Grok / SAM XAI 2.1 stays at medium.
SAM 1.5.8 IS THE MAIN MODEL. Picker default is now SAM 1.5.8 (Meta Muse Spark 1.2), described as Frontier SAM Model. SAM XAI 2.1 (Grok 4.6) stays in the picker as the second-opinion option. Cached lookups and auto-grade still skip the LLM.
SAM 1.5.8 NOW RUNS META MUSE SPARK 1.2 (was gemini-3.7-flash). Flagship stays SAM XAI 2.1 (Grok 4.6) and remains the picker default. Label stays SAM 1.5.8. The 1.5.8 slot uses Meta Model API (muse-spark-1.2 at /v1/chat/completions) with the same OpenAI-compatible tool loop as Grok (schedule → sim → [PICK]), reasoning_effort medium. Cached lookups and auto-grade still skip the LLM. Internal provider key stays "gemini" so second opinion, failover, and Airtable Model Version keep working without a UI change.
FOOTBALL ON /stats. NFL and NCAAF were already in the graded log (40 + 6 rows) but the By Sport list labeled them only "NFL" / "NCAAF", so the word Football never appeared. Those rows now read Football (NFL) and Football (NCAAF). Same records, same win rates.
LOGIN RESTORED ON sam_bettors. Sport sites and grimaldi.tv were POSTing /api/verify-login to Pages Functions that were not running (405 empty body, then the UI said Connection error). Login now checks Supabase sam_bettors on this Worker at /api/verify-login: username or email row exists = in. Password hashes on that table were never Grimaldi.tv's unlock, so membership is the check. Network sites call https://ai.grimaldi.tv/api/verify-login so they do not depend on Pages Functions.
GRIMALDI.TV → PARLAYKINGS.NET SSO. Signing in on grimaldi.tv now follows you to parlaykings.net. The old session cookies were Domain=.grimaldi.tv, so the browser never sent them to the .net host. Launch Sports Analyzer now goes through /sso/handoff (reads the grimaldi session) then /sso/consume (sets sam_session on parlaykings.net). Token is HMAC-signed and dies in 60 seconds. Logged-out clicks still land on parlaykings.net unsigned.
RHO IS FITTED, NOT 0.28. MLB/NHL/Soccer bivariate Poisson correlation was a leftover 0.28 placeholder. Fitted from two seasons of completed ESPN scores (not Airtable): MLB 2025-26 regular season n=4510 residual r=0.02; NHL 2024-26 regular n=2632 residual r negative so clamped to 0 (this construction cannot take negative correlation); EPL 2024-26 n=760 residual r=0.01. Airtable is a pick log, not two seasons, and has no NHL rows.
SHARE CARD. Sharing parlaykings.net now shows a large Parlay Kings logo (1200x630) on iMessage, X, Facebook, Slack, and the rest. Served at /og.jpg with Open Graph + twitter:summary_large_image.
iOS HOME-SCREEN TITLE. apple-mobile-web-app-title was still UNCLE SAM under the Parlay Kings icon; it now says Parlay Kings. PWA theme/background colors match the purple.
PARLAY KINGS PURPLE BACKGROUND + .NET DOMAIN. Chat/login/stats pages drop the Uncle Sam poster for the deep purple from the Parlay Kings shield. parlaykings.net and www.parlaykings.net are Worker custom domains alongside parlaykings.app.
GOLD BAR. The red-white-blue top stripe (and the matching strip under the logo) is now gold, matching the Parlay Kings crown.
PARLAY KINGS BRANDING ON THE APP + TENNIS AUTO-GRADE ACTUALLY GRADES. (1) Header no longer shows the Uncle Sam hat, "The Best Sports Analysis On The Web", or "Sports Analysis Machine · SAM XAI 2.1". Favicon, apple-touch icon, PWA icon, and the logo up top are the Parlay Kings mark Drew bought (the PNG in AIProject). (2) Tennis auto-grade was fetching ESPN scoreboards with a dates=RANGE, which returns zero events. Live ATP/WTA boards are tournament-scoped and already include finished earlier-round matches. Grader now fetches the live board (no dates), dedupes the same competition when it appears twice, and treats a walkover as Push. Older US Open qualifying rows that sat Pending were graded in Airtable.
SECURITY + ACCURACY PASS (version held at SAM 1.5.8 per Drew's call — label unchanged). (1) TWO OPEN ADMIN ENDPOINTS CLOSED: /admin/trigger-autograde and /admin/clear-pending-mlb (plus its /api/clear-pending-mlb alias) were sitting ABOVE the session/auth checks in the fetch handler and had no auth at all — anyone who knew the path could hit them with a plain GET, and clear-pending-mlb mutates data (wipes pending MLB bets). Both now require admin auth before doing anything. Behavior change: those two URLs now need ?key=
SAM 1.5.8 NOW RUNS gemini-3.7-flash (was gemini-3.6-flash). Same paid intro rate as 3.6 ($0.75/$3.75 per 1M through end of 2026) — not a cheaper swap, a better Flash for the tool loop (schedule → sim → [PICK]). Thinking stays MEDIUM. The 25k Monte Carlo is still Worker code (run_matchup_simulation), not the model. Cached lookups and auto-grade still skip Gemini entirely. Label stays SAM 1.5.8.
$10 / 10-PACK CHECKOUT NOW CREDITS THE ACCOUNT. Stripe Checkout for /buy?tier=10 was already creating a live $10 session, but paying it never wrote credits_remaining: the success URL dropped the session id, STRIPE_WEBHOOK_SECRET is not set on the Worker, and no sam_bettors row has ever received pack credits. Return from Checkout now retrieves the paid session from Stripe and applies the same unlock as the webhook (10 credits for the $10 pack, or 30-day premium for $25), with a KV lock so webhook + return cannot double-credit.
NETWORK HAMBURGER MATCHES GRIMALDI.TV. SAM's slide-out menu now uses the same starred order as the rest of the network: HOME, ⭐ LIVE BROADCAST, ⭐ LAUNCH SPORTS ANALYZER, Football, Baseball, Basketball, UFC, WNBA, Tennis, Lib News, Trivia, then SAM-only links (Stats, Workflow, About, Version Notes).
TENNIS AUTO-GRADE CATCH-UP. Older WTA/ATP rows were sitting Pending after the match was final because (1) ESPN tennis scoreboards return the whole tournament, (2) Event Date was often the tournament start or an earlier-round date, and (3) names like Lili/Lilli and Lin Zhu/Zhu Lin missed the exact-string match. Grader now range-fetches tennis ~16 days past Event Date, matches last names and reversed given/family order, and processes oldest Pending first. get_game_schedule prefers a still-scheduled match over a completed earlier-round result so the logged Event Date is the actual match.
TENNIS IS HOLD SERVE FIRST (Drew, college tennis at Eastern Illinois). If they cannot break you they cannot beat you — the sim now uses real service-game hold % from TennisMyLife match tapes (service games minus breaks against), not a rank-guessed hold. Surface-specific hold (clay/grass/hard) wins when there are enough games on that court. Weather actually moves the hold number: wind and rain make holding harder, heat speeds the bounce, humidity fluffs the ball. Indoor still plays faster. Rank/W/L are fallback only when hold sample is thin. Writeups must cite both players' hold %. MLB unchanged.
TENNIS THIN-DATA 50/50 (MLB LEFT ALONE). Drew: leave MLB for a few more days, it is hitting; the leak is weaker tennis where rank/W/L never come back and SAM still crowns a winner off a default #80 hold. Tennis/WTA/ITF only: if either player is missing rank or season W/L, or the hold/break sim is 50-51, lock the logged pick to 50/50, say why, still write the Airtable row, still serve the cache with Line movement. Do not auto-grade Win/Loss on those rows (Push once the match is final). MLB and every other league still pick a side.
BASKETBALL IN THE SAM MENU. The hamburger sports list now has Basketball under MLB, linking to https://basketball.grimaldi.tv (NBA + NCAAB). Same network link as football, WNBA, tennis, and UFC.
UFC ACCURACY PASS (post-release patch, version held per Drew's call). (1) 5-ROUND FIGHTS WERE BEING SIMULATED AS 3-ROUND FIGHTS: runUFCMultiStatSimulation had the round loop hardcoded to 3, with no exception for title fights or scheduled 5-rounders \u2014 a real bug affecting every main event and championship bout, since decision math, finish-rate odds, and cardio all play out differently across 5 rounds. Added fetchUFCFightRounds(), which looks up the matched fight on ESPN's UFC scoreboard and reads the real round count off the competition's own description field (\u201c5 Rnd (5-5-5-5-5)\u201d) plus its title-fight flag, run in parallel with the existing power-ratings fetch so it costs no added latency. Falls back to 3 rounds if the fight can't be matched on the board (same non-blocking spirit as the rest of the UFC pipeline). (2) REACH/STANCE EDGES WERE FLAT ACROSS ALL WEIGHT CLASSES: a 4-inch reach edge was scored identically for a heavyweight and a flyweight, which isn't physically right \u2014 reach differential matters proportionally more relative to a smaller frame. Added ufcWeightClassScale(), a lookup table (1.6x at strawweight down to 0.65x at heavyweight) now applied to the reach-edge calculation, keyed off the real weightClass field ESPN's athlete endpoint already returns (added to fetchUFCFighterStats's extracted fields \u2014 was being pulled by ESPN this whole time but not captured). (3) NOT FIXED, BY DESIGN: opponent-strength normalization (raw SLpM/TD stats aren't adjusted for the quality of who they were earned against) and ring-rust/layoff time are both real remaining gaps, but neither has a reliable current data source \u2014 ESPN doesn't expose opponent quality metrics, and pulling last-fight date would mean two more sequential ESPN calls per fighter on top of the already-tuned rate-limit-safe chain (the one that took a dedicated bugfix round to get right, see below). Faking either with a heuristic would make picks look more rigorous than they are, so left alone until there's a real data source worth wiring in.
MLB 64% DEFAULT LOCKED OUT. After five straight 8/19 losses, Airtable showed the Monte Carlo at 50-58% while Confidence % and the writeup were almost always 64% (40 of ~108 recent MLB picks). Grok was rounding coin flips up because the prompt said not to hedge and that favorites clear 60-65%. Fix is in code: logged confidence and the visible split are overwritten from the actual run_matchup_simulation numbers (picked side's sim %, both sides sum to 100). Band calibration may still pull a number DOWN, never above the sim. Dummy/tiny-sample starters (Stock 12.1 IP / 3 GS) no longer sneak through the sample-weight backdoor — under 15 IP and 4 GS is unused, matching the Friday rule. Starter lock restored: 0% weight more than 6 hours before first pitch, ramp 6h→3h, full inside 3h. Prompt no longer permits inflating a 52% sim to 64%.
CONSERVATISM SHRINKAGE SET BACK TO 0.50 (0.55 → 0.50). Per Drew's call after MLB looked off. A raw 70% edge reports 60% instead of ~61%, a raw 80% edge 65% instead of ~67%. Same constant as the original conservatism pass; applies to every sport including MLB.
POST-MORTEM TAGS ARE ACTUALLY ACCURATE. Grade-time tags were leaking starter names into LOSS PATTERNS (comma-split turned "actual starter(s): Pintaro" into fake tags) and faded_market was still looking for "SAM and the market disagree" — a phrase the Market Line never writes. Parser now keeps a whitelist (wrong_probable, dummy_era, injured_qb_used, preseason_stats, thin_data_signal, high_confidence_miss, favorite_miss, coinflip_miss, faded_market). faded_market fires from the real line: SAM vs implied % (>=5) or "**+X% edge** over the market" / "**X% below** the market", never "roughly aligned". Historical losses are re-read the same way when TRACK RECORD loads, so old cards count. Autopsy line is "POSTMORTEM: tags — Winner 62% — logged Stock, Pintaro started". If faded_market or thin_data_signal repeats enough in a league, the next sim gets a hard tool-result warning (do not widen a SAM-vs-market gap; sparse/zero stats are a red flag) instead of rewriting the whole engine after one bad night.
WNBA IN THE SAM MENU. The hamburger sports list now has WNBA under UFC/MMA, linking to https://wnba.grimaldi.tv. Same network link as football, tennis, and UFC.
STRIPE PAYMENT ACTUALLY UNLOCKS SAM. Paying the $25 VIP Payment Link or /buy never flipped sam_bettors: the live Stripe webhook omitted checkout.session.completed, the Worker only handled that one event, Payment Links have no client_reference_id, and STRIPE_WEBHOOK_SECRET was never set so every POST 400'd. Webhook now accepts checkout.session.completed and invoice.paid, matches the payer by session id OR email (creates a sam_bettors row on that email if needed), sets premium_until from the Stripe period end (or +30 days / +10 credits for /buy), and verifies events via the signing secret or by retrieving the event from Stripe when the secret is missing.
GOOGLE SHEET STAYS CURRENT WITH THE LIVE REPLICA. Airtable is still the official write. Sport sites already read https://ai.grimaldi.tv/api/bets-replica.json (not Airtable). After every replica save, SAM pings GOOGLE_SHEET_SYNC_URL if set so the Drive sheet Ai Grimaldi Bets replica refreshes immediately. A one-minute Apps Script fallback also pulls the same JSON.
ALL GROK KEYS DEAD → SAM 1.5.8. After every Grok key (primary + backups) fails or is skipped, a new sim goes to Gemini / SAM 1.5.8. Already-run games still serve the stored writeup. The reply notes that SAM XAI 2.1 was unavailable. Next request skips the dead Grok keys for 15 minutes instead of waiting on them again.
GOOGLE SHEET / BETS REPLICA. Airtable stays the official log. Reads (cache, slate leftover check, track record, /stats, sport-site pick APIs) now use a Cloudflare KV copy first so they stop burning the Free-plan 1,000 Web API calls. New picks and grades still write Airtable, then update the replica. Public feed: GET /api/bets-replica.json and GET /api/bets-replica.csv (newlines stripped so Google IMPORTDATA can parse Notes). drewgrimaldi can force a one-time Airtable dump at GET /api/bets-replica/refresh.
ODDS API BACKUP KEY. env.ODDS_API_KEY_BACKUP is tried when the primary Odds API key returns 401/402/429. Only after both keys fail do we skip The Odds API for 15 minutes and fall through to ESPN moneylines. Secret is set via wrangler, never hardcoded.
ESPN MONEYLINE FALLBACK WHEN THE ODDS API IS OUT. The Odds API returned 401 on cache hits (monthly credits gone — they reset on the 1st). Line movement went silent. fetchMarketOddsData now tries The Odds API first; on 401/402/429 or a missing key it reads moneylines off the ESPN scoreboard we already fetch (ESPN BET / DraftKings / FanDuel when present). No new paid key. A 15-minute skip avoids hammering a dead Odds API key. Failed odds still do not print a warning on the user-facing pick.
CACHED REPLY ALWAYS CHECKS LIVE ODDS FOR LAST-MINUTE LINE MOVEMENT. Person B asking "Tigers vs Nationals" after Person A already ran it still gets the identical stored pick/writeup (no second sim). SAM now always fetches the current market line on that cache hit and appends a movement note: same / toward / away from the pick, opening price → now. The previous 800ms skip dropped this too often. If the first log had no Market Line, Person B still sees the current line labeled as checked just now. Opening price is parsed from the stored Market Line and also saved as a ---ODDSLINE--- snapshot on new writes so later lookups do not depend on the regex.
AIRTABLE CACHE SHORTCUT ACTUALLY FIRES FOR ALREADY-RUN GAMES. The whole point of logging picks is to serve the stored writeup instead of re-running verify/stats/25k sim. That path existed but missed most real questions after a slate run: (1) only the first 100 Pending rows were searched, no pagination; (2) two Pending rows for the same MLB series (yesterday + today) counted as "ambiguous" and skipped the cache entirely; (3) Placed Date coming back as a datetime failed the exact === todayET() check so every hit was treated as stale and re-simulated; (4) lookup required the user to type "vs" or "at"; (5) a cache hit still waited on The Odds API (tennis can fan out many 4s calls) plus a non-peek credits check before answering. Fix: paginate today's + pending records, prefer today's unique matchup, match nicknames even without vs/at, compare dates by YYYY-MM-DD, and return the stored reply immediately (line movement has 800ms to append, otherwise skip).
SCRAPINGANT ESPN FALLBACK WHEN ZENROWS IS OUT. Drew has a ScrapingAnt API key to use while ZenRows is unpaid/capped (AUTH004). Added fetchEspnViaScrapingAnt() against https://api.scrapingant.com/v2/general (x-api-key header). tryEspnBypassFallbacks now tries ZenRows first, then ScrapingAnt, then the unused Browser binding. If ZenRows returns 402/AUTH004, a sticky flag skips it for 15 minutes so we don't keep burning failed ZenRows calls. fetchEspnCoreWithRetry (UFC/tennis athlete lookups) now uses the same fallback chain instead of ZenRows-only. Secret: SCRAPINGANT_API_KEY. No-ops until that secret is set.
SAM LOGIN IS sam_bettors MEMBERSHIP. Every row in Grimaldi.tv's sam_bettors table can sign in with that username or email. Password hashes on that table were never the same as Grimaldi.tv's screen-name unlock, so requiring a SAM hash rejected real members (credits and premium included). Lookup is username OR email; if the row exists, SAM starts a session on that bettor id so free/daily, credits_remaining, and premium_until apply as stored. Not on the list = incorrect username. Supabase errors surface as such. drewgrimaldi is still the only admin for Run today.
ADMIN SLATE RUNNER — RUN TODAY'S BOARD INTO AIRTABLE. drewgrimaldi-only. GET /api/slate pulls today's ESPN scoreboards (MLB, NFL, NCAAF, NBA, NCAAB, NHL, WNBA, UFC, Tennis, WTA, EPL), expands UFC/tennis cards into individual matchups, skips completed games and anything already Pending in Airtable, and returns the leftover list. The signed-in admin UI has a Run today button that walks that list one game at a time through the existing /api/chat path (same sim, same [PICK], same Airtable write as typing the matchup). slate:true on /api/chat is rejected unless the session username is drewgrimaldi; when it is, the prediction-credit meter is not touched. One game per Worker request so ZenRows/subrequest limits stay the same as a normal typed prediction.
NFL / NCAAF / NBA / CBB ACCURACY PASS. These are the volume sports. (1) NBA preseason is ignored the same way NFL August is — season_type pre and October games before the 20th do not count as this year's offense/defense. Under 8 real NBA/CBB games (6 WNBA), last year is blended in. (2) Rest was in the lambda math but never set. ESPN last-game lookback now sets restDays for NFL/CFB/NBA/WNBA/CBB. NBA/WNBA/CBB on a 1-day turnaround is treated as a back-to-back and scales offense down. (3) Basketball key-player skips OUT/IR so a hurt star's PPG is not the factor. Tool result prints rest and data vintage.
POST-MORTEM TAGS NOW CHANGE THE SIM, NOT JUST THE PROMPT. On auto-grade of a Loss, SAM compares the logged starter/QB to the actual ESPN box-score starter and writes structured tags (wrong_probable, dummy_era, injured_qb_used, preseason_stats, favorite_miss, plus the old thin_data / high_confidence / coinflip / faded_market). Last 20 losses per league are rolled up when TRACK RECORD loads. If a tag repeats enough: wrong_probable x3 in 6+ MLB losses halves starter-lock weight; dummy_era x2 in 4+ losses raises the IP/GS bar; high_confidence_miss x4 in 8+ losses pulls that league's win% toward 50. Friday-style Stock/Pintaro would tag wrong_probable and then actually tighten the next MLB card. One bad night does not rewrite the engine.
FOOTBALL WEEK-1 HARDENING (NFL + CFB). Season is close; do not let August preseason or a 0-game 2026 slate pretend to be this year. (1) NFL BALLDONTLIE games tagged preseason, week<=0, or dated August/late July are dropped. If fewer than 4 real regular-season games exist, last year is blended in (or used outright if this year is empty). (2) ESPN NFL/CFB standings do the same: under 4 games played, blend last year instead of treating a Week 0/1 sample or empty 2026 table as a full season. (3) QB key-player: injured QBs are skipped so a backup is the input; Preseason splits are never used for QB rating. Tool result now prints the data vintage so a Week 1 pick says "2025 regular season / preseason ignored" instead of looking like live 2026 form.
MLB BULLPEN + DON'T LOCK A STALE PROBABLE. Follow-up to Friday's card. (1) BULLPEN: ESPN team pitching stats are pulled for both clubs. A real relief/bullpen ERA is preferred; if ESPN only exposes overall staff ERA it is used as a weaker staff proxy. That factor is blended into expected runs so the 8th/9th inning is not just "whoever started." (2) STARTER LOCK: listed probable ERA is not applied more than 6 hours before first pitch (0% of the starter weight), only partially from 6h to 3h, and fully inside 3 hours or after first pitch. A Stock-style 10.13 probable the morning of cannot manufacture a 62% pick anymore. Tool result states STARTER LOCKED / NOT LOCKED and the hours to first pitch. Re-run inside 3 hours to lock the actual starter.
MLB PITCHER INPUTS HARDENED AFTER A BAD FRIDAY CARD. Friday's graded MLB slate went 2-5 (likely 2-6): dummy 0.00 ERAs (Jobe), a listed probable who did not start (Stock 10.13 / Pintaro actually threw), and ERA-only weight turning 62% favorites out of mid-rotation matchups. Fix is in the sim inputs, not the prompt. Probable-pitcher lookup now pulls IP/GS/G with ERA. A 0.00, ERA over 12, under 15 IP and 4 GS, or a reliever/opener profile (starts are a small share of appearances) is NOT used to scale expected runs, and the tool result says so. Usable ERAs are sample-shrunk toward 4.20 and capped 0.86-1.16 so one ugly or tiny-sample number cannot manufacture a 62% pick. Star-bat IL weight raised (OF/DH/SS) so a Judge/Stanton/Bellinger-style pile-up actually moves the Yankees-type number instead of getting ignored as three 1.5% OF dings.
UFC TALE OF THE TAPE IS NOW PART OF THE SIMULATION OUTPUT. The fight sim returns the actual tape it used (age, height, weight, reach, stance, SLpM, strike acc, TD/15, TD acc, sub/15, last-5) plus the 25,000-trial win% split by KO/sub/decision and a why line computed from those same inputs (which tape edges moved the number, most common finish path). Not extra summary color — those rows are the simulation's inputs and results.
PREDICTION ENGINES: POSITION-WEIGHTED INJURIES + UFC MULTI-STAT + TENNIS HOLD/BREAK. Three code-level upgrades so winner accuracy comes from better inputs and better sims, not more prompt text. (1) TEAM INJURIES: ESPN reports are now scored by position and status (QB/ace/star OUT haircuts harder than a depth-chart Questionable) and applied automatically to each side's lambda, cap 18%. Simulation writeup names the notable absences. (2) UFC: fights no longer mash tale-of-the-tape into one power number plus N(0,15). New per-round model uses SLpM, strike accuracy, takedown avg/accuracy, submissions, reach, stance, and last-5 form from fightingdata.com (KO/sub finish paths + decision). (3) TENNIS/WTA/ITF: matches no longer use the same mashed fight Gaussian. New hold/break set sim (best-of-3, best-of-5 for ATP slams) with hold probabilities from rank, season W/L, TennisMyLife surface record, indoor/outdoor, recent form, and H2H. Conservatism shrinkage still applied to the raw win% on all three paths.
CONSERVATISM SHRINKAGE SET TO 0.55 (0.50 → 0.55). Per Drew's call after the fresh deploy landed back on 0.50. Same small nudge as the earlier 0.50 → 0.55 pass: a raw 70% edge reports ~61% instead of 60%, a raw 80% edge ~67% instead of 65%. Holding here.
GROK MODEL BUMPED TO 4.6, REASONING EFFORT LOWERED TO MEDIUM. Swapped GROK_MODEL from "grok-4.5" to "grok-4.6" (Drew is getting a fresh API key for it) and dropped reasoning_effort from "high" to "medium" per Drew's request. Also discovered while fixing the base-consolidation bug above that Grok's calibration/track record has zero real history to date — every prior write silently 404'd against a base that never existed — so this reasoning-effort change starts from a clean slate anyway.
REAL ROOT CAUSE FOUND VIA WRANGLER TAIL: ZENROWS CONCURRENCY LIMIT, NOT ESPN BLOCKING. Deployed the ZenRows-first change, ran wrangler tail live against a real WNBA prediction per Drew's request, and got a definitive answer instead of another round of static-code guessing: 8 athlete-overview calls fired, 5 succeeded, 3 came back AUTH006 (Too many concurrent requests) directly from ZenRows' own API -- not ESPN, not a subrequest-ceiling error, a distinct problem from everything chased earlier this session, and one the Cloudflare Workers Paid upgrade does nothing for since it's a ZenRows-side limit, not a Cloudflare-side one. Traced the exact mechanism: getKeyPlayerFactor's candidate-stats loop used Promise.all across 4-6 players, AND the outer call site ran both teams' getKeyPlayerFactor through Promise.all too -- a nested fan-out that, for a WNBA matchup, meant up to 8 simultaneous ZenRows sessions competing for the same concurrency slot the instant today's earlier change (routing every ESPN call through ZenRows first, no direct attempt) went live. That earlier change fixed the original latency problem but directly caused this one -- more calls landing on ZenRows at once, more collisions. Fixed both nesting levels to sequential (await one candidate/team fully before starting the next) rather than parallel. Also converted tennis's two-player Promise.all to sequential pre-emptively, same mechanism at smaller scale (2 concurrent vs. up to 8) -- Drew confirmed tennis is also still broken and this is a cheap, low-risk fix for the same class of bug rather than waiting for a second tail session to prove it independently. NOT YET touched: getInjuryReport's both-teams Promise.all (2 concurrent) -- not implicated in the log, left alone for now; worth revisiting if injury-report-specific errors show up in a future tail session.
ALL ESPN CALLS NOW GO STRAIGHT TO ZENROWS, EVERY SPORT, NO DIRECT ATTEMPT FIRST. Drew confirmed the error was \u201cSAM's having trouble connecting right now\u201d \u2014 the generic message shown when the whole model call fails, not an ESPN-specific error surfacing cleanly. That pointed at latency/timeout rather than a clean 403: every blocked direct fetch was paying a full connect-attempt-then-fail cost before ever falling back to ZenRows, and with WNBA/tennis routing several ESPN calls per prediction, that failed-attempt tax could plausibly stack up into the whole request timing out \u2014 which would explain a generic connection failure instead of a traceable ESPN error. Drew's ask: skip the 403 entirely, wire it all up with ZenRows first. Removed the conditional gating in both fetchEspnWithRetry() and fetchEspnCoreWithRetry() (isTightMarginEspnUrl, isEspnDirectLikelyBlocked's sticky flag, and the forceZenRows flag added earlier today for tennis) \u2014 now both functions try ZenRows immediately whenever env.ZENROWS_API_KEY exists, for every sport, every endpoint, with the direct fetch demoted to a last-resort fallback only if ZenRows itself comes back empty or erroring. isTightMarginEspnUrl() and isEspnDirectLikelyBlocked() are now unused (left in place rather than risk over-editing; harmless dead code, not wired into anything). Cleaned up the now-meaningless forceZenRows argument from all four tennis call sites since every call gets that behavior unconditionally now. This trades a real increase in ZenRows usage (every ESPN call, not just the previously-flagged tight-margin sports) for the fastest possible path per call \u2014 Drew is already on the 45k/month paid tier anticipating this kind of volume.
TENNIS ZENROWS COVERAGE GAP FOUND AND FIXED; WNBA CONFIRMED ALREADY FULLY COVERED. Drew reported still getting errors on both. Pulled the actual live deployed worker directly from Cloudflare first (not just the local file) to rule out a stale-deploy mismatch \u2014 it wasn't one, the live code already has every WNBA/tight-margin fix from this session. Traced every WNBA data path (team resolve, roster, standings, athlete search, athlete stats, schedule, injuries, grading) and confirmed all of them go through fetchEspnWithRetry(), which does have the guaranteed-ZenRows check (isTightMarginEspnUrl) \u2014 no gap found there. Tennis is a different story: its entire pipeline (fetchTennisPlayerStats \u2014 athlete lookup, rank list, rank detail, statistics, four separate calls per player) runs exclusively through fetchEspnCoreWithRetry(), which never got the WNBA-style guaranteed-first-try fix at all \u2014 it only fell back to ZenRows reactively, after a 403 had already been hit once this invocation (the sticky flag), or inline on that same call after a failed direct attempt. Since three of the four tennis calls use ESPN-returned $ref URLs rather than predictable literal paths, pattern-matching them the way isTightMarginEspnUrl does for WNBA would be fragile, so added an explicit forceZenRows parameter to fetchEspnCoreWithRetry() instead and set it true at all four tennis call sites \u2014 tennis has zero BALLDONTLIE coverage and its whole stat pipeline sits on this one function, so it needs the same guarantee WNBA already has, not URL-guessing. NOT YET CONFIRMED: whether this was the actual cause of the WNBA errors specifically, since that path checked out clean \u2014 need the real error text/message from Drew to pin that one down rather than guessing further.
CHAT BACKGROUND SWAPPED TO THE NEW UNCLE SAM POSTER. Replaced BACKGROUND_BASE64 with the newly uploaded promotional artwork. The original PNG upload was 2.5MB \u2014 too large to embed directly without pushing the worker script size up significantly, so it was resized to roughly match the previous background's dimensions (1100px wide, same aspect ratio) and re-encoded as JPEG at quality 78, landing at about the same final size as what was already in place (~316KB vs. the previous ~280KB) rather than ballooning the deployed script. Verified the new base64 decodes back into a valid JPEG before shipping.
NCAA.COM ADDED AS A GENERAL-PURPOSE STAT SOURCE FOR CFB, WIRED IN AS A QB-RATING FALLBACK. Drew asked for \u201call of the stats\u201d from NCAA.com's individual leaderboards, used as a fallback only when ESPN's college football data is thin. Checked the page directly first: it's a plain, simple server-rendered HTML table (no JS, no bot-check) with a consistent structure across every category, and found all 41 category IDs from the site's own dropdown (Rank/Name/Team/Cl/Position/G plus category-specific columns) \u2014 confirmed category 453 in Drew's link is specifically Passing Yards, not Passing Efficiency (the real QB Rating equivalent, which is category 8, \u201cPass Eff\u201d column). Built this as reusable infrastructure rather than a one-off scrape: parseNcaaStatsTable() generically parses any category's table by reading the header row and mapping columns dynamically, so adding coverage for a different stat later is a one-line addition to NCAA_FOOTBALL_STAT_CATEGORIES, not a new parser. fetchNcaaFootballStatCategory() fetches one category by name; fetchNcaaFootballPlayerStat() searches a category's full leaderboard for a specific player by name, disambiguating by team when multiple players share a name. Wired the Passing Efficiency category specifically into getKeyPlayerFactor() for CFB: when ESPN's per-QB stat fetch comes back empty for every candidate on the roster (common for less-marquee college programs ESPN doesn't track closely), falls back to looking up the top-of-roster QB by name and team on NCAA.com's real Passing Efficiency leaderboard, using \u201cPass Eff\u201d as the QBRating-equivalent value and \u201cPass Att\u201d as the usage stat \u2014 same downstream math, just a real number instead of a missing one. Verified the full path (table parsing, name lookup, team disambiguation, and a real not-found case) against live NCAA.com data before shipping.
TWO NEW UFC FALLBACK DATA SOURCES ADDED, BOTH VERIFIED SCRAPABLE WITHOUT ZENROWS. Drew's original ask was whether ufcstats.com could be added now that ZenRows exists \u2014 checked it directly first: it's not a simple IP block like ESPN, it's a client-side JS proof-of-work challenge (computed SHA-256 in-browser), a fundamentally harder bot-check that a plain fetch() can never solve regardless of where it runs from. Left that one alone pending real verification against ZenRows' antibot+js_render mode. Drew then suggested two alternatives instead: ultimatefightingstats.com and fightingdata.com. Confirmed ultimatefightingstats.com's fighter stats are on its free tier (not paywalled) before touching anything \u2014 it's a Next.js app whose fighter pages are server-rendered but the actual data sits inside a React Server Component payload with inconsistent backslash-escaping depth, which made SLpM/Str. Acc./TD Avg. reliably extractable via a backslash-tolerant regex but left reach/age/takedown-accuracy/submission-avg in a more deeply nested structure not worth force-parsing. fightingdata.com turned out to be the stronger source: plain server-rendered HTML with a genuinely simple table structure (
SWITCHED TO NCAAF/NCAAB AS THE CANONICAL COLLEGE LABELS, PER DREW. Found the real control point while making this change: extractPick() already force-normalizes any CFB/NCAAF wording the model writes into one fixed string before it's ever logged to Airtable \u2014 so the model's own phrasing barely matters, this hardcoded rewrite is what actually decides the League field. It was writing \u201cNCAA Mens Football\u201d/\u201cNCAA Mens Basketball\u201d, not the \u201cCollege Football\u201d label added to the [PICK] enum in the previous entry \u2014 those would never have matched each other regardless of what the enum said. Updated both the strict-parse and lenient-fallback versions of this rewrite to output \u201cNCAAF\u201d/\u201cNCAAB\u201d instead. Also found a real functional gap while doing this: normalizeLeague() recognized \u201cNCAAM\u201d as a college-basketball alias but never \u201cNCAAB\u201d at all \u2014 fixed. Updated the [PICK] tag enum and the two tool-parameter descriptions that said \u201cNCAAM\u201d to say \u201cNCAAB\u201d for consistency. Net effect: since nothing has actually been logged for CFB yet (confirmed with Drew, season starts end of August), this lands clean \u2014 no legacy data to reconcile, every CFB/CBB pick from here forward will consistently land under NCAAF/NCAAB in Airtable, TRACK RECORD, and calibration.
GETTING AHEAD OF COLLEGE FOOTBALL BEFORE THE SEASON STARTS. Drew clarified CFB hasn't actually been used yet (no graded picks exist), season starts end of August \u2014 wanted to be proactively ready rather than wait for a live failure the way WNBA's problems had to get diagnosed one at a time. Checked what CFB actually shares with WNBA's risk profile: BALLDONTLIE_SPORT_PATHS only covers NFL/NBA/MLB/NHL \u2014 CFB gets zero BALLDONTLIE coverage, same as WNBA, so its standings always fall through to ESPN. The good news: CFB's key-player loop is inherently lighter than WNBA's \u2014 it filters to QB-only candidates (2-4 per roster), not the 6-player scan NBA/WNBA/CBB use, so it was never going to hit the same subrequest wall from that angle. But the missing-BALLDONTLIE trait on standings is identical, so extended the guaranteed-ZenRows routing proactively: renamed isWnbaEspnUrl() to isTightMarginEspnUrl() and added a check for \u201c/football/college-football\u201d alongside WNBA's \u201c/basketball/wnba\u201d. Verified against real CFB vs. NFL vs. NBA URLs before shipping (CFB/WNBA both match, NFL/NBA correctly don't). Also confirmed against ESPN directly that the season's real schedule is already live and pulling correctly \u2014 99 games found on the default scoreboard starting 2026-08-29, so get_game_schedule's verification gate has real data to check against starting day one, not an empty board. CBB shares the exact same no-BALLDONTLIE / 6-player-loop profile as WNBA too but wasn't in scope for what was asked here \u2014 worth the same treatment if it starts showing the same symptoms.
COLLEGE FOOTBALL: CHECKED WHETHER IT'S WIRED UP LIKE WNBA. Short answer: the simulation/prediction pipeline was already fully wired (ESPN_LEAGUE_PATHS, TEAM_ROSTER_LEAGUES, and KEY_PLAYER_SPORT_PATH all already include CFB, and the odds-mapping aliases for \u201ccollege football\u201d/\u201cncaaf\u201d/\u201cncaa mens football\u201d were already present) \u2014 CFB never had WNBA's missing-odds or missing-ZenRows problems. But checking it directly surfaced a real, separate bug: the [PICK] tag's league enum \u2014 the literal instruction telling the model what to write in the League field \u2014 never included College Football or College Basketball at all, only listing NBA/NFL/MLB/etc. and a generic \u201cOther\u201d catch-all. Confirmed this is a real, active problem by checking the live Airtable League field's schema directly: it has BOTH \u201cCollege Football\u201d and \u201cNCAA Mens Football\u201d as separate select options (same for basketball) \u2014 meaning the model has been writing whichever phrasing it felt like each time, splitting the same sport's graded history across two different League buckets in the TRACK RECORD/byLeague breakdown the model itself reads back on every request. Added \u201cCollege Football\u201d and \u201cCollege Basketball\u201d explicitly to the [PICK] tag enum so future picks converge on one label each. Also added \u201ccfb\u201d/\u201ccbb\u201d as defensive aliases in the odds-mapping table (fetchMarketOddsData's leagueToSport), since pick.league is written to Airtable and used for the odds lookup completely unnormalized \u2014 raw model text, not run through normalizeLeague() the way the internal simulation code paths are. NOT fixed: the existing historic split between \u201cCollege Football\u201d and \u201cNCAA Mens Football\u201d records already in Airtable \u2014 that's a data cleanup Drew would need to do deliberately (merging the two labels), not something to silently rewrite.
TWO REAL FIXES: WNBA MARKET LINE, AND BOTH-TEAM WIN PERCENTAGES. (1) WNBA odds were simply missing \u2014 checked fetchMarketOddsData's leagueToSport mapping (used to translate a league into The Odds API's sport_key) and \u201cwnba\u201d was never in it, so every WNBA lookup hit the null-sportKey guard and silently returned nothing, every single time, regardless of ZenRows/caching/anything else touched today. Added \u201cwnba\u201d \u2192 \u201cbasketball_wnba\u201d (The Odds API's real sport key). (2) Drew wanted both teams' win percentages shown, not just the winner's \u2014 the system prompt's rule 3 example only demonstrated stating one side (\u201cgiving the Blue Jays a 55% win probability\u201d), which is likely why the model usually only stated the winning side. Updated to explicitly require both (\u201cgiving the Blue Jays a 55% win probability to the Royals' 45%\u201d).
WNBA ESPN CALLS NOW GO STRAIGHT TO ZENROWS, ALWAYS \u2014 DETERMINISTIC INSTEAD OF PROBABILISTIC. Drew's ask after the sticky-block flag: don't leave it to a heuristic, just always route WNBA through ZenRows. Fair call \u2014 the sticky-block optimization from the previous entry only helps probabilistically (depends on Cloudflare reusing a warm isolate across requests), which still leaves room for the exact \u201cworked once, then didn't\u201d flakiness Drew was seeing. Added isWnbaEspnUrl() (checks for \u201c/basketball/wnba\u201d in the request URL \u2014 every WNBA ESPN call goes through fetchEspnWithRetry with that path, so this catches all of them: scoreboard, roster, injuries, standings, and key-player stats alike) and wired it into fetchEspnWithRetry so a WNBA call skips the direct-fetch attempt unconditionally and goes straight to ZenRows every single time \u2014 guaranteed 1 subrequest per call instead of a possible 1-or-2. Deliberately scoped to WNBA only, not all ESPN traffic: WNBA is the specific, confirmed tight-margin sport (no BALLDONTLIE coverage the way NBA/MLB/NFL/NHL get), and unconditionally routing every sport through ZenRows would multiply usage against the new 45k/month plan for leagues that aren't actually having a problem. Verified the URL match against real WNBA vs. NBA vs. core-API (UFC/tennis) URLs before shipping \u2014 no false positives. The sticky-block flag from the previous entry stays in place for every other sport.
ESPN \u201cSTICKY BLOCK\u201d DETECTION \u2014 CUTS BLOCKED-CALL COST ROUGHLY IN HALF FOR THE REST OF A REQUEST. Drew's WNBA symptom (\u201cworked once, then didn't, then worked again\u201d) is the signature of a request sitting right at the edge of Cloudflare's subrequest ceiling \u2014 not a hard bug, a budget that's genuinely borderline depending on which specific ESPN calls happen to get blocked that particular request. Every ESPN call that gets 403'd was costing 2 subrequests (the failed direct attempt, then the ZenRows retry), for every single call, even though the block is IP-based and effectively all-or-nothing \u2014 once one call gets blocked, the rest almost certainly will too. Added a lightweight, fail-safe optimization: a module-level espnDirectBlockedUntil timestamp. The first ESPN call in an invocation still attempts a direct fetch as normal (so a genuinely-unblocked window isn't wasted); if that gets a 403/520/429, every subsequent ESPN call in that same request (and likely nearby ones too, since Cloudflare commonly reuses warm isolates across close-together requests) skips the doomed direct attempt entirely and goes straight to ZenRows for the next 4 minutes \u2014 1 subrequest instead of 2. For a WNBA prediction making 15-20+ ESPN calls, if even half hit this path that's roughly 7-10 fewer subrequests, real headroom back under the ceiling. This is a soft, probabilistic optimization, not a guarantee \u2014 isolate reuse across requests isn't something Cloudflare promises, so the flag may not always carry over between separate invocations, but it can only help within a single request and never makes anything worse (the very next direct attempt after the sticky window expires re-establishes ground truth normally).
WNBA FRESH PREDICTIONS FAILING OUTRIGHT \u2014 A REAL REGRESSION FROM TODAY'S OWN ZENROWS EXTENSION, NOW MITIGATED. Drew reported the same \u201chaving trouble connecting\u201d error but this time for a brand-new (non-cached) WNBA prediction, not a lookup. Root cause is a real tradeoff introduced earlier today: extending ZenRows to rosters, injuries, standings, and player stats fixed those endpoints returning null on a 403, but it also means every blocked call now costs 2 subrequests (the failed direct attempt plus the ZenRows retry) instead of 1. WNBA was already documented as the tightest-margin sport for Cloudflare's per-invocation subrequest ceiling (no BALLDONTLIE coverage the way NBA/MLB/NFL/NHL get, and an earlier session had already cut its key-player candidate fetch from 15 to 6 for exactly this reason) \u2014 today's fix traded silently-degraded data for a request that fails outright once the doubled cost tips it over. Two real fixes, not guesses: (1) found that both teams' roster lookups, and separately both teams' standings lookups, were being fetched in Promise.all \u2014 fully parallel \u2014 even though the underlying team-list and standings calls are the exact same URL for both teams in a given league. Parallel execution meant neither could benefit from the other's cache write (both check cache, both miss, both fetch independently). Switched both to sequential awaits so the second team's lookup now hits the cache the first one just populated, cutting a full duplicate ESPN call (and its possible ZenRows retry) out of every single team-sport prediction, not just WNBA. (2) Cut WNBA's key-player candidate cap specifically from 6 to 4 (NBA/CBB keep 6, they have more budget headroom) since that stat-fetch loop was already flagged as the single biggest subrequest sink before today, and today's change quietly doubled it. Not able to confirm this fully resolves it without a live invocation to test against \u2014 flagged to Drew as the next thing to watch.
FOUND THE ACTUAL REASON WNBA KEPT GETTING STUCK: A THIRD, SEPARATE \u201cat\u201d BLIND SPOT, THIS TIME IN GRADING ITSELF. Drew reported WNBA \u201cstill struggling\u201d a day after the fast-path and dupe-check \u201cat\u201d fixes. Checked live Airtable data directly: all 4 of the previous night's WNBA games were confirmed STATUS_FINAL on ESPN, but only 2 of 4 Pending records had been graded \u2014 the exact 2 that stayed Pending were both logged with full team names and \u201cat\u201d (\u201cPhoenix Mercury at Atlanta Dream\u201d, \u201cDallas Wings at Washington Mystics\u201d), while the \u201cvs\u201d-phrased ones graded fine. Root cause: splitMatchupLabel(), the function gradePendingBets() uses to break a Bet Label into two team names before calling findFinalResult(), only split on \u201cvs\u201d/\u201cv\u201d/\u201cv.\u201d \u2014 same blind spot as the earlier two fixes, but in a completely different function nobody had checked yet. When it returned null for an \u201cat\u201d-phrased label, gradePendingBets() just did \u201cif not sides, continue\u201d \u2014 silently skipped that record, forever, every single cron cycle, with no error or log line to notice by. Fixed the same way as the other two: widened the separator regex to also accept \u201cat\u201d. Verified against the real stuck labels before shipping. Manually graded the two records stuck from this bug directly in Airtable so Drew didn't have to wait for the next cron cycle: Phoenix Mercury at Atlanta Dream \u2192 Win (Atlanta Dream 96-82), Dallas Wings at Washington Mystics \u2192 Loss (Mystics won 96-92). LESSON FOR NEXT TIME: there are now three independent places that had to each separately learn to recognize \u201cat\u201d as a matchup separator (fast-path lookup, write-time dupe check, and this grading parser) because each was written as its own regex rather than calling one shared parser. If a fourth one turns up, the real fix is consolidating all matchup-label parsing into a single shared function instead of patching each site individually as it's discovered.
REMOVED REMAINING USER-FACING GEMINI/GROK BRANDING. Drew asked whether provider names had already been scrubbed from the model switcher; checked the live deployed code directly and found four spots that still said it outright: the model-picker dropdown subtitles (\u201cGemini-powered \u2014 the original model\u201d / \u201cSecond opinion \u2014 powered by Grok\u201d), the cross-model failover notice shown to a user when their requested provider is down (\u201cGemini was unavailable, so Grok answered this one instead\u201d), the \u201cget a second opinion\u201d button's tooltip (\u201cGet Grok's take\u201d), and the Workflow page's step-by-step explainer (\u201cGemini or Grok, whichever you picked\u201d). All four now reference SAM 1.5.8 / SAM DS 2.1 instead of the underlying provider names. Internal-only references \u2014 function names (callGemini/callGrok/requestGemini/requestGrok), changelog history, and code comments \u2014 were intentionally left alone; those aren't visible to a user.
SAME BUG, DEEPER ROOT CAUSE: MATCHUP DEDUPE WAS EXACT-STRING-MATCH EVERYWHERE, CAUSING BOTH DUPLICATE PENDING RECORDS AND MISSED CACHE HITS. Drew hit the same \u201cSAM's having trouble connecting\u201d error again on \u201cPhoenix Mercury at Atlanta Dream\u201d even after the earlier \u201cat\u201d-regex fix. Investigating the live Airtable data directly surfaced the real, deeper problem: there were TWO separate Pending records for the exact same real game \u2014 \u201cMercury vs Dream\u201d and \u201cPhoenix Mercury at Atlanta Dream\u201d both existed as distinct rows, same for Dallas/Washington and Storm/Liberty. Root cause: THREE separate places all compared Bet Label text with strict, brittle exact-string equality \u2014 checkExistingPrediction (fast path), findExistingPendingRecord (post-model backstop), and critically the write-time dupe check inside _logBetToAirtableInner itself. Whenever the model phrased a matchup differently between two requests (full team names vs short nicknames, \u201cvs\u201d vs \u201cat\u201d), the write-time dupe check failed to recognize the existing Pending record and silently created a second one for the same real game \u2014 which is also exactly why the \u201cat\u201d-regex fix alone wasn't enough: even with a matching hint, there could now be a differently-worded duplicate row that a strict equality check still wouldn't find. Fixed properly this time with a shared, non-string-literal comparison: teamNickname() extracts each team's last significant word (Seattle Storm to storm, Atlanta Dream to dream), matchupNicknameKey() builds a sorted pair from both sides of a vs/v/at-separated matchup string, and sameMatchup() compares two labels by that key instead of raw text. Wired into all three call sites, replacing every prior exact LOWER({Bet Label}) = \u201c...\u201d Airtable formula with a broader Result=Pending (+ Event Date where available) fetch filtered client-side by sameMatchup(). Verified against the real duplicate pairs pulled from Airtable before shipping: \u201cMercury vs Dream\u201d / \u201cPhoenix Mercury at Atlanta Dream\u201d, \u201cStorm vs Liberty\u201d / \u201cSeattle Storm vs New York Liberty\u201d, and \u201cDallas Wings vs Washington Mystics\u201d / \u201cDallas Wings at Washington Mystics\u201d all now correctly recognize as the same game, while unrelated matchups (e.g. Mercury/Dream vs Wings/Mystics) correctly stay distinct. KNOWN FOLLOW-UP: the duplicate Pending records already sitting in Airtable from before this fix were not automatically merged \u2014 Drew was flagged to review and clean those up manually so grading doesn't double-count them.
WNBA \u201cALREADY LOGGED\u201d LOOKUP FAILING FOR \u201cAT\u201d-PHRASED QUERIES; TRACED TO A REAL BUG, NOT A REGRESSION FROM TODAY'S OTHER WORK. Drew reported \u201cSAM's having trouble connecting\u201d for \u201cSeattle Storm at New York Liberty\u201d even though that exact matchup was already logged Pending. Root cause: the fast-path cached-lookup regex (vsMatch) only recognized \u201cvs\u201d / \u201cv\u201d / \u201cv.\u201d as a team separator in the user's raw message \u2014 it never matched \u201cat\u201d at all, so any \u201cTeam A at Team B\u201d phrasing (extremely common, and exactly how Drew typed it) skipped the cheap cached-reply path entirely and fell through to the full prediction pipeline every time, even for an already-logged matchup. Combined with the pre-existing, previously-documented fact that WNBA is unusually subrequest-hungry and runs close to Cloudflare's per-invocation ceiling, this made WNBA specifically prone to blowing the budget on a request that should have been a near-free cache hit. Fixed by widening the separator regex to also accept \u201cat\u201d as a separator, verified against real phrasings before shipping (\u201cSeattle Storm at New York Liberty\u201d now correctly resolves to the stored \u201cSeattle Storm vs New York Liberty\u201d Bet Label). Accepted a minor, low-cost tradeoff: a stray \u201cat\u201d in an unrelated sentence can now trigger one harmless extra Airtable lookup that simply finds no match and falls through normally. As a second, independent precaution, removed today's new line-movement odds fetch from the post-model backstop path specifically (findExistingPendingRecord) \u2014 that path only runs after the full expensive pipeline has already executed, the worst place to add one more external fetch; the fast path (now fixed) is the safe, cheap place for that feature and covers the vast majority of real cache hits. Also found and fixed a related pre-existing display bug while investigating this record directly: some older logged replies carry a legacy trailing \u201c---ODDSLINE---\u201d block containing a raw, unformatted odds JSON blob, which extractFullReplyFromNotes() was including verbatim in the text shown to a second user asking about the same matchup. Now stripped.
MISSION STATEMENT: 99% TARGET REPLACED WITH A REALISTIC FAVORITES-BASED BENCHMARK. Drew asked to remove a specific line tying the 99% accuracy target to \u201call the tools you're given\u201d \u2014 that reasoning sentence was cut first. He then asked to add real-world context: most favorites win 60-65% of the time, so consistently landing 70-75% is a genuine win. Once both were in, the paragraph still stated a flat \u201c99% of the time\u201d job target sitting right next to a line calling 70-75% \u201cgenuinely winning\u201d \u2014 an internal contradiction. Fixed by replacing \u201cat least 99% of the time\u201d with \u201cas often as you possibly can\u201d in the JOB line itself, so the 60-65%/70-75% favorites framing is now the one stated benchmark instead of competing with a separate 99% figure. No other part of the mission paragraph (verify-before-you-speak, calibration, decisiveness-over-hedging) was touched.
LINE MOVEMENT COMMENTARY ON CACHED (DUPLICATE-MATCHUP) REPLIES. Drew wanted a second user asking about an already-logged matchup (e.g. person B asking about Tigers vs Royals after person A already triggered a prediction) to get the identical cached message person A got, but with a comment on how the market line has moved since it was first logged, appended at the very end. Implementation is deterministic arithmetic, not a second LLM call \u2014 line movement is just \u201cprice then vs. price now,\u201d so re-invoking Gemini/Grok for it would only add latency and cost without adding judgment. Added parseOpeningLineFromText() to pull the originally-logged odds back out of the stored reply's own \u201c**Market Line (Bookmaker):**\u201d line (already persisted in Notes, no schema change needed), refactored fetchMarketOdds into a structured fetchMarketOddsData() plus a thin formatter (zero behavior change for the fresh/first-logged path), and added buildLineMovementComment() to phrase the comparison \u2014 moved toward the pick, moved away, or steady, with a \u201cnotably\u201d qualifier only on \u00b15+ point implied-probability swings. Wired into both places a cached reply can be served: the pre-model fast path (checkExistingPrediction) and the post-model backstop (findExistingPendingRecord). Fails open on both ends \u2014 no movement line is added if the original message never had odds logged (e.g. no ODDS_API_KEY at the time) or if a current line can't be found now; the cached message is served exactly as before in that case.
ESPN RESPONSE CACHING ADDED (Cache API, no new bindings needed). Motivation: Drew's ZenRows free-tier usage was tracking toward ~6,000 requests/month against a 5,000 cap after just 2 days at ~200/day \u2014 he's moving to the paid 45k tier, but caching still cuts real duplicate load rather than just paying it away. The clearest waste: gradePendingBets() re-fetches the same day's ESPN scoreboard once per pending bet, every 6-hour cron tick \u2014 five WNBA bets pending on the same date meant five identical scoreboard fetches per run for data that hadn't changed. Added getEspnCached()/putEspnCached() using Cloudflare's built-in caches.default (no Workers KV namespace or wrangler.toml change required \u2014 works immediately on deploy) and wired both fetchEspnWithRetry() and fetchEspnCoreWithRetry() to check the cache first and populate it on any successful response, whether that response came from the direct fetch or the ZenRows bypass. TTL varies by endpoint volatility via getEspnCacheTtlSeconds(): scoreboard URLs get a short 5-minute TTL (games can be live mid-fetch, so this can't be cached too long, but it still collapses the same-run duplicate-bet case above); athlete/roster/team-list/statistics URLs get a 1-hour TTL since a player's season stats or a roster don't meaningfully change within a day. Fails open by design \u2014 if caches.default is ever unavailable (e.g. certain local dev setups), getEspnCached() catches and returns null rather than breaking the real fetch, so this can only reduce ZenRows/ESPN load, never introduce a new failure mode. This reuses and finally implements the fetchEspnWithRetry cacheable option, which existed in the function signature already but was dead \u2014 never actually wired to anything.
CONFIDENCE CALIBRATION SWITCHED FROM GLOBAL PLATT SCALING TO PER-BAND CALIBRATION. Drew noticed the visible confidence number almost never moved from the raw simulation win% and correctly suspected the calibration wasn't really \u201clearning.\u201d Tested the actual live fitPlattScaling/applyPlattScaling against 391 real graded Gemini picks pulled straight from Airtable before touching anything: the 2-parameter global logistic fit (a=0.975, b=-0.301) was genuinely fitting, not null or broken \u2014 but SAM's raw confidence values cluster extremely tightly in the 51-65% range (a side effect of CONSERVATISM_SHRINKAGE keeping SAM cautious), and the fitted curve's zero-correction crossover point landed almost exactly on 56%, one of the single most common raw values. So most real predictions saw a 0-2 point shift or none at all, which reads as \u201cnot learning\u201d even though the math was working. Root problem: one global curve averages away band-specific miscalibration. Replaced with fitBandCalibration()/applyBandCalibration(): buckets graded picks into the same <50/50-59/60-69/70-79/80-89/90%+ bands already used for the CONFIDENCE CALIBRATION prompt block, computes each band's actual historical hit rate, and blends it with the raw stated confidence via Bayesian shrinkage (prior weight 10 \u2014 a band with only a handful of graded picks stays close to raw; a band with dozens converges toward its real hit rate). Verified against the same 391-pick dataset: this immediately surfaced that the 60-69% band (98 graded picks) was actually only hitting 52% \u2014 invisible under the old global fit because the well-calibrated 50-59% band (245 picks, 57.1% actual) dominated the aggregate average. High bands with only 1-2 graded picks (80-89%, 90%+) correctly stay close to raw rather than overfitting to noise. fetchPredictionHistory()'s returned shape renamed plattParams \u2192 calibrationParams end to end (both provider call sites in the /api/chat handler, the failover path, and the post-extractPick application block) for clarity; behavior at the call site is otherwise unchanged \u2014 still applied immediately after extractPick(), still overwrites pick.confidence before Airtable logging and the odds lookup, still rewrites the visible reply text so what the user reads matches what gets logged.
STANDALONE \u201cCONFIDENCE: NN%\u201d LINE REMOVED FROM VISIBLE REPLIES. Drew flagged that the labeled confidence callout was just repeating the same number as the win% split \u2014 true by design (rule 6 defines pick.confidence as \u201cyour final win percentage\u201d) \u2014 and asked for the redundant line gone. Removed the requirement from system prompt rule 3 and from buildCachedReplyText()'s cached-reply text for consistency. The win-percentage split still appears naturally in the narrative (e.g. \u201cgiving the Blue Jays a 55% win probability\u201d); only the separate labeled line is gone. Underlying confidence tracking, Airtable logging, and calibration are unaffected \u2014 this only changes what's displayed. USER-FACING SUMMARY (Drew's own words, worth keeping verbatim): \u201cFor a while it was not showing the results of the Monte Carlo simulation because the confidence % was the same as the winner %.\u201d The standalone line was never a separate, calibrated read on the pick \u2014 it was the same simulation output shown twice under two different labels. Two changes together actually fixed this: (1) this entry, removing the redundant line so only the natural win% narrative shows; and (2) the later switch from one global Platt-scaling curve to per-band calibration (see that entry below), which lets the number genuinely move now instead of landing back on the same value it started from.
ZENROWS ESPN-403 BYPASS EXTENDED TO EVERY REMAINING RAW ESPN FETCH; WNBA GRADING BUG FOUND AND FIXED. Antigravity's prior ZenRows integration only covered the main scoreboard call and UFC/MMA core-API lookups; roughly a dozen other ESPN call sites (fetchAthleteOverviewStats \u2014 the actual QB rating/PPG/ERA fetch used for every non-UFC sport, rosters, injuries, standings, team search, tennis, golf, MLB probable pitcher, and critically findFinalResult, the function gradePendingBets() calls to check whether a game is final) still used a raw fetch() with zero fallback. Root cause of \u201cWNBA not grading\u201d: findFinalResult had no ZenRows fallback, so a 403 on that specific scoreboard call silently left the bet Pending forever with no retry \u2014 now routed through fetchEspnWithRetry() like everything else. Also caught that fetchUFCFighterStats was already calling fetchEspnCoreWithRetry() but never passed env through, meaning its ZenRows fallback silently never fired either \u2014 fixed by threading env through the full call chain (espnAthleteSearch \u2192 espnAthleteSearchRaw, getUFCMatchupPowerRatings, getTennisMatchupPowerRatings \u2192 fetchTennisPlayerStats, runFieldSimulation \u2192 fetchGolfField/fetchGolfPlayerScoringAvg, getKeyPlayerFactor \u2192 fetchAthleteOverviewStats, and every other affected function). Also stubbed the previously-undefined fetchEspnViaBrowserRender() so tryEspnBypassFallbacks() can't throw a ReferenceError if a BROWSER binding is ever added without an actual implementation.
"PREDICTIONS REMAINING" ON PAGE LOAD WAS ALWAYS ONE LOW — REAL BUG, DISPLAY-ONLY. Drew noticed the badge read "2 Predictions Remaining" immediately on signing in, before making any prediction at all. Root cause: checkPredictionAllowance() is written to answer "if a prediction happens right now, how many are left after it" — correct for its real call sites (right before actually processing a prediction, in /api/chat), but the page-load render was calling that exact same function purely to DISPLAY the current count, so it was always showing the number one prediction already spent, whether or not that was true. Confirmed this was purely cosmetic before touching anything: the page-load call site only ever read .predictionsLeft off the result and discarded everything else, including the deduction object and (for guests) the guestCookie — so nothing was actually being written back to Supabase or set as a cookie on page load; no real prediction was ever silently consumed by this, only the number shown was wrong. Fixed by adding a peek option to both checkPredictionAllowance() and checkGuestPredictionAllowance(): peek:true reports the true current count with no "as if consuming one" adjustment and returns no deduction/cookie at all, while every real prediction-processing call site (three of them, all unchanged) keeps the exact same default behavior as before. The one page-load call site now passes peek:true. Verified the corrected arithmetic in isolation across four cases (fresh account showing 3, mid-usage showing the right number after real predictions, and the paid-credits phase after the daily free ones are gone) before shipping.
MISSION STATEMENT: EXPLICIT GOAL SENTENCE ADDED. Drew asked whether the goal of "pick the right winner consistently" was stated plainly anywhere — it wasn't; the existing mission paragraph only had the calibration framing ("be calibrated, not just confident"), which is related but not the same claim. Talked through the risk before adding it as-is: a model told its goal is raw win-rate has an easy way to look successful at that without being honest (inflate confidence, avoid flagging genuinely close calls) that a calibration-based goal doesn't allow, since a Brier score punishes over- and under-confidence equally and can't be gamed by sounding more sure. Landed on combining both in one sentence rather than picking one: "The goal is simple to state and hard to fake: be right, and be right exactly as often as you say you will be — the second half of that isn't a softer, lesser version of the first, it's what keeps the first one honest, since a stated confidence you can't back up with a real hit rate isn't actually being right, it's just sounding right." Placed right after the existing "reason for existing" sentence, before the three-point breakdown (verify/calibrate/show work), which is unchanged.
EXPLICIT "Confidence: NN%" LINE REQUIRED IN EVERY PREDICTION. The win-percentage split was already required in the narrative (rule 3, unchanged), but only woven into prose (e.g. "giving the Blue Jays a 55% win probability") — there was no guaranteed standalone, unmistakable callout of the number. Added a requirement that every prediction end with a plainly labeled "Confidence: NN%" line, same final number as the rest of the reply and the logged pick, not a second/different figure. No code change needed to keep this consistent with Platt scaling: the existing text-rewrite that corrects the raw confidence number to the calibrated one already does an exact string replace of "{number}%" everywhere it appears in the reply, so this new labeled line gets corrected right along with the rest of the text automatically.
STATS PAGE NOW REQUIRES SIGN-IN + MORE VISIBLE SIGNUP PROMPT. (1) /stats now checks the same session used everywhere else on the site (already computed once per request, reused here rather than recomputed) — no session redirects to /login with a return path back to /stats, same pattern /buy already used. (2) A "Don't have an account? Create one now!" line was added directly under the chat input on the main page, shown only when signed out (disappears entirely once session exists — confirmed via two separate rendered-HTML checks, one per state). The only prior way to find sign-in was a small all-caps nav link at the very top of the page, easy to miss — this puts it right where people are actually looking, in normal-weight readable text with an underlined link, not tiny mono nav styling. Confirmed no layout regression on the fixed-height chat page: input bar, send button, and the existing footer credit line all still render in their normal positions with the new line in between.
RESPONSIBLE-USE DISCLAIMER GATE RESTORED. Also built in the same separate conversation as Recent Picks and also apparently never actually deployed, for the same reason. This is NOT the old "Before you enter" click-through gate (that removal was real and intentional, from a much earlier version, unrelated) — it's a newer, simpler one-time overlay: exact text as originally specified ("This site is not gambling advice. It is a simulation of specific sporting events in the most realistic way possible. Please use responsibly." plus a line about thumbs up/down feedback), a single "I Understand — Continue" button, no email or data collection, dismissed permanently per browser via a localStorage flag (sam_disclaimer_ack_v1) — separate from and layered on top of the real sign-in gate, which still controls actual access regardless of this dismissal. Verified with an actual functional test this time, not just a syntax check: rendered the real page through Node (evaluating the outer template literal properly, not a naive text extraction, after that gave a false-positive syntax error earlier tonight) and confirmed the gate shows on first visit, dismisses on click, and correctly stays dismissed after a reload.
RECENT PICKS: PREDICTED WINNER LABEL MADE EXPLICIT. The Picked Winner field was already being pulled correctly from Airtable and displayed in each Recent Picks row — verified the field name my code requests ("Picked Winner") matches the real Airtable field exactly, confirmed against the actual table schema. It was just labeled "Pick: {team}", which reads ambiguously (a pick the user made vs. SAM's prediction). Relabeled to "Predicted Winner: {team}" so there's no ambiguity.
RECENT PICKS RESTORED ON /stats + PREMIUM BUTTON FINISHED. Two things closed out here. (1) The "Recent Picks" section on /stats — last 10 bets from Airtable sorted by Placed Date, showing Bet Label, Picked Winner, Confidence %, and a color-coded Result badge (green win / red loss / amber push / blue pending) — had gone missing from the live worker. It was built in a separate conversation and, best evidence available, was never actually deployed before a later "pull exactly what's live" step in this thread established the deployed code as the working base going forward, which didn't have it. Re-added fetchRecentPicks() next to fetchPublicStats(), wired into the /stats handler, and re-added the section between "By Sport" and "Visual Breakdown" with matching CSS — confirmed against a fresh pull of the live worker plus the original pre-edit file to make sure nothing else was actually missing (Platt scaling, simulation score averages, premium_until pricing, footer branding, and calibration/Brier all confirmed present and untouched throughout). (2) The Premium buy button's gold+bold "PREMIUM VIP UNLIMITED" relabel, requested right before this, hadn't made it to production yet either — finished here on the same confirmed-live base: gold (#E0B84C) and bold (800 weight) on both the stars and the label text, in the buy menu and the in-chat paywall button.
PRICING RESTRUCTURED — 10-PACK PRICE CHANGE + NEW EXPIRING PREMIUM TIER. CREDIT_TIERS is now just two options: 10 Predictions for $10 (was $5), and a new Premium tier — unlimited predictions for 30 days, $25. The old 100-for-$50 and lifetime-for-$200 tiers are removed from sale; anyone who already bought lifetime access keeps it (the lifetime_access check in checkPredictionAllowance was left untouched, purely grandfathered, just no longer purchasable). PREMIUM IS TIME-LIMITED, NOT PERMANENT, WHICH IS A NEW CONCEPT: added a premium_until date column directly to the existing sam_bettors Supabase table (chose to extend the table already used for sign-in/billing rather than a separate premium_subscribers table that turned out to exist in the same project but isn't the one in use here) — checkPredictionAllowance now grants unlimited (999) predictions whenever premium_until is today or later, same shape as the lifetime_access check right above it. The Stripe webhook sets premium_until to 30 days out on a successful Premium purchase, extending from the existing premium_until if it's still active rather than resetting from today, so an early renewal doesn't cost the customer days. Since Drew separately created a Payment Link directly in the Stripe dashboard for this tier (buy.stripe.com/...), which won't carry the metadata.tier field our own dynamically-created Checkout Sessions do, the webhook now falls back to matching CREDIT_TIERS by exact amount_total when metadata.tier is absent — so a $25 charge from either path resolves to the same "premium" tier and correctly credits whichever bettor's client_reference_id is attached. The in-app buy button itself still points at our own /buy?tier=premium dynamic-checkout route rather than the raw Payment Link, since that path already guarantees correct metadata and client_reference_id with no matching logic needed — the fallback exists specifically for the separate dashboard-created link. UI: the Premium option in both the buy menu and the in-chat paywall now uses the same gold styling the old Lifetime tier had (border/price color — #D4AF37/#E0B84C), with a gold star on each side of the label. Paywall copy updated from "credits never expire" to distinguish the two tiers correctly, since one of them now does expire. Blog access and ufc.grimaldi.tv access, the other two things bundled into this $25 tier, live entirely outside this worker — nothing here can gate access on those properties; this change only covers what SAM itself controls (unlimited predictions).
FOOTER BRANDING UPDATED, CONSISTENT ACROSS EVERY PAGE. The old "Built for The Drew Grimaldi Podcast — comedy, politics, and entertainment live every Saturday at 11AM ET on grimaldi.tv." line (previously only on the /about and /stats pages) is replaced everywhere with a single consistent line: "Proprietary Technology built by GRIMALDI.TV". Applied to /about, /stats, /workflow, /changelog (the latter two had no footer at all before this), the login and signup pages, and the main chat app itself. The main app required care since it's a fixed 100vh flex-column layout (not a scrolling page like the others) — confirmed via a static Playwright render at desktop and mobile viewport sizes that the new line sits cleanly below the existing footer-links row without overlapping the input bar or send button, and that the scrollable message area simply absorbs the ~20px via its existing flex:1 sizing.
SIMULATION SCORE AVERAGES NOW SURFACED IN THE PREDICTION, AND MADE GENUINELY REAL IN THE PROCESS. Two parts to this: (1) run_matchup_simulation's tool result already carried an "avg simulated score" per side, but nothing told the model to actually state it in the reply — it sat in the tool result unused. Added an explicit instruction (system rule 2) so any team sport with a real score line (NFL, NBA, WNBA, MLB, NHL, Soccer, College) now states the average simulated score for each side alongside the win percentage (e.g. "projected score: Lakers 112.4 - Celtics 108.7"). Doesn't apply to UFC, NASCAR, Bowling, Darts, Tennis/WTA/ITF, or Horse Racing, none of which produce a team score. (2) WHILE WIRING THIS UP, FOUND THE UNDERLYING NUMBER WASN'T ACTUALLY SIMULATED: runTeamSportSimulation() was reporting the pre-simulation lambda input (the Poisson/normal distribution's rate parameter, known before a single trial ever ran) back out as "avg simulated score" — mathematically close in expectation, but not an actual product of the 25,000 trials, which runs against this whole project's standing rule of reporting only what's really computed. Fixed by accumulating the real per-trial scoreA/scoreB values during the existing simulation loop and computing the true empirical average afterward. Confirmed this isn't just a technicality before shipping: for bivariate Poisson sports (MLB/NHL/Soccer) the two numbers converge closely at 25,000 trials as expected, but for the normal-distribution sports (NFL/NBA/WNBA/College) negative-score trials get clamped to 0 before the average, which measurably pulls the true average away from the raw lambda in low-scoring, high-variance matchups (tested case: lambda 8.0 in, true empirical average 9.8 after clamping) — exactly the kind of case this fix now reports correctly. expectedScoreA/expectedScoreB keep the same field names and same downstream usage, only the computation changed.
PLATT SCALING IMPLEMENTED — CONFIDENCE IS NOW ACTUALLY RECALCULATED IN CODE, NOT JUST NUDGED BY A PROMPT INSTRUCTION. This is the real version of the calibration idea discussed across several earlier patches: instead of just telling the model "be more conservative" and trusting it to comply, this mathematically transforms the model's raw stated confidence into a corrected number before it's ever shown to a user or logged, using real historical (confidence, outcome) pairs already sitting in Airtable. IMPLEMENTATION: fitPlattScaling(records) fits a 2-parameter logistic regression (calibrated_prob = sigmoid(a × raw_confidence + b)) via plain gradient descent (1500 iterations, no external ML library — this is a Cloudflare Worker, no numpy/sklearn available) directly on the same records array fetchPredictionHistory() already fetches for the TRACK RECORD block, so this required zero new Airtable calls. Requires 20+ graded picks with a logged confidence to fit at all; returns null below that, which cleanly no-ops everything downstream (raw confidence passes through unchanged — there was real historical data to correct against yet). fetchPredictionHistory()'s return signature changed from a bare string to { text, plattParams } to carry the fitted parameters out alongside the existing prompt text; both call sites in the /api/chat handler (the primary provider call and the cross-provider failover path) updated to destructure the new shape and track a single finalPlattParams that follows whichever provider's call actually succeeded. APPLIED IMMEDIATELY AFTER extractPick(): the model's raw stated confidence (parsed from the hidden [PICK] tag) is run through applyPlattScaling() and the result overwrites pick.confidence before anything downstream sees it — Airtable logging, the odds-lookup step, everything. Also rewrites the VISIBLE reply text (an exact string replace of "{rawConfidence}%" with "{calibratedConfidence}%") so what the user actually reads matches what gets logged — without this, a user could read "91%" in the chat while Airtable silently logged "61%", which would be a real, confusing inconsistency. Sanity-checked in isolation before touching the live file: a systematically overconfident synthetic history (records saying "92%" but only winning 60% of the time) correctly pulled a fresh 91% claim down to 61%, and the visible text and logged value matched exactly after the rewrite. TWO HONEST TRADEOFFS, DOCUMENTED RATHER THAN HIDDEN: (1) The text-replace is an exact string match on "{number}%", not a semantic understanding of the reply — in the rare case another stat in the same reply happens to share the exact same percentage integer (e.g. a shooting percentage that happens to equal the confidence number), that unrelated mention would also get rewritten. Considered a targeted regex or position-based replace instead, but a blind match on the final PICK confidence is the same number the model already tries to state prominently and consistently, so the actual collision risk in practice is low; flagging it as a known limitation rather than pretending it's impossible. (2) Confidence % in Airtable now stores the CALIBRATED number going forward, not the model's raw output — there's no separate "raw confidence" field preserved. This means future re-fits are calibrating on top of already-corrected history rather than pure raw model tendency. In practice this is self-stabilizing (once well-calibrated, a, b converge toward the identity transform and stop correcting further) rather than harmful, but it does mean if the underlying model's raw tendency drifts later (e.g. a prompt or model change), that fresh drift will take longer to show up clearly in the numbers than it would against an untouched raw signal. Didn't add a second Airtable field to preserve the raw value, since that's a live schema change and Drew's explicit preference this session was reusing existing data over adding new infrastructure — noted here so it's an easy thing to revisit later if drift-masking ever becomes a real practical problem rather than a theoretical one. NOT CHANGED: teamAWinPct/teamBWinPct (the raw per-team simulation split reported in the [PICK] tag) are left untouched — those represent the simulation's raw output by design, not the "final calibrated" number, so only the single main confidence field gets the Platt transform.
BRIER SCORE MADE ACTIONABLE, NOT JUST REPORTED. Drew asked about implementing an auto-weighting "reward" system using the Brier score (routing between Gemini/Grok based on which is better-calibrated per sport) — flagged before building it that true real-time per-sport routing would mean calling both models on every request to compare them live, which directly reopens a cost problem already solved once before (see the earlier changelog note: automatic background second-opinion calls were removed because they silently doubled every request's API cost). Building real auto-routing would also need a new small Airtable table to cache a periodically-computed "current best provider" value cheaply (recomputing from ~1000 records on every page load isn't viable), which is a live-schema change bigger than anything else done this session. Drew's follow-up cut through this cleanly: instead of routing between models, just have the Brier score directly shape the confidence number on the CURRENT pick — which needs zero new infrastructure, since the score is already computed from data already being fetched. Implemented as a tiered instruction attached directly to the existing Brier score line inside computeConfidenceCalibration(): Brier ≥ 0.25 (at or worse than pure coin-flip guessing) → explicit instruction to pull confidence toward the middle (55-65%) on new picks rather than trusting the simulation's raw output; Brier 0.20-0.25 (better than coin-flip but loose) → lean somewhat more conservative, especially above 80%; Brier < 0.20 (genuinely good) → explicit instruction to keep calibrating the same way rather than second-guessing a working approach. This only fires once there are 8+ graded picks with a logged confidence (same threshold used elsewhere in this block), computed per-provider like everything else here. Sanity-checked all three tiers against representative Brier values (0.10, 0.19, 0.22, 0.25, 0.41) before shipping — each mapped to the intended guidance tier correctly. No new tables, no new endpoints, no change to model-selection/routing logic — purely a smarter instruction layered onto data already flowing into the prompt.
BRIER SCORE ADDED TO CONFIDENCE CALIBRATION. Drew asked whether SAM used a Brier score (the standard proper scoring rule for probabilistic forecasts: mean of (stated_confidence - actual_outcome)² across all graded picks, 0 = perfect, 0.25 = pure coin-flip guessing) — it didn't, so it was added. Computed inside the existing computeConfidenceCalibration() using the exact same records loop already reading Confidence % and Result per pick, so no new Airtable calls or storage. Unlike the bucket breakdown (which shows WHERE SAM is over/underconfident), the Brier score is a single overall summary number, and being a proper scoring rule it can't be gamed by a model shading its stated confidence away from its true belief — a model that always says 50% looks "calibrated" in a trivial sense but gets punished by Brier score for being unhelpful/unsharp, same as an overconfident model gets punished for being wrong. Sanity-checked against known reference cases before shipping: a perfect forecaster scores 0.0, pure coin-flip guessing scores exactly 0.25, an overconfident model (says 90%, only wins 50%) scores 0.41 (worse than coin-flip, correctly penalized), and a well-calibrated 70% forecaster (wins 7 of 10) scores 0.21 (better than coin-flip, correctly rewarded) — all matched expected values exactly. Reported alongside the existing bucket lines in the same hidden CONFIDENCE CALIBRATION block (only surfaced once there are 8+ graded picks with a logged confidence, same small-sample threshold used elsewhere in this block), computed per-provider like everything else in this context (Gemini and Grok get their own separate Brier score from their own separate bases). This is a diagnostic/summary number only — no new instruction was added telling SAM to directly act on it beyond what the existing calibration instruction already covers, since the bucket breakdown is still the more actionable, specific signal for adjusting behavior band-by-band.
MISSION/PROBLEM STATEMENT ADDED TO SYSTEM PROMPT. Drew's observation: the prompt jumped straight from the one-line identity statement into a wall of procedural rules (never reference odds, verify before predicting, etc.) without ever telling Gemini/Grok WHY those rules exist — no articulation of the actual problem SAM is solving. Confirmed this was true by re-reading the prompt's opening. Added a new paragraph immediately after the identity line, before rule 1: names the real problem (most sports takes are either uninformed gut calls or shaped by sportsbook financial incentives, neither trustworthy), states SAM's actual answer to it (real computed simulation on independently-verified live data, never odds, with an honest self-correcting track record instead of unearned confidence), and gives an explicit fallback heuristic for situations the numbered rules don't cleanly cover ("would a real analyst with real data, real accountability, and no incentive to mislead say this?"). Distilled into three concrete priorities referenced back to existing mechanisms already in the prompt: verify-before-speaking (the existing hard verification gate), calibration over raw confidence (the existing TRACK RECORD/CONFIDENCE CALIBRATION context), and showing real work (naming actual players/stats so reasoning is checkable). This doesn't change any procedural rule or add new behavior \u2014 it's meant to give the model something to generalize from when a specific rule doesn't obviously apply, rather than only ever pattern-matching against an enumerated list.
STADIUM INTELLIGENCE EXTENDED TO TENNIS/WTA/ITF — A REAL SCHEDULE-MATCHING BUGFIX, A GEOCODING BUGFIX, AND PHYSICS VERIFIED AGAINST REAL SOURCES BEFORE SHIPPING, NOT ASSUMED. Full record of everything done in this session, per Drew's request for maximum detail. (1) THE BUG THAT WAS ACTUALLY THERE: Drew asked to bring Stadium Intelligence weather to tennis too. Before writing any code, the ESPN tennis API shape was fetched live and compared against the team-sport shape already in use. Team sports (confirmed live via MLB, NFL, NBA scoreboard fetches) return a flat event.competitions[0] object with venue = {fullName, address: {city, state, country}, indoor: true/false}. Tennis (confirmed live via the ATP scoreboard) returns something structurally different: the top-level event IS the tournament (e.g. "Mifel Tennis Open by Telcel Oppo"), individual matches live nested three levels down at event.groupings[].competitions[], and each nested competition has its own venue = {fullName: "City, Country", court: "Court name"} — no address sub-object, no indoor boolean, nothing matching the team-sport shape at all. Cross-checking this against getGameSchedule's actual matching code (competitorNameHits, and the events.filter() block) confirmed the real, pre-existing consequence: since that code only ever reads e.competitions (singular, flat), and tennis events have zero entries there, get_game_schedule could NEVER match an individual tennis player pairing — only the tournament name itself, which never contains a player's name. This directly contradicts a claim made in an earlier changelog entry that tennis schedule lookup had "no known gap" — it did, silently, the whole time, and this patch is the actual fix, not a cosmetic add-on. This also matches and explains why rule 8 had already been written to make Tennis/WTA/ITF skip the schedule-verification gate entirely and rely on ESPN player-search verification instead (rule 2) — that carve-out was a correct workaround for a real underlying limitation, not an arbitrary exception. (2) THE FIX: new flattenScheduleEvents(rawEvents), called immediately after the ESPN fetch inside getGameSchedule, before any matching logic runs. For events with a real top-level competitions[] array (team sports, UFC, NASCAR, golf-shaped events), it passes them through completely unchanged — zero risk to any already-working sport. For events shaped like tennis (empty/missing top-level competitions but populated groupings[]), it walks every grouping (Men's Singles, Women's Singles, etc.) and every nested competition inside it, pulls both competitors' names from competitor.athlete.displayName (falling back to fullName), builds a synthetic pseudo-event per match with name/shortName set to "Player A vs Player B" for display, preserves the real tournament name separately as a new tournamentName field (needed later for the indoor-tournament check), and copies the nested competition's own date/status/venue up so the existing per-match code downstream (date formatting, ISO_DATE tagging, score display, and now Stadium Intelligence) works completely unmodified on the flattened result. This was the single highest-leverage fix in the session: it repairs tennis schedule lookup generally, not just for weather purposes. (3) WEATHER GEOCODING FOR TENNIS, INCLUDING A REAL BUG CAUGHT DURING TESTING: tennis venue.fullName is a flat "City, Country" string (e.g. "Los Cabos, Mexico", "Washington, USA") rather than separate city/state fields, so a new geocodeCityCountry(city, country) was added alongside the existing (unmodified) geocodeCity(city, state) used for team sports. First implementation queried Open-Meteo's free geocoding API by city name alone and took the first/most populous result — live-tested against "Los Cabos" before trusting it, and the very first result returned was a small village in Asturias, Spain (population data absent, elevation 89m, timezone Europe/Madrid) ranked ahead of the real Mexican resort city in Baja California Sur where the actual ATP tournament is played. This is exactly the kind of silent-wrong-answer failure mode the rest of this codebase has repeatedly tried to eliminate (see prior CANNOT_VERIFY/hard-verification-gate changelog entries), so geocodeCityCountry() now filters all candidate results to ones whose country field actually matches (case-insensitive substring both directions) the country ESPN provided, sorts the survivors by population, and returns null (triggering an honest "weather lookup failed" message) rather than falling back to an unrelated country's same-named town if nothing matches. Re-tested afterward with the country filter in place and confirmed it correctly resolves to Mexico. (4) INDOOR DETECTION FOR TENNIS, DOCUMENTED AS A KNOWN LIMITATION: ESPN's tennis API exposes no indoor/outdoor flag at all (confirmed absent in the same live API dump used for item 1), unlike team sports where venue.indoor is reliably present. Rather than guessing or defaulting silently, added a short, explicitly best-effort KNOWN_INDOOR_TENNIS_EVENTS array matched against the tournament name: ATP Finals, WTA Finals, Next Gen ATP Finals, Paris Masters/Rolex Paris Masters, the European Open (Antwerp), Erste Bank Open (Vienna), Swiss Indoors (Basel), Stockholm Open, and the Dallas Open. Any indoor tournament not on this list will incorrectly get treated as outdoor and attempt a real (if irrelevant) weather lookup — this is a known, named gap the same way Bowling/Darts' missing schedule source is documented elsewhere in this changelog, not something papered over. (5) PHYSICS — CHECKED AGAINST REAL SOURCES RATHER THAN ASSUMED, IN BOTH DIRECTIONS: Drew raised two separate physical claims in this session; both were checked against real reporting/research before writing a single word of system-prompt guidance, because getting either one backwards would have made every future hot/humid prediction worse, not better, in a way neither Drew nor the model would have any way to notice afterward. • TENNIS + HEAT (confirmed correct, refined): multiple sources (physics-focused tennis blogs, sports-science writeups) agree hot air is measurably less dense than cool air — one source's own numbers: going from 10°C to 38°C drops air density roughly 10%, adding an estimated 3-5 km/h of ball speed; separately reported that temperatures above roughly 32°C also harden clay specifically, raising bounce and further speeding play. Net effect: heat makes the ball fly faster and bounce higher, favoring aggressive/power hitters over grinders — the opposite direction from Drew's original framing ("hot and humid... ball moves slower"), so this half of the original claim was corrected rather than encoded as stated. • TENNIS + HUMIDITY (Drew's instinct was directionally reasonable but via the wrong mechanism, now modeled correctly): humidity's effect on air density is real but tiny and actually points the same direction as heat (water vapor, molecular weight 18, is lighter than the nitrogen/oxygen it displaces, molecular weight 28/32, so humid air is very slightly less dense — multiple physics-focused sources independently confirm this counterintuitive point and describe the pure air-density effect on ball speed as "almost imperceptible" on its own). The real, well-documented, and separate mechanism behind players and commentators calling humid conditions "heavy" is that the ball's felt absorbs moisture and physically "fluffs up" over the course of a match, adding real mass and drag — directly reported by pros at the 2026 US Open ("You'll feel it get heavier, which means it doesn't move as fast... a little slower because it's getting fluffier"). This is now modeled as its own distinct advisory (fires at ≥70% average humidity) rather than being merged into or treated as canceling out the heat effect, since they are physically different mechanisms operating on different parts of the system (surrounding air vs. the ball's own material). • TENNIS + INDOOR COURT SPEED (Drew's claim, from personal college-tennis experience, checked and confirmed): cross-referenced against tennis-surface writeups, coaching sites, and direct player quotes. Consistent majority finding: removing wind and sun lets players hit with more confidence and commit fully to lines, and indoor hard courts are frequently set up faster besides — one widely-quoted characterization of 1990s-2000s indoor events as "crazy fast" surfaces where "big servers dominated indoors." One dissenting professional viewpoint was also surfaced during research (a top server arguing arena conditions actually help his opponents more, by removing the wind-disruption edge aggressive hitters get outdoors) — noted here for completeness rather than silently discarded, though it's the minority view. Net: Drew's claim is well-supported and is now surfaced on every indoor Tennis/WTA/ITF result, worded as "generally favoring" rather than an absolute, matching the real state of the evidence. • MLB + HUMIDITY (Drew's instinct, checked and found backwards, correctly NOT implemented): the same water-vapor-is-lighter-than-N2/O2 mechanism applies to baseball. Multiple independent sources — an AccuWeather meteorologist quoted in sports reporting, a University of Arizona engineering professor, and a peer-reviewed physics paper specifically studying MLB's humidor — agree humid air is very slightly less dense, so a batted ball carries marginally FARTHER in humidity, not less. The peer-reviewed paper's own numbers: raising a stored ball's relative humidity from 30% to 50% increases fly-ball distance by about 2 feet from aerodynamics alone — the opposite direction from the offense drop MLB's humidors are actually used to produce. The real mechanism behind humidors suppressing home runs is separate: pre-conditioning the ball in humidified storage for days before use increases its mass and reduces its coefficient of restitution (how lively it is off the bat), which is a bat-contact effect, not a through-the-air drag effect — and it's now a constant, standardized equipment policy at all 30 MLB parks since 2022, not something that varies with any single day's forecast the way wind or rain does. Historical evidence cited in sourcing: Coors Field pitchers ran a 6.50 ERA from 1995-2001 (over 2 runs worse than the rest of the league) before a humidor was introduced there in 2002, after which ERA at altitude dropped roughly a full run — a real, large, but storage-based effect, unrelated to game-day weather. Conclusion: no humidity advisory was added to MLB's Stadium Intelligence output, since the ambient-weather effect is both backwards from Drew's framing and, per the same research, too small (a couple of feet) to be worth surfacing at all — correctly doing nothing here was the right call, not an oversight. (6) CODE CHANGES SUMMARY: flattenScheduleEvents() added and wired into getGameSchedule immediately after the ESPN fetch. getStadiumIntelligence() signature extended to (venue, isoDate, tournamentName, league); now branches on venue shape (team-sport address-based vs. tennis flat-string) rather than assuming one shape. New geocodeCityCountry() added alongside the untouched geocodeCity(). New KNOWN_INDOOR_TENNIS_EVENTS list + isKnownIndoorTennisEvent() helper. Open-Meteo hourly query extended to also request relative_humidity_2m (previously temperature/precipitation/windspeed/winddirection only). Advisory logic restructured from a single if/else-if into an accumulating array so multiple genuinely-independent advisories (e.g. rain AND humidity) can appear together instead of one silently overriding another. New sport-aware thresholds: heat advisory for Tennis/WTA/ITF only, fires at ≥85°F average; humidity advisory for Tennis/WTA/ITF only, fires at ≥70% average humidity; existing wind (≥15mph) and precipitation (≥50%) advisories now use tennis-specific wording (toss/serve/shot-control) when the league is Tennis/WTA/ITF, generic wording (ball flight/passing/kicking) otherwise. Rule 8 in the system prompt updated to describe all of the above to the model, including the indoor-tennis exception (mention the fast-court note even though there's no forecast) and explicit instruction not to conflate the heat and humidity effects or treat them as canceling out by default. (7) LIVE TESTING BEFORE SHIPPING: fetched a real upcoming ATP match (Gea vs Wong, Mifel Tennis Open, Los Cabos) end-to-end through the new code path — correctly flattened out of the tournament's groupings, correctly geocoded to Los Cabos, Mexico (not Spain), and returned a real forecast: 70°F, 77% humidity, 3mph wind SW, 89% precipitation chance — correctly triggering both the rain advisory and the humidity/ball-fluffing advisory together (confirming the accumulating-advisory-array change works as intended). Separately tested a known-indoor case (Turin, standing in for the ATP Finals) to confirm the indoor branch correctly fires the fast-court note with no forecast data attached, and that the KNOWN_INDOOR_TENNIS_EVENTS name-matching works. Both team-sport paths (MLB/NFL/NBA venue shape) were re-confirmed unaffected by these changes since flattenScheduleEvents() passes their already-flat event shape through untouched.
STADIUM INTELLIGENCE (WEATHER) — HEADLINE FEATURE OF THIS RELEASE. SAM previously had zero weather awareness — no wind, temperature, or precipitation data fed into any matchup, which matters for outdoor sports (wind at MLB parks affecting fly balls, wind/rain at NFL and outdoor events affecting passing/kicking/footing). Branded as "Stadium Intelligence" per Drew's direction, alongside the existing tennis surface-adjustment logic (getTennisSurfaceAdjustment), which is conceptually the other half of the same idea — course/venue-specific conditions affecting the matchup. Implementation: getGameSchedule's per-match ESPN response already includes a venue object with fullName, address {city, state}, and (confirmed via live testing) an "indoor" boolean — so no new tool call or hardcoded stadium list was needed. Added getStadiumIntelligence(venue, isoDate), called automatically for every non-completed match returned by getGameSchedule: indoor venues short-circuit immediately ("no weather impact"), outdoor venues get geocoded via Open-Meteo's free geocoding API (results cached in-memory per city/state for the life of the isolate) and checked against Open-Meteo's free 16-day hourly forecast (no API key required for either endpoint). Reports average temp, max wind speed + compass direction, and max precipitation probability across the game-time window (1pm-10pm local, falling back to full-day if the game's exact hours aren't in that band), with a plain-language advisory appended when wind ≥15mph or precip chance ≥50%. Games more than 15 days out get an explicit "forecast unavailable yet" line instead of silently omitting weather. Live-tested against Wrigley Field (outdoor, correct real forecast), Videotron Centre (indoor, correctly skipped), and Lambeau Field 3 days out (correct real forecast) before deploying. Failure modes (geocode miss, forecast fetch failure) degrade to a short explanatory line rather than throwing, consistent with the rest of getGameSchedule's error handling. FOLLOW-UP (same release): Stadium Intelligence data was reaching the model but staying backstage — useful for its own reasoning, invisible to the actual user reading the prediction. Per Drew's call ("bring up the weather in the final prediction too, that can impress some people"), rule 8 now explicitly tells SAM to work real forecast numbers into the visible reply when they're plausibly relevant to the matchup (wind for MLB/NFL/Soccer/Golf, meaningful rain chance for footing/ball security), naming the actual figure rather than a vague "windy" — and to say nothing at all when the venue is indoor or conditions are unremarkable, so this doesn't turn into clutter on every single pick. Both get_game_schedule tool descriptions (Gemini and Grok schemas) updated to flag that outdoor-venue results now carry this data, so the model knows to look for it.
RECENT FORM (STREAK) CONTEXT ADDED, DELIBERATELY INFORMATIONAL-ONLY (post-release patch, version held at 1.5.7 per Drew's call). Drew asked about rewarding SAM for win streaks; discussed it first rather than building it blind, since a naive "streak = raise confidence" mechanic would just be the hot-hand fallacy baked into the prompt — one game's outcome doesn't make the next one more likely to hit, and it would directly fight the calibration feature added in the previous patch. Landed on a strictly informational version instead: added computeRecentForm(), which reuses the same graded-records array fetchPredictionHistory already pulls (no new Airtable calls) and reports, per league, the last 5 results as a W/L sequence plus the current same-result streak length (e.g. "NBA: last 5 = W-W-L-W-W (currently 1W in a row)"). This is appended as a new "RECENT FORM" section in the same hidden TRACK RECORD context, explicitly labeled informational-only with an instruction never to adjust confidence based on streak length since each event is statistically independent. Leagues with zero graded picks are omitted. No confidence math, calibration bands, or logging behavior were touched — this sits alongside the existing calibration block, it doesn't feed into it.
CONFIDENCE CALIBRATION ADDED TO TRACK RECORD CONTEXT (post-release patch, version held at 1.5.7 per Drew's call). Previously fetchPredictionHistory's TRACK RECORD block only showed raw win/loss counts (overall and by league) plus the 30 most recent graded picks with their stated confidence — useful, but it never actually told the model whether its stated confidence numbers have been trustworthy. Added computeConfidenceCalibration(), which reuses the exact same graded-records array already being fetched (no new Airtable calls, no new storage) and buckets every pick by its logged Confidence % into 50-59% / 60-69% / 70-79% / 80-89% / 90%+ bands, computing the ACTUAL win rate within each band. This is appended as a new "CONFIDENCE CALIBRATION" section in the same hidden TRACK RECORD context injected into every prompt, computed separately per provider since Gemini and Grok log to separate Airtable bases and build separate track records. If a band is running well below its stated number — e.g. the model has been saying 90%+ but that band is actually only hitting 65% — the injected instruction tells it to state a more conservative number next time it lands in that range, rather than repeating a high confidence just because the analysis feels strong. Bands with fewer than 8 graded picks are flagged "small sample, weigh lightly" so early, thin data in newer sports (WNBA, Darts, Bowling, etc.) doesn't overcorrect calibration off a handful of results; bands with zero graded picks are omitted entirely rather than shown as 0/0.
UFC NO LONGER GATED BY get_game_schedule (the actual fix). Drew's AI-thoughts trace showed the real failure clearly: run_matchup_simulation succeeded perfectly (Oban Elliott 31% / Michael Oliveira 69%, real ESPN tale-of-the-tape stats — confirming both are real, currently-competing fighters), but the model THEN ran a redundant get_game_schedule check per rule 8, that check didn't surface the fight, and the model declined based on it — throwing away an already-successful, fully-verified simulation. Two problems: (1) ordering was backwards — a real simulation that pulled live fighter stats is a STRONGER real-existence proof than the schedule board, since a sim literally can't run for fake fighters; (2) ESPN's UFC scoreboard often lists only a card's headline/main-card bouts and omits prelims and early prelims, so a completely real scheduled fight (this was a prelim bout) can be missing from the board entirely — making the schedule check a false-negative machine for undercard UFC fights specifically. FIX (system prompt rule 8): UFC now skips the schedule-verification gate as a required step, exactly like Tennis/WTA/ITF/Bowling/Darts already do, and relies on run_matchup_simulation's fighter-stats verification instead. If the sim returns a real result, that IS verification — proceed to the [PICK], and a get_game_schedule miss can no longer veto it. UFC only declines when run_matchup_simulation itself returns CANNOT_VERIFY. get_game_schedule may still be called for UFC to fetch the event date, but it's now a date lookup, not a veto gate. Also removed UFC from rule 8's "required schedule coverage" league list so the prompt isn't self-contradictory.
FUZZY FIGHTER-NAME FALLBACK (partial fix). Drew's AI-thoughts trace revealed the real cause of the last round of UFC declines: the MODEL itself passed misspelled fighter names to the tool — "Odan Elliott" (for Oban) and "Michal Oliveira" (for Michael). ESPN's search is strict and returns nothing for a garbled name, so the tool's CANNOT_VERIFY was actually correct behavior on bad input — not a code bug. Added a typo-tolerant fallback to espnAthleteSearch: if the exact full-name search finds nothing, it retries with the LAST name alone (far less likely to be typo'd), fuzzy-matches the intended full name against those results via Levenshtein edit distance, and accepts the closest result only if it's genuinely within a typo or two (never a loose same-surname match). Verified live: "Odan Elliott" now correctly resolves to Oban Elliott (edit distance 1), while genuinely fake names ("Xzyqwp Fakefighter", etc.) still correctly decline — the hard verification requirement is preserved. KNOWN LIMITATION: this can't fix every misspelling. When the surname is very common (e.g. "Oliveira" — ESPN returns 7+ MMA Oliveiras and the last-name search's top-10 may not even include the intended fighter), a misspelled first name like "Michal Oliveira" still can't be resolved safely without risking grabbing the wrong fighter, so it still declines. The more complete future fix would be to resolve fighter names against the correctly-spelled roster that get_game_schedule already pulls from ESPN's event data, rather than re-searching from the model's possibly-typo'd input — noted for later, not done here.
SECOND UFC BUGFIX: SCHEDULE MATCHING WASN'T BIDIRECTIONAL. Elliott vs Oliveira was still declining after the rate-limit fix, but this time from get_game_schedule (rule 8's schedule check), not the fighter-stats check — confirmed by tracing the exact failure point. Root cause: get_game_schedule's name matching only checked one direction (does the ESPN competitor's own name contain the search query). That works fine when the query is a single fighter's name, but for a two-competitor matchup the model naturally passes the combined "Fighter A vs Fighter B" string — which is LONGER than either fighter's individual name, so a competitor's short name can never "contain" it. Confirmed live: querying "Oban Elliott" alone found the event fine, but "Oban Elliott vs Michael Oliveira" (the realistic query shape) matched zero events, for every single UFC matchup tested, regardless of fighter. This wasn't UFC-only either — the same shared getGameSchedule function backs every individual-competitor sport (tennis, WTA, ITF, darts, bowling, NASCAR), so any of them could have hit the identical failure. FIX: matching is now bidirectional both ways — checks whether the competitor's name contains the query AND whether the query contains the competitor's name — with an explicit non-empty guard on both sides (a naive bidirectional check without that guard would make an empty-string competitor name silently match every single event, since "anything".includes("") is always true in JS). Re-tested live against every combination from Drew's actual matchups (combined "A vs B" strings, single names, with/without a period after "vs") — all now correctly resolve to the right event.
UFC FOLLOW-UP: FIXED RATE-LIMIT/CONCURRENCY BUG. After switching to ESPN, all 3 test matchups Drew tried still failed to verify. Root cause found by replicating the exact production logic live: fetchUFCFighterStats fetches both fighters via Promise.all (concurrently) — confirmed that firing both fighters' full ESPN core-API lookup chains at once reliably triggers a 503 rate-limit response, while the exact same requests spaced out by even a second succeed every time. This explains why it wasn't fighter-specific — every UFC request hit the same concurrency burst regardless of who was fighting. FIX: (1) the two fighters are now fetched sequentially instead of via Promise.all, (2) added fetchEspnCoreWithRetry — one retry after a short pause specifically for 503/429 on ESPN's core API, catching any remaining transient hits, (3) added hasRealFightStats() validation — previously, if ESPN's stats endpoint came back completely empty for a fighter, the code would silently produce a blind 50/50 power rating and present it as "real data used," which defeats the entire point of the verification requirement; now it explicitly requires at least one real fight stat (not just bio data like age/reach) before treating a fighter as verified. Re-tested live against all 3 of Drew's matchups after the fix: Elliott/Oliveira and Medić/Rodriguez now both verify successfully with complete real stats; Spasić/Luciano correctly still declines — but for a legitimate reason this time, not a bug: Spasić is making her literal UFC debut and has zero recorded fight stats anywhere (ESPN or ufcstats.com), which is exactly the "one or two declining is fine" case Drew said he's okay with.
WNBA SUBREQUEST FIX. Found the actual biggest cost driver in the whole pipeline: getKeyPlayerFactor's basketball branch (NBA/WNBA/CBB) was fetching individual ESPN stats for up to 15 players PER TEAM — 30 total for both teams combined — just to identify the single highest-usage player to factor into the simulation. Every other sport's key-player check is far cheaper (NFL/CFB only checks QBs, NHL only checks goalies — 1-3 candidates each), so basketball was a massive outlier. NBA gets away with this because BALLDONTLIE saves it subrequests elsewhere in the pipeline that WNBA doesn't have access to, so WNBA was eating this cost in full — the confirmed main reason it kept tipping over Cloudflare's 50-subrequest-per-invocation Workers Free plan ceiling while NBA/MLB/NFL/UFC stayed comfortably under it. Trimmed the basketball candidate pool from 15 down to 6 per team, which still reliably covers a team's actual rotation/starters. Combined with the earlier removal of the automatic background second opinion, this should give WNBA (and CBB, which has the same cost profile) enough headroom to complete reliably. Also confirmed WNBA was already a registered League option in the Bets table's select field — nothing needed there, it just had zero records against it since no WNBA pick had ever logged successfully before now.
UFC DATA SOURCE SWITCHED FROM ufcstats.com TO ESPN. Root cause of continued UFC declines after the diacritic fix: ufcstats.com is now serving a JavaScript bot-detection "Checking your browser..." challenge page for every single query — confirmed by testing it against Jon Jones, arguably the most famous fighter in UFC history, who still has a real profile there and still got the challenge page instead of real data. This meant the worker's fetch() could no longer get real fighter stats for ANY UFC matchup, not just debuts, and is not something to work around (bot-detection is there deliberately, so no attempt was made to defeat it). Found a full replacement instead: ESPN's own core API (sports.core.api.espn.com) hosts the exact same tale-of-the-tape data ESPN's public MMA Fightcenter page displays — confirmed byte-for-byte against a live screenshot (Uroš Medić vs Daniel Rodriguez, UFC Belgrade Aug 1 2026): significant strikes landed/min, strike accuracy, takedown average/accuracy, submission average, plus reach/age/stance, all as clean structured JSON on a domain already used everywhere else in this app. Rebuilt fetchUFCFighterStats() on ESPN (espnAthleteSearch + sports.core.api.espn.com athlete + statistics endpoints) instead of scraping ufcstats.com. Also rebuilt the power-rating formula per Drew's own manual methodology — previously each fighter's rating was computed independently against a fixed baseline; now it's a genuine head-to-head comparison (computeUFCPowerRatings, taking both fighters together), crediting whoever's better on each stat by the size of the margin, with reach and age (younger favored, per Drew) weighted as explicit factors alongside the volume/accuracy stats — sanity-tested against the real Medić/Rodriguez numbers and produces a sensible, non-extreme 48/52 split. Updated all CANNOT_VERIFY messaging, the tool description text (both Gemini/Grok copies), the NASCAR comment analogy, and the workflow page to reference ESPN instead of ufcstats.com.
DIACRITIC NAME-MATCHING BUGFIX. Root cause found and confirmed live: Marina Spasić vs. Stephanie Luciano (a real, scheduled UFC Belgrade fight on Aug 1, 2026 — confirmed directly against UFC.com and ESPN's own live mma/ufc scoreboard, both fighters present) was getting declined as "could not be verified," even though the fight is 100% real. Cause: get_game_schedule's local name-matching does a plain JS .includes() substring check, which is accent-sensitive — ESPN stores her name with the Serbian diacritic ("Spasić"), but the match query (typed the normal way, "Spasic", no accent) can never match it, since ć and c are different characters to a raw string comparison. This isn't UFC-specific — it silently affects any fighter, tennis player, soccer team, or darts player with an accented name (Serbian, Polish, Croatian, Brazilian-Portuguese, etc. — extremely common in MMA and tennis). FIX: added a shared stripDiacritics()/matchName() helper (Unicode NFD normalization + strip combining marks) and applied it to every local name-matching comparison across the file: get_game_schedule (the confirmed root cause), injury report team matching, BALLDONTLIE team matching, ESPN standings team matching, football-data.org soccer standings matching, darts player-list matching, TennisMyLife surface-stats name matching, fetchMarketOdds team/winner matching, and — importantly — findFinalResult's grading logic (both the side-matching and the Win/Loss winner-name comparison), since an accented name could have been silently causing incorrect grades or stuck-Pending picks even after successfully passing verification. UFC's own ufcstats.com fighter-stats check and ESPN's tennis player search are unaffected by this fix since they send the name to those services' own server-side search rather than doing local substring matching.
REMOVED AUTOMATIC BACKGROUND SECOND OPINION. Every Gemini-answered question used to silently trigger a full background Grok analysis (ctx.waitUntil) on every single request — its own Gemini/Grok tool-calling round trip, its own ESPN calls, its own Airtable logging. That was quietly costing every normal request roughly double its actual subrequest count, which is very likely what tipped some WNBA requests (already more subrequest-hungry than BALLDONTLIE-backed sports like NBA/MLB/NFL/NHL) over Cloudflare's 50-subrequest-per-invocation Workers Free plan ceiling, producing the generic "SAM's having trouble connecting" error. Removed the trigger and deleted the now-dead triggerGrokSecondOpinion function entirely. This isn't a loss of functionality — the "get a second opinion" button added to every message already covers this on demand through the exact same /api/chat path (including the same Airtable logging), so Grok's track record now only grows from picks people actually asked for, per Drew's call.
ITF TENNIS GRADING FOLLOW-UP (post-release patch, version held at 1.5.7 per Drew's call). Live-tested ESPN's tennis/atp and tennis/wta scoreboard endpoints directly after the ITF prediction fix shipped: neither one ever surfaces ITF World Tour matches — both only carry main-tour (250/500/1000-level) events, confirmed on a live date with real tournament names returned for each (ATP: Mifel Tennis Open, Mubadala DC Open; WTA: Odlum Brown VanOpen, Memphis Classic, etc., none ITF-level). The previous patch had mapped ITF to ESPN_LEAGUE_PATHS.tennis/atp as a "best-effort" grading source — that was wrong and actively misleading: it silently never found a match, so every ITF pick sat in Pending forever with no visible sign anything was broken. REMOVED that mapping. ITF now has no ESPN grading path at all, same permanent, documented gap as Bowling and Darts (findFinalResult short-circuits immediately via the existing "if (!path) return null" check instead of wasting a scoreboard fetch that could never succeed). ITF predictions themselves are unaffected — they still verify and run correctly via ESPN player-search per the earlier fix; only auto-grading of completed ITF picks is unavailable, and will need to stay a manual Airtable update until a real ITF data source is found.
ITF TENNIS BUGFIX (post-release patch, version held at 1.5.7 per Drew's call). ROOT CAUSE: ITF Tennis was never actually wired in anywhere despite being a listed sport — it was missing from the [PICK] league enum, missing from every "Tennis and WTA" verification/injury/simulation check in both the system prompt and the tool code, and missing from ESPN_LEAGUE_PATHS entirely. The practical effect: whenever a matchup was tagged ITF, rule 8 treated it as a league "with real schedule coverage" (since ITF wasn't in the Tennis/WTA schedule-skip exception list) and required get_game_schedule to confirm it — but get_game_schedule had no ESPN path for ITF at all, so it could never find the matchup and the request hit the hard verification gate and got declined. This is why it looked like only "some" ITF matches worked: any that got asked about using the TENNIS or WTA league tag directly (bypassing the ITF label) still went through the working ATP/WTA path and succeeded, while ones correctly tagged ITF always failed. FIX: added ITF as a first-class alias everywhere Tennis/WTA are handled — the [PICK] enum, the tool schema league lists, the injury-report individual-sport check, the run_matchup_simulation verification/CANNOT_VERIFY branch (reuses the same ESPN player-search + ranking + win/loss lookup as Tennis/WTA), the TML surface-adjustment skip (TennisMyLife only covers ATP tour level, so ITF now skips it the same way WTA already did), and rule 8's schedule-gate exception list (ITF now skips the schedule check and relies on ESPN player verification only, same as Tennis/WTA, since ITF has no reliable ESPN schedule board either). Workflow page's verified-sports list updated to include ITF alongside Tennis/WTA. (No ESPN_LEAGUE_PATHS entry was added for ITF — see the follow-up changelog entry above; there's no real ESPN source to grade ITF picks against.)
VERIFICATION HARD-GATE + WRONG-PITCHER BUGFIX + WORKFLOW PAGE UPDATE (post-release patch, version held at 1.5.7 per Drew's call). (1) NO MORE FAKE/THEORETICAL MATCHUPS: closed every "if we can't verify this, estimate/guess and proceed anyway" fallback across the whole prompt and tool layer. Team sports now decline outright (no [PICK], no log) if get_game_schedule can't find the real matchup or if live team stats can't be pulled — previously a stats-fetch failure fell back to a qualitative guess that still got logged. UFC/Bowling/Darts/Tennis/WTA: if their real-data check fails (ufcstats.com, pba.com, PDC tour cards, ESPN rankings), it's now a hard decline instead of the model estimating a power rating and proceeding — removed at the CODE level, not just the prompt, so it can't be talked into estimating anyway. Horse Racing is disabled entirely: there's no schedule source and no stats source for it at all (Equibase blocks scraping), so unlike every other sport there was never any way to verify a race or named horses were real — rather than leave that wide open, predictions for it now always decline. NASCAR keeps its necessary estimate (no live per-driver stats exist, full stop) but now requires get_game_schedule to first confirm both drivers are actually entered in a real race — caught a real ESPN limitation while building this: entry lists are only published for the immediate next race, not races further out, so this verification is reliable for the next race and weaker beyond that. (2) BUGFIX — get_game_schedule couldn't actually match individual athletes: its matching only checked competitions[0] and only team-style fields, so it could never match UFC fighters or NASCAR drivers at all (they use .athlete, nested across many sub-competitions per event, not .team). Fixed to search every sub-competition and match both team and athlete name fields — required for the verification gate in item 1 to work correctly instead of wrongly blocking every real UFC/NASCAR matchup. (3) DECLINED PREDICTIONS NO LONGER COST A FREE PICK: found and fixed a side effect of item 1 — the daily/paid prediction allowance was being deducted before the model even attempted verification, so a correctly-declined fake matchup was still burning one of the user's 3 free daily predictions for nothing. Deduction now only happens when a real [PICK] actually resulted. (4) WRONG-PITCHER BUGFIX (the important one): getMLBProbablePitchers picked the FIRST upcoming game between two teams within a 10-day window, with no awareness of which specific night was being asked about. Verified live and reproduced exactly — with the Tigers mid-series against the Orioles (games on consecutive nights, each with a different starter), the tool returned Game 1's pitcher (Keider Montero) regardless of which night was actually asked about, when the real answer for a later night in the series was Tarik Skubal, a completely different pitcher. Fixed by adding an eventDate parameter to run_matchup_simulation — SAM now passes the exact date it already confirmed via get_game_schedule (rule 8) straight into the pitcher lookup, which filters to that specific game instead of guessing the first one. Verified fixed live: without eventDate the tool returned Game 1's starter, with eventDate set to Game 3's date it correctly returned Game 3's actual starter. Applied the same fix to getHomeAwayContext/findScoreboardEvent (home-field advantage), which had the identical latent bug for any two teams that meet more than once in a season (e.g. divisional rematches) — not MLB-specific, just most visible there since series are so common. Confirmed the NFL/college key-player mechanism (QB rating) isn't exposed to this specific bug since it doesn't do date-based schedule matching at all, only home-field advantage shared the underlying issue and is now fixed the same way. (5) WORKFLOW TAB UPDATED: the in-app /workflow page ("How SAM works") was rewritten to match everything shipped recently instead of describing the old pipeline — added a distinctly-colored new step for the verification hard-gate (deliberately NOT reusing the existing red "logged pick" style, which would have been misleading), updated the Airtable-cache step to explain same-day-cached vs next-day-fresh-rerun behavior, updated the team-sport branch card to list home-field/injury/pitcher/key-player factors, updated the individual-sport branch card to explain the verify-or-decline policy per sport, and updated the intro/closing copy to state the no-fake-matchup guarantee alongside the existing no-odds guarantee. Legend updated with the new indicator color.
SIMULATION ACCURACY + LOGGING RELIABILITY PASS (post-release patch, version held at 1.5.7 per Drew's call). (1) CHATTINESS REMOVED: system prompt rule 4 no longer tells SAM to follow a prediction with podcast-style color commentary, storylines, and an invitation for the user to keep chatting — it now stops cleanly right after the prediction. (2) DOUBLE-LOGGING RACE FIXED: logBetToAirtable's dupe check (query Airtable, then write) was a classic check-then-act race — two near-simultaneous requests for the same matchup could both see "nothing pending yet" and both write, producing duplicate Pending records. Fixed with an in-worker lock keyed on provider:matchup:eventDate so overlapping calls for the same pick now queue and serialize instead of racing. (3) STALE PLAYER NAMES FIXED: SAM was naming specific players (a departed pitcher, in one reported case) from its own training knowledge with no live verification. Added a get_team_roster tool plus a stronger system-prompt rule requiring any named player to actually appear in a live ESPN roster pull. This is now deterministic, not prompt-compliance-dependent — run_matchup_simulation automatically fetches and bundles both teams' current rosters at the top of its own response for every team-sport matchup, so the data arrives before the model writes anything and can't be skipped by forgetting a separate tool call. (4) IDENTICAL CACHED REPLIES GUARANTEED: previously the "already logged" shortcut only fired on a strict "X vs Y" phrasing match before calling the model; any other phrasing skipped it and got a freshly-generated (and differently worded) reply even though the pick was already Pending in Airtable. Added a shared buildCachedReplyText() template used by both the pre-model fast path and a new post-model backstop (findExistingPendingRecord, matched on matchup + event date) — any rephrasing that slips past the fast path still gets caught after the model responds and swapped for the exact same canned message, byte-for-byte, with no re-log and no fresh odds/commentary appended. (5) ROSTER FETCH BUG FIXED FOR NBA/WNBA/SOCCER/CBB: ESPN returns MLB/NFL/NHL/CFB rosters grouped by position ({position, items:[...]}) but returns NBA/WNBA/Soccer/CBB rosters as a flat player list with no items wrapper at all — the original get_team_roster only handled the grouped shape, so it was silently returning "no roster data" for every NBA/WNBA/Soccer/CBB team since the roster-verification feature (item 3) shipped. Fixed via two new shared primitives, resolveEspnTeam() and fetchTeamRosterRaw(), that normalize both shapes into one flat player list; verified live against real rosters (Celtics: 16 players, Liverpool: 39, Rangers: 26) before shipping. (6) MLB STARTING PITCHER FOLDED INTO THE ACTUAL SIMULATION MATH: previously roster/player data only shaped the writeup, never the win%. New getMLBProbablePitchers() pulls each team's live probable starter + current ERA directly from ESPN's scoreboard (a field ESPN publishes specifically for MLB, verified live), converted via pitcherERAToFactor() into a lambda multiplier that suppresses the opposing team's expected runs — a true ace now measurably lowers the opponent's simulated scoring, not just gets a mention. (7) INJURY REPORTS NOW FEED THE MATH TOO, ALL TEAM SPORTS: new injuryReportToFactor() counts real Out/Doubtful/Questionable entries from the already-fetched injury report and scales that team's own offense/defense accordingly (capped, never a boost, only ever neutral-or-worse). This runs for every league that reaches the generic team-sport branch — NFL, NBA, WNBA, MLB, NHL, Soccer, CFB, CBB — not just MLB. (8) KEY-PLAYER QUALITY FOLDED IN FOR NFL/CFB/NBA/WNBA/CBB/NHL: new getKeyPlayerFactor() identifies each team's actual key contributor (QB for NFL/CFB via QBRating, leading scorer for NBA/WNBA/CBB via PPG, starting goalie for NHL via save%) and fetches their real current ESPN stat to adjust the simulation. Important correctness note from live testing: ESPN roster order is NOT depth-chart order (verified — Dak Prescott, Dallas's clear starter, was listed behind two backups), so "first player at the position" would have been wrong. Fixed by fetching usage (attempts/games/starts) for every same-position candidate and picking whichever one actually has current-season snaps. A second bug caught in the same testing pass: the first version of that fix let a bench player's stale career total outrank the real starter's smaller current-season total — usage is now strictly current-Regular-Season-only with zero fallback, so an inactive player correctly scores 0 usage instead of borrowing career numbers. If the identified key player is currently listed Out/IR, their stat boost is suppressed entirely (factor reverts to neutral) rather than crediting a team with a bench-level replacement's numbers — the separate injury-count factor (item 7) still applies its own penalty for that same absence, so it isn't double-ignored. Soccer was intentionally left out of this specific mechanism at Drew's call (low value while the Premier League is in its off-season with no current club-competition stats published yet, plus not worth the added per-request ESPN calls for a sport he doesn't prioritize) — Soccer still gets items 7 and 9. (9) REAL HOME-FIELD/COURT/ICE ADVANTAGE ADDED, ALL TEAM SPORTS: this was completely unmodeled before today. New findScoreboardEvent()/getHomeAwayContext() pull the actual confirmed home team for the specific matchup from ESPN's schedule (order-independent — works whether the home team is passed as teamA or teamB, verified live both ways) and feed it into the sim's existing (previously always-neutral) travel-tier multiplier. Applies universally, including MLB and Soccer, unlike item 8. (10) KNOWN REMAINING GAP: rest-days and recent-form-differential multipliers in calculateLambdaMultiplier are still unpopulated dead code, same as before this patch — not addressed today, flagged for a future pass if wanted.
AUTO-GRADE DATE-MATCH BUGFIX (post-release patch, version held at 1.5.7 per Drew's call): findFinalResult's exact-date check only ran when multiple candidate ESPN events matched a pick's team/fighter names (allMatches.length > 1); a single fuzzy name match skipped the date check entirely and was accepted as-is even if it fell on the wrong day within the \xB11-day search window. This let loosely-matched UFC undercard fighters (common surnames, generic names) get graded against the wrong event on the wrong date, sometimes producing a real Win/Loss off a real 'winner' flag but with a score summary of 'N/A - N/A' since MMA competitors don't carry numeric scores the way team sports do \u2014 that blank score summary was the tell. Fix: the exact Event Date check (America/New_York) now always runs whenever eventDateStr is available, regardless of how many candidates were found; if the single match doesn't land on the exact date, the pick is left Pending instead of being graded. No behavior change for the common case where the correct event is the only candidate and already falls on the right date. TENNIS + GOLF + CASINO BUTTON REMOVAL + REAL CALIBRATION SYSTEM. (1) TENNIS ADDED: real stats, not just a power-rating guess — run_matchup_simulation now looks up both named players via ESPN's player search (site.web.api.espn.com/apis/search/v2), pulls their current ATP/WTA ranking and season singles win/loss record from ESPN's core API, and blends the two into a 20-100 power rating for runFightSimulation() — same head-to-head mechanic as UFC/Bowling/Darts. Falls back to a model-estimated power rating if either name doesn't resolve to an ESPN tennis player. Schedule lookups also now work for Tennis (ATP) and WTA via the existing get_game_schedule tool, routed through ESPN's tennis/atp and tennis/wta scoreboards — no known gap here, unlike Bowling/Darts. (2) GOLF ADDED: a genuinely different shape of prediction since golf is a 100+ player field, not a two-competitor matchup — so this isn't run_matchup_simulation. A new run_field_simulation tool pulls this week's full PGA or LPGA tournament field from ESPN's scoreboard, fetches each player's real current-season scoring average per round from their ESPN overview page (in parallel), and runs a 10,000-trial Monte Carlo simulating 4 rounds per player off their own scoring-average-centered distribution — win% is the share of trials where that player posts the field's lowest 4-round total. Players ESPN has no scoring-average profile for (rookies, unranked amateurs) are simulated off the field-wide average scoring rate instead of being dropped, so the field size stays accurate. Trial count is 10,000 rather than the usual 25,000 given the cost of simulating 100+ players per trial instead of 2 — flagged in the tool's own response text, not hidden. Returns the top 10 contenders by win% plus the tournament name; the model should present this as a ranked list of contenders rather than a single [PICK], since there's no single opposing side to log a win/loss grade against. (3) SPORTSBOOK BUTTONS REMOVED: the row of 4 sportsbook/casino links (DraftKings, FanDuel, BetMGM, Caesars) under the chat bar is gone — HTML, CSS, and the openSportsbookWindow popup-window JS all removed. Version label was intentionally held at "SAM 1.5.6"/"SAM DS 2.0" through development of this batch of changes, then bumped to "SAM 1.5.7"/"SAM XAI 2.1" once the whole batch (Tennis, Golf, sportsbook removal, calibration system) was ready to ship as one release. (4) REAL CALIBRATION SYSTEM ADDED: rule 6's soft "please self-calibrate based on this history text" instruction is now backed by an actual computed number instead of relying on the LLM's own judgment. New Supabase table sam_calibration (same project as sam_bettors) holds one row per (provider, league): a running bias in percentage points, plus n_graded/wins/losses. Every time gradePendingBets grades a Win or Loss, it now also runs one online-update step — delta = 3% × (actual outcome − stated confidence), added to that league's running bias, clamped to ±15 points so no single bad streak can run away with it. run_matchup_simulation reads the current bias for the calling provider+league before returning its result and shifts both win percentages by that exact number (mirrored so they still sum to 100) — done as a single wrapper around the existing dispatcher's output text (every branch already returned "TeamName: NN% win rate" in the same shape, so one regex-based post-processor covers UFC/Bowling/Darts/Tennis/NASCAR/team-sports/etc. without touching each branch). When a bias is applied, the tool response tells the model plainly that the number is already calibrated and not to layer rule 6's soft self-calibration on top of it, so the two mechanisms don't double-count. Golf's run_field_simulation is untouched — no single win/loss to calibrate against in a full-field event. (5) CALIBRATION BACKFILL ENDPOINT: a new GET /admin/backfill-calibration?key=SESSION_SECRET route (reuses the existing session secret, no new env var) walks every already-graded pick in both Airtable bases, buckets them by league, computes each bucket's actual-hit-rate-minus-stated-confidence as a batch average (not a slow one-by-one replay), and seeds sam_calibration with that number plus the real win/loss counts — so leagues with real history (MLB, UFC) start calibrated on day one instead of drifting up from a blank bias=0 one graded pick at a time. Safe to re-run — upsert on (provider, league) means it just recomputes and overwrites, doesn't double-count. (6) KNOWN CAVEAT — SMALL-SAMPLE NOISE: Tennis, Golf, Darts, and Bowling only have a handful of graded picks each right now, so their calibration bias is currently dominated by sampling variance, not a reliable read on real miscalibration — the standard error of an observed hit rate shrinks with sample size (roughly 1/√n), so at n=4 a 3-1 or 1-3 stretch (pure chance) can swing the bias by several points, while a mature bucket like MLB (150+ graded) barely moves per pick. Don't read much into those four leagues' bias numbers until each has ~20-30+ graded picks — this isn't a bug, there's no code fix for it, it just needs more games to play out. (7) SCOPE CLARIFICATION: the calibration system (item 4) corrects confidence calibration — how far a stated win% is off from real historical hit rate for that league — not pick accuracy. It cannot fix a pick where the simulation favored the wrong side; it only tightens or loosens the confidence number attached to whichever side was already picked. The Monte Carlo simulation's underlying stats and math are unchanged and exactly as accurate (or inaccurate) as before this release. (8) POST-PREDICTION MATCHUP COMMENTARY & TRANSPARENT SIMULATION Q&A: System prompt rules updated to (a) instruct SAM to deliver lively, podcast-style color commentary, key storylines, X-factors, and game-script hot takes immediately following a prediction split, inviting user dialogue; and (b) empower SAM to answer user follow-up questions openly and transparently about how its 25,000-trial Monte Carlo simulation engine, recent-form weighting (last 8 games), injury metrics, and conservatism shrinkage work without being evasive or logging duplicate pick tags. UI + FOLLOW-UP CHAT PATCH (post-release patch, version held at 1.5.7 per Drew's call): (9) Brain emoji restored to the "Thinking..." indicator shown while a request is in flight. (10) The "SAM's Reasoning" thought bubble is now collapsible — clicking its label toggles a ▾/▸ chevron and shows/hides the reasoning box; still renders by default same as before. (11) FOLLOW-UP SIMULATION Q&A: the chat can now actually answer follow-up questions about a simulation it just ran, not just describe its methodology in the abstract. The client keeps the last 6 exchanges (user message, SAM's reply, and that turn's internal simulation trace) in an in-memory array only — nothing persisted to Airtable, Supabase, or browser storage, gone on refresh — and sends it with each new message. Server-side, a new buildHistoryContext() formats that into a plain-text block and prepends it to the augmented message sent to Gemini/Grok (both the primary-provider and cross-provider-failover paths), so a follow-up like "why the Lions and not the Bears" or "what was the split before calibration" gets answered from the real numbers/trace of that turn instead of being re-simulated or guessed at.
PROVIDER SWAP + AUTO SECOND OPINION + SPORTS EXPANSION. (1) DEEPSEEK REPLACED WITH GROK: the second model slot now runs on xAI's Grok (grok-4.5) via its OpenAI-compatible /v1/chat/completions endpoint, instead of DeepSeek. Same tool schema, same multi-hop tool-calling loop, same [PICK] parsing, same reasoning_content shape — only the endpoint, model name, and GROK_API_KEY (+ backups) env vars changed. Now labeled 'SAM DS 2.0' in the model picker and in logged picks (was 'SAM DS1.4'). (2) AUTO SECOND OPINION: whenever Gemini answers as the primary model (the default path), the exact same question now also gets fired at Grok in the background — no added latency or change to the user-visible reply. If Grok lands on its own clean [PICK], it's logged as a separate record (its own Model Version, its own calibration against its own track record) in the same base that used to hold DeepSeek's picks — so that base now fills with Grok's independent second opinions on every matchup Gemini handles. If the user explicitly picks Grok as primary, no redundant self-second-opinion runs. Any second-opinion failure (missing key, API error, no clean pick) is swallowed and logged server-side only — it never affects or delays the primary response. (3) SERIES/PLAYOFF BUGFIX: multi-game series between the same two teams (e.g. Twins vs Guardians, game 2 of 5) were breaking prediction logging, all for the same root cause: nothing checked the actual game date, only the team names. A pre-model cache shortcut matched on team names alone, so once game 1 of a series was Pending, every later question about the same two teams got served game 1's stale cached result instead of ever reaching Gemini, the simulation, or a new Airtable write — now it only takes that shortcut when exactly one Pending record matches; 2+ (a series) skips straight to the real pipeline. The write-time dupe guard had the same blindness and now also checks Event Date so each game of a series logs as its own record. Auto-grading had the same risk too — it could've graded game 2's Pending record with game 1's score — so it now requires an exact Event Date match before grading, and leaves the record Pending rather than guessing if none match exactly. (4) WNBA ADDED: full team-sport support (schedule, injuries, simulation) — routes through the same bivariate/normal scoring pipeline as NBA (stdDev 12, league-average baseline 83.0 pts). BALLDONTLIE doesn't cover WNBA, so it goes straight to the ESPN standings fallback for offense/defense — same behavior NFL/NBA/MLB/NHL get whenever BALLDONTLIE itself is unavailable, not a new code path. (5) NASCAR ADDED: schedule lookup works like any other ESPN-backed league (racing/nascar-premier). Since NASCAR is a 30-40 car field, not a two-team matchup, and there's no equivalent of ufcstats.com to auto-pull real per-driver stats, run_matchup_simulation treats it like the UFC power-rating fallback — a head-to-head finish-ahead-of simulation between the two named drivers, using power ratings you estimate from recent finishes/track fit/form (reuses runFightSimulation() unchanged). No team injury report for NASCAR, same reasoning as UFC (individual competitors, not team rosters). (6) GEMINI MODEL BUMP: primary model string updated from gemini-3.5-flash to gemini-3.6-flash. Grok (the DS 2.0 slot) is unaffected. (7) MODEL LABEL FIX: the model-picker label, Airtable test-ping value, and page display text were still hardcoded to "SAM 1.5.5" from before this release — all instances now correctly read "SAM 1.5.6" so logged picks and the visible label match the actual running version. (8) SIGN-IN GATE: the whole site now requires a login — a small, Drew-approved allowlist stored in the Bettors table (Username + Password Hash fields, PBKDF2-SHA256, never plaintext). Sessions are stateless signed cookies (HMAC-SHA256, 30-day expiry) — no KV/D1 needed, just one new SESSION_SECRET env var. Static assets (favicon, background image) and the login page itself stay reachable pre-auth; everything else redirects to /login without a valid session. Logged-in sessions also now carry the Airtable record ID of whoever is signed in, so every pick — including the Grok background second opinion — links back to whoever triggered it via the Bettor field, instead of logging anonymously like before. (9) BILLING: 3 free predictions per bettor per day (resets at midnight ET), then paid packs via Stripe Checkout — $5/10 predictions, $50/100 predictions, $200/lifetime unlimited. Paid credits never expire. Usage state lives on the Bettors record (Credits Remaining, Lifetime Access, Free Predictions Used Today, Free Reset Date) — same table the sign-in gate already uses. A cache hit or a failed call never counts against the daily/paid allowance — only a real, completed prediction does. Hitting the limit shows a paywall message in-chat with one-click links to each tier; a Stripe webhook (signature-verified, same HMAC-SHA256 pattern as session cookies) credits the right bettor the moment a purchase completes. Two new secrets required: STRIPE_SECRET_KEY and STRIPE_WEBHOOK_SECRET. (10) AUTH + BILLING MOVED TO SUPABASE: sign-in and credit/usage tracking now live in a new public.sam_bettors table in the existing grimaldi.tv Supabase project, not Airtable — real Postgres, no per-request rate limit, RLS enabled with no policies so only the worker's service-role key can touch it, never a browser. Picks/bets logging stays on Airtable exactly as before; since a cross-system linked record isn't possible, the old Bettor linked-record field on Bets is replaced by a plain-text Bettor Username field. Two new secrets required: SUPABASE_URL and SUPABASE_SERVICE_ROLE_KEY. (11) SELF-SERVE SIGNUP: anyone can now create their own account at /signup (username 3-32 chars, password 8+ chars, name/email optional) — no more Drew-approval gate on account creation. New accounts start on the standard 3-free-predictions-per-day tier like everyone else, same paywall applies. Login and signup pages now cross-link to each other. (12) BOWLING ADDED: real stats, not just a power-rating guess — run_matchup_simulation now pulls each bowler's actual most-recent-season PBA Tour scoring average directly from their pba.com profile page (PBA has no search endpoint, so this guesses the name-based URL slug directly, e.g. "EJ Tackett" -> ej-tackett) and converts it to a 20-100 power rating for runFightSimulation() — same head-to-head mechanic as UFC/NASCAR. Falls back to a model-estimated power rating only if a profile isn't found. No schedule source is wired in yet (PBA's schedule isn't on ESPN) — event date comes back as unknown until that's built, a known gap, not a silent one. (13) HORSE RACING ADDED: same power-rating head-to-head model as NASCAR (runFightSimulation, two named horses compared directly) — but unlike NASCAR, no real-stats attempt is made at all. Equibase, the official North American horse racing data source, explicitly prohibits automated scraping in its own Terms of Use and sits behind Imperva bot protection regardless, so this always uses a model-estimated power rating, by design, not as a temporary gap. (14) DARTS ADDED: real stats, using dartsdatabase.co.uk since the official pdc.tv is a JS-only app with no server-rendered data to fetch at all. No search endpoint exists on dartsdatabase.co.uk, so run_matchup_simulation fetches the real Tour Card Holders list (all 128 current PDC pros), matches the given names against it, then pulls each matched player's real current-season 3-dart average from their profile page and converts it to a 20-100 power rating for runFightSimulation() — same head-to-head mechanic as UFC/Bowling. Falls back to a model-estimated power rating if either name doesn't match a current tour card holder. No schedule source for darts either — same known gap as Bowling, event date comes back unknown until that's built separately. (15) OLD DISCLAIMER GATE REMOVED: the click-through "Before you enter" card that showed on every page load (predating the real sign-in system) is gone — HTML, CSS, and the unlockSAM/showSAM JS all removed. It was a leftover gate from before login existed and was stacking as a redundant second gate after every real sign-in. The initial greeting now fires directly on page load instead of waiting on that click. (16) PURCHASE CONFIRMATION ADDED: the /?purchase=success redirect Stripe sends users back to after Checkout used to do nothing at all — the page loaded plain, with the only signal a purchase went through being that the next prediction happened to work. Now the tier is passed through the redirect and the page shows an explicit in-chat confirmation naming what was added ("10 predictions", "100 predictions", or "lifetime access"), with a matching message on /?purchase=cancelled too. The query params are stripped via history.replaceState right after so a refresh or back-button press doesn't replay the message.
Simulation engine supercharge (Grok recommendations). (1) BIVARIATE POISSON: Team sport simulations now use a bivariate Poisson distribution with correlation 0.28 — scores are no longer independent, capturing the real-world phenomenon that high-scoring games tend to be high-scoring for both teams (e.g. NFL shootouts, NBA pace). The shared Poisson component pulls the two teams' outputs together slightly, producing more realistic joint score distributions. (2) 25,000 TRIALS: Trial count raised from 20,000 to 25,000 for lower statistical variance and more stable win percentages across runs. (3) IMPROVED LAMBDA CALCULATOR: A dedicated calculateLambda() function now adjusts each team's expected score before simulation using four real-world factors — rest days (back-to-back penalty, well-rested bonus), travel distance tier (home/short/medium/long haul), recent form (last-5-game scoring trend vs season average), and injury severity (starter-out vs key-player-questionable penalties). These contextual modifiers stack multiplicatively on top of the existing offense/defense matchup math. (4) BRIDGING FUNCTION: A bridgeToTeamSim() function translates the adjusted lambdas and contextual metadata into the bivariate simulation, cleanly separating lambda calculation from simulation execution. (5) UFC UNCHANGED: runFightSimulation() is untouched — power-rating model still applies there. (6) STATS PAGE LINK: /stats page's Airtable link fallback updated to a fresh invite link. (7) THINKING BUBBLE EMOJI: restored the missing 🧠 on the 'Thinking...' indicator, which had been dropped entirely rather than just mis-encoded. (8) AUTO-GRADING: the existing Cron Trigger (every 6 hours) previously had no scheduled() handler to call, so it fired into nothing — added one. It now pulls every Pending record from both bases, checks ESPN's scoreboard for a completed game matching the matchup and date, and writes Win/Loss/Push into Result plus a final-score line into Notes. Anything ESPN hasn't finished yet, or where a winner can't be confidently matched, is left Pending for the next run rather than force-graded.
Reliability and accuracy pass. (1) CROSS-MODEL FAILOVER: if every key for the requested provider fails, the worker now automatically retries the other provider (Gemini vs DeepSeek) before giving up, using whichever keys are already configured. The reply gets a short notice when this happens, and the Airtable log/Model Version field correctly credits whichever model actually answered, not the one originally requested. (2) SIMULATION TRACE EMOJI FIX: Gemini's path was using an invalid JS escape sequence for the dice-emoji run_matchup_simulation marker in the visible reasoning, so it rendered as garbled text instead of the emoji, making it look like the simulation hadn't run even when it had. Now matches DeepSeek's already-correct version. (3) TRIAL COUNT WORDING FIX: two places (system prompt rule 2, and the UFC power-rating-fallback tool response) still described the simulation as 10,000-trial in the text even though the actual simulations have run at 20,000 trials for a while; both now correctly say 20,000, matching the real trial count everywhere else. (4) DEEPER TRACK RECORD READ: fetchPredictionHistory's page cap raised from 3 pages (300 records) to 10 pages (1,000 records), so a full season's worth of graded picks feeds the calibration instead of just the most recent 300.
Bugfix on the dual-model build. SCHEDULE LOOKUP DATE-WINDOW FIX: get_game_schedule was calling ESPN's scoreboard endpoint with no dates range, which defaults to "today only." NFL/NBA/MLB have games most days so this rarely showed, but sparse-schedule sports — UFC especially, roughly one card every 1-2 weeks — would come back as "no games found" for any fight card not happening that exact day. Fixed by passing an explicit 45-day-forward dates=YYYYMMDD-YYYYMMDD range to the same ESPN call, matching the fix already shipped on ai2/DS1.0. DEEPSEEK LABEL BUMP: the DeepSeek option in the model picker now reads "SAM DS1.2" (was DS1.0) to reflect this fix; Gemini's side stays at 1.5.3 since it was unaffected — the schedule tool is shared code, but this changelog only bumps the model-facing version numbers that actually changed behavior for the user.
Conservatism pass on top of 1.5.2, plus a later dual-model update. CONSERVATISM: (1) DEEPER SHRINKAGE: CONSERVATISM_SHRINKAGE lowered from 0.70 to 0.50, so every raw Monte Carlo win% is pulled further back toward a 50/50 coin flip before it's shown, logged, or handed to the model (e.g. a raw simulated 80% edge now reports as ~65% instead of ~71%). (2) SLOWER RECENT-FORM REACTIVITY: the BALLDONTLIE offense/defense blend for team sports (NFL/NBA/MLB/NHL) changed from 65% last-8-games / 35% season average to an even 50/50 split, so a hot or cold short stretch no longer dominates the simulation's inputs before shrinkage even applies. Together these two changes compound — flatter inputs feeding a flatter output curve — intentionally trading some upside on confident calls for fewer overconfident misses. DUAL-MODEL UPDATE: (3) MODEL PICKER: Added a dropdown next to the chat input so anyone can switch between SAM 1.5.3 (Gemini, gemini-3.5-flash) and SAM DS1.0 (DeepSeek, deepseek-v4-flash with thinking mode) per message — both share the same tool set, Monte Carlo simulation engine, and conservatism/calibration logic, just through each provider's own tool-calling format. Gemini remains the default; DeepSeek is opt-in via the picker. (4) MODEL-TAGGED LOGGING: Picks now record which of the two models made each call in the Model Version field, so Gemini and DeepSeek picks can be graded and compared side by side in the same Airtable base. (5) EMOJI FIX: Corrected long-standing mojibake-corrupted emoji across the suggestion chips and thinking/reasoning indicators — swapped in a plain arrow style for suggestion chips and restored a icon for the thinking/reasoning display. (6) MODEL PICKER STYLING: A two-line title/subtitle dropdown with a blue selection checkmark. (7) DEEPER THINKING: When DeepSeek is selected, its reasoning_effort runs at high for more careful reasoning on every pick; Gemini stays at its existing medium thinking level given its output tokens (including reasoning) bill at a much higher rate. (8) SAVE TO HOME SCREEN: Added a standalone button at the bottom of the page — triggers a real one-tap native install on Chrome/Android, falls back to manual steps only where no native install API exists (iOS Safari).
Accuracy, polish, and market-integration patch on top of 1.5.1. (1) LIVE SCORES & TIMES: Enhanced schedule lookup tool to fetch and display live/final scores and current status of matches. (2) REVERT LOGO: Restored the original favicon logo branding. (3) 20,000 SIMULATION TRIALS: Upgraded the Monte Carlo simulation trial count from 10,000 to 20,000 to increase resolution and lower statistical noise. (4) COLLEGE SPORTS SUPPORT: Configured dedicated scoring models, fallbacks, and paths for College Football (NCAAF/CFB) and College Basketball (NCAAM/CBB). (5) STANDINGS COLLISION FIX: Resolved a bug where short abbreviations like 'gp' collided with longer field names (e.g. 'avgpointsfor'), corrupting statistics and resulting in random coin-flip predictions. (6) INTEGRATED ODDS CHECKING: Allowed SAM to analyze sports betting lines, spreads, and gambling odds from the very beginning of the prediction process. (7) ODDS MISMATCH DETECTION: Instructed the AI to identify value opportunities (+EV edge) by comparing Monte Carlo win percentages against market implied probabilities. (8) PARLAY PREDICTIONS: Equipped SAM to suggest value-parlays by pairing high-edge picks together, with dedicated suggestion chips on the home and matchup views. (9) ODDS REMOVED: Reverted item (6)/(7) above — SAM no longer references, invents, or compares against sports betting lines, spreads, or gambling odds anywhere in the prediction process (there was never a real odds data source wired in, so this was a hallucination risk, not a real market signal). Picks are stat-based only again, per the original 1.0 design. (10) COLLEGE LEAGUE NAMING FIX: Added a league-alias normalizer so "College Football", "NCAA Football", "CFB", and "NCAAF" (and the basketball equivalents) all resolve to the same internal key regardless of which wording the model uses in a tool call. The [PICK] tag now logs college picks under the exact labels "College Football" / "College Basketball" to match the Airtable tracking sheet, instead of fragmenting into NCAAF/CFB/NCAAM/CBB as separate values.
Three-part accuracy and polish patch on top of 1.5's engine work. (1) CONSERVATISM: picks are now more conservative across the board — added a CONSERVATISM_SHRINKAGE constant (0.70) that pulls every raw Monte Carlo win% partway back toward a 50/50 coin flip before it's ever shown, logged, or handed to the model (e.g. a raw simulated 80% edge now reports as 71%, a raw 55% reports as ~54%). Applies uniformly to both the team-sport simulation and the UFC fight simulation, reflecting real uncertainty in season-average inputs rather than overstating confidence on a single computed run. It's one tunable constant (0 = always report a flat coin flip, 1 = old/no-shrinkage behavior), and stacks underneath the existing rule-6 track-record calibration — shrinkage sets a more conservative floor first, calibration can still move it from there. (2) RECENT-FORM PATTERN RECOGNITION: added to the BALLDONTLIE stats pipeline (NFL/NBA/MLB/NHL). Previously every completed game fetched for a team's season (up to 25) was averaged flat, so a hot or cold recent stretch got diluted back toward the full-season number by games from months ago. Games are now explicitly sorted newest-first (the API's own ordering wasn't guaranteed chronological), and offense/defense is computed as a weighted blend: the last 8 games count for 65% of the final number, with the fuller season average filling in the remaining 35% as a stabilizing baseline. Early in a season, before a team has more than 8 games played, it falls back to the plain season average since there's no meaningful recent-vs-season split yet. When the blend is used, the simulation's reasoning trace says so explicitly (sample size and weighting), so a recent-form-driven pick is visibly distinguishable from a plain season-average one. (3) LOGO FIX: Uncle Sam hat promo graphic. Deliberately cropped to exclude the graphic's "Betting Edge / Live odds" subtitle text so it fits the new styling and avoids conflicting with our no-gambling-odds disclaimer. Served via a dedicated route below, not inlined per-page.
Initial launch — patriot gate screen, dual-key Gemini fallback, stat-based predictions only (no gambling lines, spreads, or odds, ever).
Proprietary Technology built by GRIMALDI.TV