BUGS.md keeps every finding's original description and gains, per entry, what was actually done and where its regression test lives — including the two entries fixed differently from the plan (B-15 validates at startup, B-09 kept both callers plus a bounded retry) and the one only partially fixed by decision (B-16, where shipping the guide was deferred). It also gains a Runtime verification section, which is the part worth reading: what the live Docker deployment actually demonstrated (startup validation on the real .env, the listener connecting and holding, rounds cycling, the migration applied, and the reconciler's missing-tx heuristic checked against the real server's error message) separated from what has no runtime evidence at all — nothing has spent money since the restart, so the two-phase write, the RBF retargeting, the reconciler's actual behaviour and the dust path are unit-tested only. A green suite is not a working deployment, and the file now says so. CLAUDE.md documents the two things a reader would otherwise have to reverse- engineer: the transaction lifecycle (why rows are written before broadcasting, what each PendingTransaction status means, why spent_txid must track the current txid, and that one-active-round is now a DB invariant) and the Electrum connection's rotation/keepalive/timeout behaviour. Its Known gaps list is rewritten to say what is still open now that transaction-level state self-heals but round-level state does not. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
34 KiB
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Language
The user communicates in Italian in chat — reply to them in Italian. Everything written to the repository (code, comments, commit messages, docs, this file) must be in English. Reasoning/thinking should also be done in English.
Project status
All 10 build-order stages from /home/davide/.claude/plans/scalable-mixing-sloth.md are code-complete and unit-tested (137 tests green): project skeleton, DB schema + Alembic migrations, auth, HD wallet derivation, Electrum client, deposit detection, bet flow, round/draw engine, payout, withdrawal, RBF fee-bump, admin config + audit log. Beyond the original 10 stages: a Docker + Caddy deployment (see below), a full admin dashboard (/admin), a static test UI for the user-facing flow (/), a pending-inclusive balance display (see "Balance display" below), and a Server-Sent Events push channel layered on top of the original polling (see "Real-time updates" below).
Real-money verification on mainnet, done so far: registration + address derivation, deposit crediting (1-conf), a real 10 PLM bet (broadcast, confirmed, change credited back), and a full round cycle — close → draw (real block hash) → payout (70/30 split, exact sat math verified against the broadcast tx) → confirmation → round closed → next round auto-opened. Withdrawal and the RBF bump path are unit-tested but have never been exercised against a live broadcast. See "Known gaps" below before treating this as production-ready.
A full-codebase audit on 2026-07-26 found 24 bugs — five of them critical, including a dropped Electrum connection that hung the whole server with no reconnect, an RBF fee bump that wedged a round forever, and no way for the system to recover a broadcast that never confirmed (funds frozen). All 24 are fixed; BUGS.md is the record, with each one's root cause, what was actually done, and where its regression test lives. Read it before assuming any behaviour here predates those fixes.
Before writing code, always read the "Architecture" section below in full, plus the diagrams in flowchart/: platform-overview.mmd for the whole 5-phase flow, and round-lifecycle.mmd for the round/draw lifecycle in detail. Every node in these diagrams corresponds to a behavior that must be implemented exactly as described, including the labels on the edges (conditions, retries, loops). Regenerate their companion PDFs with flowchart/render-pdf.sh <file>.mmd after editing either one.
Human-facing guides live in docs/ (Italian, per explicit request — an exception to this file's English-only rule below): setup.md, running-the-server.md, guida-utente.md, guida-admin.md.
Commands
The server itself — in development and in production alike — always runs via Docker (see "Deployment" below); there is no supported way to run uvicorn directly against this codebase. The venv (.venv/) is only for local tooling: running tests, authoring Alembic migrations, and running the one-time scripts that generate the secrets/key material that end up referenced from .env.
source .venv/bin/activate # venv already created at .venv/
pip install -e ".[dev]" # install/update deps
alembic revision --autogenerate -m "message" # generate a new migration after editing app/db/models.py (applied automatically by the container's startup command — see Deployment — never run `alembic upgrade head` manually)
PYTHONPATH=. python scripts/generate_master_key.py # one-time: create+encrypt the server's master xprv (requires XPRV_ENCRYPTION_KEY in .env; see Deployment for where MASTER_KEY_PATH should point)
PYTHONPATH=. python scripts/decrypt_master_key.py # ops recovery: decrypt+print the existing master xprv (asks for confirmation first)
PYTHONPATH=. python scripts/encrypt_master_key.py # ops bootstrap: bring your own externally-generated xprv instead of generating one (getpass prompt, --overwrite to replace)
python -m pytest # run all tests
python -m pytest tests/unit/test_hd.py # run one test file
python -m pytest tests/unit/test_hd.py::test_derivation_is_deterministic # run a single test
.env (gitignored) holds real secrets for local dev; .env.example documents the required keys and how to generate them.
Deployment (Docker + Caddy)
The app is always run via Docker — dev and prod alike use the same docker-compose.yml, just with a different SITE_ADDRESS (see below); there's no separate dev-mode compose file or bare-uvicorn workflow. docker-compose.yml runs two containers: app (this codebase, built by Dockerfile, runs alembic upgrade head then uvicorn) and caddy (reverse proxy + automatic TLS). .env holds the app secrets; docker-compose.yml overrides DATABASE_URL/MASTER_KEY_PATH inside the container to point at the bind-mounted ./data/ (db, encrypted master key, logs — all gitignored, persist across container restarts). Set MASTER_KEY_PATH in .env itself to the host-side equivalent, ./data/keys/master.xprv.enc, so the venv-run key-generation scripts above (see "Commands") write to the exact same file the container reads — one source of truth for the key, whichever way it was generated.
mkdir -p data/db data/keys data/logs # one-time: host dirs bind-mounted into the app container
# one-time: generate the master key via the venv script above (scripts/generate_master_key.py),
# not via `docker compose run` — MASTER_KEY_PATH in .env already points at ./data/keys/
docker compose up -d --build # build + start app and caddy — same command for dev and prod
docker compose logs -f app # tail app logs (also written to ./data/logs/app.log)
docker compose down # stop
Caddy's site address comes from SITE_ADDRESS (env var on the host, read by docker-compose.yml):
- Dev, no domain: leave it unset (defaults to
localhost). Caddy detects it isn't a public hostname and issues a self-signed cert from its own internal CA — browsers will warn on first visit, expected for local testing (curl -kor click through). - Production, with a domain:
SITE_ADDRESS=lottery.example.com docker compose up -d(DNS must already point at the server, ports 80+443 reachable). Caddy automatically requests and renews a real Let's Encrypt certificate — no other config needed.
Known risk: docker-compose.yml sets restart: unless-stopped on app, so a crash mid-round auto-restarts the container — which hits the scheduler-resume gap below (a round stuck in closing/drawing/paying_out at restart stays stuck). Don't treat this as unattended-safe until that gap is closed.
Tech stack (MVP)
- Backend language: Python.
- PLM node access: Electrum protocol only (no full node/P2P). Bootstrap server for development:
santantonio.sytes.net:50002(SSL). - Auth: Argon2 password hashing + JWT sessions.
- Secrets: master xprv encrypted at rest with a symmetric scheme (AES-GCM/Fernet); the encryption key itself lives in an env var, never in the DB or in git.
- Operational config: every business/round parameter (fee address, bet amount, round duration, round cooldown, draw animation duration, minimum amount, network fee rate, RBF timeout) lives in the
round_configDB table (single row,app/rounds/config.py) and is only editable live via the admin dashboard (/admin) or its API — no env var involved at all, no redeploy or restart needed. Defaults for a brand-new instance are hardcoded column defaults on theRoundConfigmodel (app/db/models.py), notapp/config.py. Secrets and infra wiring (master key, JWT secret, Electrum host, admin token, database URL) stay env-var-driven in.envsince those genuinely need a restart. - Round cooldown:
round_cooldown_seconds— gap after a round closes before the next one opens, so players have time to see the outcome (default 30s). Not in the original flowchart; added afterwards as an explicit design decision. - Maintenance pause:
RoundConfig.paused(defaultfalse), toggled viaPOST /admin/pause/POST /admin/resume(a dedicated "Manutenzione" card in/admin's Parametri section, not a plain config field — it's a deliberate operator action, audit-logged aslottery_paused/lottery_resumed). When set,rounds/service.py:open_new_round_if_neededstops opening a next round once the current one closes — it never interrupts a round already in progress (that one still closes, draws, and pays out its winner normally).GET /rounds/currentexposes it aslottery_pausedso the user-facing page (/) shows a maintenance banner.
PLM network parameters
Source of truth: PalladiumWallet repo, ChainProfiles.cs and PalladiumNetworks.cs — always re-check that repo if a value is needed that isn't listed here, rather than guessing.
Mainnet:
- BIP44/84 coin type:
746(i.e. HD pathm/84'/746'/0'/0/index) - Bech32 HRP:
plm - P2PKH address version byte:
55(addresses start withP) - P2SH address version byte:
5 - WIF prefix:
0x80 - Block time: 120s
- BIP32 extended key headers (Legacy/native-segwit
zprv/zpubetc.): seeExtKeyHeadersinChainProfiles.cs
Electrum connection (rotation, keepalive, timeouts)
One connection serves everything — deposit credits, broadcasts, confirmations, the chain tip the draw waits on — which makes it the platform's biggest single point of failure. Three things keep it honest:
- Server rotation.
ELECTRUM_HOST/ELECTRUM_PORTis the primary;ELECTRUM_FALLBACK_SERVERSis a comma-separated list ofhost:port[:notls]extras (parsed byelectrum/client.py:parse_endpoints, which rejects malformed entries at startup rather than during the outage when the fallback is needed). The listener tries the next server after any failed or dropped session, and only sleeps on the backoff once every server has had a turn — so one dead server costs a single attempt, not an outage. - Every request is bounded (
_REQUEST_TIMEOUT_SECONDS, 15s) and a timeout tears the connection down. Unbounded waits used to hang aPOST /betswhile holding the per-user lock, and could stop the confirmation poller permanently. - The drop is observable.
client.wait_closed()resolves when the read loop dies, andlistener._run_onceraces it against the notification consumers and a 60sserver.pingkeepalive. Without this the listener sat on queues nobody would ever fill again and never reconnected — whilelistener.clientstill looked alive to everything else.
Balance display
place_bet/request_withdrawal (app/bets/service.py, app/withdrawals/service.py) select whole UTXOs to cover the amount (select_utxos, largest-first) and mark every selected UTXO spent_txid immediately at broadcast time — well before the tx has any confirmations. User.cached_balance_sats (recompute_balance, app/wallet/balance.py) only sums confirmed, unspent UTXOs, so right after a bet/withdrawal it understates the user's real balance by the entire unconfirmed change amount, which is often far larger than the amount actually moving.
compute_pending_balance (app/wallet/balance.py) fixes the displayed number without touching what's actually spendable: it decodes the raw tx of every in-flight (status="pending") bet/withdrawal PendingTransaction belonging to the user and sums whichever outputs pay back to the user's own address, adding that to cached_balance_sats. GET /users/me returns both balance_sats (confirmed-only — still what withdrawal-max and internal spend logic use, since only confirmed UTXOs are actually spendable) and pending_balance_sats + has_pending (what the frontend displays, colored green when settled and amber while has_pending is true).
Real-time updates (SSE)
GET /rounds/stream (app/api/routes/rounds.py) is a Server-Sent Events channel layered on top of the original polling loops in app/static/index.html/admin.html — polling is the fallback, not replaced, so a blocked/dropped SSE connection just degrades to the pre-existing behavior. The channel carries no payload and needs no auth: it's purely a "something changed, go refetch" ping; personalization (e.g. user_played below) still lives entirely in the normal per-user REST endpoints.
app/rounds/events.py's RoundEventBroadcaster (module-level singleton broadcaster) is a simple in-process pub/sub — one asyncio.Queue (maxsize 1, so redundant notifications coalesce) per connected SSE client. broadcaster.publish() is called from every point that changes something a dashboard would want to know about: a new round opening (rounds/service.py), every round status transition (rounds/scheduler.py: closing/drawing/paying_out/closed), a bet or withdrawal broadcast (bets/service.py, withdrawals/service.py), any pending tx confirming — bet/withdrawal/payout (tx/confirmation.py), a deposit credited (deposits/service.py), and a new block tip arriving (electrum/listener.py — the exact moment the "drawing" phase is waiting on).
Deliberate scope decisions, not oversights:
- Single-process only, no cross-worker fan-out. Fine for the current deployment (one uvicorn process, see
docker-compose.yml). A multi-worker/multi-container deployment would need a shared channel (e.g. Redis pub/sub) instead — don't add that speculatively before it's actually needed. - Generic broadcast, not a per-user channel. Every connected client refetches on every event, even ones irrelevant to them. Acceptable at the expected scale (~100 concurrent users); a targeted per-user channel would need auth on the SSE endpoint and server-side knowledge of who's affected by each event — real engineering work, only worth it well past current expected concurrency.
MAX_SUBSCRIBERS(default 500,app/rounds/events.py) is a defensive cap only — past it,GET /rounds/streamreturns 503 instead of opening a stream, and the client'sEventSourcejust falls back to polling. Not a substitute for the app-wide "no rate limiting anywhere" gap (see Known gaps).
Frontend: both index.html and admin.html open an EventSource('/rounds/stream') and, on an update message or on open (which fires on the initial connection and every automatic reconnect), immediately re-run the same refresh calls polling would eventually do — this matters most right after a dropped connection reconnects, closing most of the "missed while disconnected" gap.
MVP business parameters
- Bet cost per round: 10 PLM by default, admin-configurable (
RoundConfig.bet_amount_sats) — not a fixed constant. - Prize split: 70% winner / 30% fees, hardcoded in
rounds/scheduler.py(winner_share = pool_amount_sats * 70 // 100) — unlike bet amount, this ratio is not inRoundConfigand would need a code change, not an admin-panel edit. - Minimum withdrawal amount: equal to the current bet amount (
RoundConfig.bet_amount_sats), enforced inapp/withdrawals/service.py— not a separate admin-configurable field. Deposits have no server-side minimum check. - Confirmations required for all tx types (deposit, bet, payout, withdrawal): 1, hardcoded in
tx/confirmation.py— not configurable, per the design decision below.
What is PLM Lottery
A periodic-round lottery system built on a Bitcoin-like coin (PLM, mainnet). Each user gets a dedicated P2WPKH address (server-side HD wallet); they deposit PLM to that address, place a fixed-cost bet to enter the current round, and when the round closes a winner is drawn who receives 70% of the prize pool (the remaining 30% goes to fees).
Architecture (from the flowchart subgraphs)
The flow is organized into 5 phases (see flowchart/platform-overview.mmd for the full-platform diagram, and flowchart/round-lifecycle.mmd for the round/draw phase in detail):
- REG (Registration): on signup the server derives a new P2WPKH address via BIP84 (
m/84'/coin'/0'/0/index, one index per user) from a master xprv encrypted at rest. This address is permanent and serves as both the deposit address and the address that receives winnings and withdrawals. - DEP (Balance top-up): an ElectrumClient/SPV subscribes to the user's address scripthash. Internal balance (DB) is credited after 1 confirmation only — the reorg risk at 1-conf is knowingly accepted in v1, with no rollback logic.
- PLAY (Bet): fixed cost per round, at most one active bet per user at a time in v1. The server builds a PSBT user-address → pool-address for the fixed amount, with a change output back to the same user address (the user's balance must never exactly equal the bet amount). Fee minimized (~1 sat/vB), deducted from the bet amount. If the tx doesn't confirm within a timeout, fee-bump (RBF) and rebroadcast.
- DRAW (Periodic draw): configurable timer (default 10 minutes). The round's own deadline (
opened_at + round_duration_seconds) is the authoritative "yellow light" cutoff for new bets — not the DB status transition.place_bet(app/bets/service.py) callsrounds/service.round_accepts_bets(round_, round_duration_seconds), which rejects the bet once the deadline has passed even ifstatusis still"open"in the DB (theRoundSchedulertick that flips it to"closing"runs every_TICK_INTERVAL_SECONDS= 5s and can lag a few seconds behind the deadline). This closes the race where a bet placed in that lag window would otherwise still be accepted. Once a round leavesopen(closing/drawing/paying_out), no new bets are accepted for it either, and a new round can't open until the current one is fullyclosed(see round cooldown below). Round closing waits for all already-broadcast bets to confirm before proceeding (avoids losing bets at the round boundary) — this is the "yellow light" behavior: no new entries once the timer hits zero, but bets already in flight are still given time to confirm before the round actually closes and draws. The next round only opens once the previous round's payout tx is confirmed — rounds never overlap in v1. v1 draw algorithm (deliberately simple, meant to be replaced later): wait for the first block confirmed after round closing, use its hash as seed,index = seed mod participant_countover the participant list ordered by broadcast timestamp (this is also the tie-break when two bets confirm in the same block). Every participant has equal probability regardless of bet amount (consistent with the fixed bet amount). The payout (70% winner / 30% fees) is signed with the pool address key; the payout fee is deducted from the winner's 70%, the 30% fee share stays intact. Same timeout → RBF → rebroadcast pattern here too. The frontend shows a generic "drawing" status box (phase label, e.g. "Pagamento al vincitore in corso…") to every viewer on every dashboard for the whole closing/drawing/paying_out phase — this one is purely cosmetic status text, driven directly bystatus, no gating. Independently and additively (not instead of it), a personalized "Hai vinto!/Non hai vinto" box appears only for users whereGET /rounds/current'suser_playedfield is true (computed viaapp/auth/dependencies.py:get_optional_user, since this endpoint is reachable logged-out too) — everyone else has nothing to reveal and never sees it. That reveal is additionally delayed by at leastdraw_animation_seconds(admin-configurable, default 20s) for cosmetic suspense, anchored to the round's server-providedcloses_attimestamp rather than a client-side "first seen" time (so reloading the page can't reset the countdown), and decoupled from the real (and much longer, ~block-time) wait forwinner_user_idto actually be set. Once revealed, the result is persisted in the browser'slocalStorage(plm_persisted_result) so it survives a page refresh even after the round moves pastpaying_outintoclosed— at which pointget_active_roundstops returning that round at all andwinner_user_iddisappears fromGET /rounds/currententirely.GET /users/me/last-round-result(app/api/routes/users.py) is a durable, DB-backed backstop for a user who reloads on a browser/device that missed the live reveal window completely: it looks up the most recent closed round the user has aRoundParticipantrow in. Seeapp/static/index.html'srefreshRound/checkLastRoundResultfor the full reveal logic. - WITHDRAW (Withdrawal): the only way to move funds out of the platform to an external address. PSBT user-address → external-address + change back to the user address, fee deducted from the withdrawn amount, same RBF retry pattern.
PLAY and WITHDRAW share a per-user DB lock: a user can never have a bet-build and a withdrawal-build in flight at the same time, since both would otherwise spend from the same UTXO set on the user's dedicated address.
Three separate on-chain confirmations, not one, between the timer hitting zero and the payout landing — a common point of confusion, worth spelling out explicitly:
- Last bet's confirmation (
scheduler.py's_tick, thepending_countcheck before_close_and_draw) — the round doesn't even flip to"closing"until every already-broadcast bet has its 1st confirmation. This can already have happened before the timer expired; it's the earliest of the three and not necessarily tied to the deadline at all. - The draw block (
_wait_for_next_block, waits fortip_height > tip_at_close, wheretip_at_closeis recorded only once step 1 is done) — by construction this must be a later, different block than whichever one confirmed the last bet in step 1. - Payout confirmation —
_trigger_payoutbroadcasts only after step 2's block is known, then registers aPendingTransaction(kind="payout")that the same genericConfirmationPoller(app/tx/confirmation.py) waits on independently — this needs yet another, later block than step 2's, since the payout can't be built before the winner is known.
So worst case (last bet confirms right at the deadline) is ~3 block times end-to-end; best case (all bets already confirmed before the timer hit zero) is ~2 (draw block + payout block). At PLM's 120s block time that's roughly 4–6 minutes worst case, 2–4 minutes best case — independent of draw_animation_seconds, which only sets a cosmetic minimum for the frontend animation.
Internationalization (user-facing page only)
app/static/i18n.js holds every user-facing string of / in 7 languages (en, it, es, fr, de, ru, zh) as one flat TRANSLATIONS table — no build step, no fetch, loaded before app.js so t() is available everywhere. Language comes from localStorage.plm_lang, falling back to navigator.language, falling back to en; the switcher lives in the chain-bar, not the navbar, deliberately — the navbar is hidden until login, which would leave the landing page and the login form untranslatable for exactly the users who need the switch.
- Static markup is translated by attribute (
data-i18n, plus-html,-placeholder,-title,-aria-label,-alt), applied byapplyStaticTranslations(root?)onDOMContentLoadedand on every switch. Anything rendered from server data is built witht()inapp.jsinstead, and re-rendered byonLanguageChange()— an element must be in one camp or the other, never both, or the two mechanisms overwrite each other (this is why#bet-btnhas nodata-i18n: its label carries the admin-configurable bet amount, sorenderBetButton()owns it). - Every language must have exactly the same key set. There is no fallback beyond
en, and a missing key renders as the raw key string. /adminis intentionally not translated (operator-facing, Italian only), and neither is/guida(servesdocs/guida-utente.md).
API error contract (app/api/errors.py): the API is single-language by design. User-facing failures answer with a structured detail — {"code", "message", "params"} — where message is English for non-dashboard consumers and code is what the frontend maps onto error.<code> in i18n.js (falling back to message for an unknown code). Domain exceptions (BetError, WithdrawalError) subclass ApiError and carry the code from where the failure actually happens; str(exc) is still the English message. When adding a user-facing error: give it a code, add error.<code> to all 7 languages, and pass interpolated values through params (amounts as *_sats — the frontend derives a *_plm sibling automatically) rather than baking them into the English text.
Admin dashboard and test UI
Two static single-page apps, served directly by FastAPI (app/main.py mounts app/static/ and adds a dedicated GET /admin route) — no build step, no framework. Each page's HTML/CSS/JS are separate files (index.html/style.css/app.js, admin.html/admin.css/admin.js), served as plain static files (no bundler):
/(app/static/index.html): the end-user test UI. Register/login, then a menu-driven dashboard (Deposito with a QR code of the address viaGET /qr/{address}, Bet, Prelievo) with a persistent round-status card (GET /rounds/current: id/status/timer/participant count/jackpot) above the menu./admin(app/static/admin.html): gated by a token screen (not a real login — just checksX-Admin-TokenagainstADMIN_TOKENfrom.env), then a navbar-driven dashboard with five sections, each backed by its own/admin/*endpoint (app/api/routes/admin.py): Parametri (RoundConfigCRUD), Utenti (list + per-user WIF privkey export, audit-logged), Round (history), Transazioni pendenti (in-flight RBF candidates), Audit log./adminis deliberately not linked from/in either direction — reachable only by knowing the URL.
Both pages talk to the same JSON API everything else uses; there's no separate "admin API" vs "user API" boundary beyond the require_admin dependency.
Non-obvious domain decisions
These choices were made explicitly during design (not derivable from reading a single file) and must be respected in any implementation:
- Private keys (xprv) are generated and held server-side — this is not a non-custodial system: the user never controls their own keys until they make an explicit withdrawal.
- The user's personal deposit address always doubles as the winnings-receiving address: there is no separate "winner address".
- 1 confirmation is the chosen threshold for all tx types (deposits, bets, payouts, withdrawals): don't introduce different thresholds (e.g. 3 or 6 confirmations) without an explicit decision.
- The draw algorithm (node R) is deliberately simple and should be treated as a replaceable/pluggable component, not the final design — don't architect around its current implementation.
- The admin panel can export any user's raw WIF private key (
GET /admin/users/{id}/privkey,app/wallet/hd.py:derive_user_wif). This is intentional, not a vulnerability to fix: the server already holds the master key everything derives from (custodial by design, see above), so this only exposes through the API something an operator could already do via a script. Every access is written toaudit_log(admin_privkey_accessed) — don't remove that logging when touching this endpoint. - RBF fee bumps are paid by whoever's change output the tx pays back to — the user for bets/withdrawals, the pool for payouts — never by the fixed counterparty amount (recipient/winner/fee-address outputs are untouched; only the sender's own change shrinks). See
bump_feeinapp/tx/broadcast.py.
Transaction reconciliation and the tx lifecycle
Everything that spends money is written before it is broadcast, and resolved against the chain afterwards. This is what makes the system recover on its own instead of needing manual DB edits (BUGS.md B-04/B-08).
PendingTransaction.status is the lifecycle: building → pending → confirmed, or
failed.
buildingis written first, with the UTXOs already markedspent_txid, and committed before the broadcast (bets/service.py:place_bet,withdrawals/service.py). A crash in that window therefore leaves evidence rather than coins spent on-chain with no record.- If the broadcast is refused, the service releases the reserved UTXOs, restores the
balance, removes the participant (or marks the withdrawal
failed), audit-logs it, and raisesbroadcast_failed— answered as 502, since the network refused it, not the caller. app/tx/reconcile.py(PendingTransactionReconciler, every 120s and once at startup) asks the chain about anything stillbuilding/pending. Tx present → promote; tx gone → markfailedwith afailure_reason, release the inputs, roll the domain row back, audit-logpending_tx_abandoned. Grace periods differ by state (120s forbuilding, 6h forpending, so the RBF bumper gets its attempts first), and a transport failure never abandons anything — only a server that positively doesn't know the tx does.
Because of this, UtxoEvent.spent_txid must always equal the current txid of the tx
reserving it: bump_fee retargets it (along with RoundParticipant.bet_txid,
Withdrawal.txid and Round.payout_txid) on every fee bump. Confirmation handlers
deliberately key off immutable ids (round_id/user_id, withdrawal_id) rather than the
txid, which changes under them.
At most one active round is a database invariant, not just a code convention:
ix_rounds_single_active (a unique index over the constant expression (1), restricted to
the active statuses) makes a concurrent second insert fail cleanly, and
open_new_round_if_needed recovers by using the winner's round.
Known gaps / TODO
Not blockers for reading the code, but must be addressed before this is production-ready. The 24 findings of the 2026-07-26 full-codebase audit are all fixed — see BUGS.md, which keeps each one's root cause, fix and regression test as the record. What remains open:
- Scheduler doesn't resume mid-flight rounds after a restart.
rounds/scheduler.py's_tick()only acts on rounds withstatus == "open". If the process restarts while a round isclosing/drawing/paying_out, it's permanently stuck — nothing re-enters_wait_for_next_blockor retries_trigger_payout. Needs a startup routine that inspects in-progress rounds and resumes (or a periodic "unstick" check) before this can run unattended. Note this is round-level state: in-flight transactions do now recover on their own (see "Transaction reconciliation" below). - RBF bump only handles one case: a single change output, paying back to the tx's own sender address, large enough to absorb the fee increase. No additional-input selection fallback — an exact-amount tx (no change) or a change output too small to absorb the bump raises
RbfError. The consequence is no longer permanent, though: a tx that can't be bumped and never confirms is eventually abandoned and its UTXOs released (see "Transaction reconciliation"), so the funds come back instead of being frozen. - Payout retry: if
_trigger_payoutfails (insufficient pool UTXOs, a badfee_address, Electrum disconnected), it logs, writes apayout_failedaudit entry, and returns — the round stays inpaying_outwith no automatic retry. The audit entry makes it visible in/admin; acting on it is still manual. - Withdrawal and RBF bump have never been exercised against a live broadcast — only deposit and bet flow are verified end-to-end with real PLM. Both paths have unit coverage, including their failure and rollback branches, but unit tests are not a live network.
- No general user-facing history endpoints (list my own bets / withdrawals / past rounds) —
GET /users/me/last-round-resultcovers exactly one case (the outcome of the most recent closed round the user played in, as a reveal-persistence backstop; see DRAW above), not a real history. The admin side has more (/admin/rounds,/admin/pending-transactions,/admin/audit-log), but there's still no "my own full history" equivalent for a logged-in user. A failed withdrawal now leaves astatus="failed"row the user cannot see anywhere — an argument for closing this gap. - Admin auth is a single shared bearer token (
ADMIN_TOKEN,X-Admin-Tokenheader) — no per-admin identity:audit_logrecords what changed (config edits are now logged too, asconfig_updated, with before/after values) but never which operator did it. This token gates the user list, private key export and round/audit history, so its blast radius if leaked is large. - No rate limiting / abuse protection on any endpoint (register, bet, withdrawal, admin).
/guidais not served in Docker.GET /guidareadsdocs/guida-utente.md, and theDockerfiledeliberately does notCOPY docs— the guide is pending a rewrite, so it isn't shipped yet. The endpoint answers a clean 404 (guide_unavailable, translated) and logs an error rather than crashing, but the navbar help link leads nowhere untilCOPY docs ./docsis added back.- No automated integration tests against a live Electrum connection — all live-network verification so far has been manual (ad hoc scripts + real mainnet transactions), not part of the
pytestsuite. - Single-process assumptions: the SSE broadcaster (
rounds/events.py) and the per-user locks (tx/locks.py) are both in-process only. Fine for the current one-uvicorn-process deployment; a multi-worker one needs a shared channel and a DB/Redis lock. Note the round-uniqueness invariant is not in this category any more — it's enforced by a DB index (see below).