Fix the round-open race, the advertised jackpot, and payout error handling

Round opening (B-09). open_new_round_if_needed now handles the IntegrityError
from ix_rounds_single_active (previous commit) by rolling back and using the
winner's round. Deviation from the plan in BUGS.md, which proposed making the
scheduler the only writer: that would mean the first bet after a cooldown
couldn't open a round, so both callers stay and a bounded retry was added
instead — a conflict where nothing is active yet just means the winner hadn't
committed, and a bet must not fail on that timing. get_active_round also logs
loudly if it ever sees more than one active round rather than silently picking
the newest.

The jackpot (B-11). It was participant_count * the *current* bet_amount_sats,
which overstated the pool (each stored bet is already net of that bet's network
fee) and silently rewrote the advertised jackpot of a round in progress whenever
an operator edited the bet amount. It now sums the participants' stored
bet_amount_sats. The remaining imprecision — the payout tx's own fee, deducted
from the winner's share and unknowable until the payout is built — is documented
in the code rather than promised away, since the comment there claimed exactness.

Payout (B-05, B-18). _trigger_payout is split into read / build+broadcast /
persist, so no DB session is held across a network call (on SQLite that meant
holding the write lock for two unbounded round-trips). That restructuring is also
what makes the error handling placeable: it now catches Exception around the
chain work and writes a payout_failed audit entry, where a malformed fee_address
used to raise EmbitError all the way to the scheduler's catch-all, leaving the
round stuck in paying_out with nothing recorded about why. Automatic payout retry
remains an open gap.

The scheduler also counts "building" participants as in-flight when deciding
whether a round may close, matching the two-phase bet write.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-27 00:32:22 +02:00
co-authored by Claude Opus 5
parent b4d70385a6
commit daf66fd6bc
5 changed files with 279 additions and 44 deletions
+55 -11
View File
@@ -1,20 +1,43 @@
import logging
from datetime import datetime, timedelta, timezone
from sqlalchemy import select
from sqlalchemy.exc import IntegrityError
from sqlalchemy.ext.asyncio import AsyncSession
from app.db.models import Round
from app.rounds.config import get_round_config
from app.rounds.events import broadcaster
logger = logging.getLogger(__name__)
_ACTIVE_STATUSES = ("open", "closing", "drawing", "paying_out")
# Bounded: a conflict means someone else is opening a round right now, so a couple
# of retries is plenty. Unbounded retries could spin if the invariant were ever
# broken in a way we don't anticipate.
_OPEN_ROUND_ATTEMPTS = 3
async def get_active_round(session: AsyncSession) -> Round | None:
"""The round currently in progress (in any non-closed state), if any. Rounds
never overlap: a new round only opens once the previous one is fully closed
(payout confirmed, or no participants to pay out)."""
return await session.scalar(select(Round).where(Round.status.in_(_ACTIVE_STATUSES)).order_by(Round.id.desc()))
(payout confirmed, or no participants to pay out).
The database enforces "at most one active round" (ix_rounds_single_active, see
app/db/models.py), so the ordering below is belt-and-braces; if it ever does
see two, that's a broken invariant and worth a loud log rather than silently
picking one."""
active = (
await session.scalars(select(Round).where(Round.status.in_(_ACTIVE_STATUSES)).order_by(Round.id.desc()))
).all()
if len(active) > 1:
logger.error(
"invariant violated: %s rounds are active at once (ids=%s) — using the newest",
len(active),
[r.id for r in active],
)
return active[0] if active else None
def round_accepts_bets(round_: Round, round_duration_seconds: int) -> bool:
@@ -55,12 +78,33 @@ async def open_new_round_if_needed(session: AsyncSession) -> Round | None:
if datetime.now(timezone.utc) < closed_at + timedelta(seconds=config.round_cooldown_seconds):
return None
round_ = Round(status="open")
session.add(round_)
await session.flush()
# Published pre-commit (the caller commits right after) — acceptable: this
# only tells subscribers "go refetch", and by the time an SSE client's
# refetch request actually lands, this in-process commit (microseconds
# away) has essentially always already happened.
broadcaster.publish()
return round_
for attempt in range(_OPEN_ROUND_ATTEMPTS):
round_ = Round(status="open")
session.add(round_)
try:
await session.flush()
except IntegrityError:
# Another caller (the scheduler tick, or a concurrent place_bet) got
# there first — ix_rounds_single_active turns what used to be two live
# rounds into a clean failure here. Roll our insert back and use theirs.
# Safe to roll back: this runs before its callers have written anything
# else in this session.
await session.rollback()
existing = await get_active_round(session)
if existing is not None:
logger.info("lost the race to open a round; using round %s", existing.id)
return existing
# Nothing active *and* the insert conflicted: the winner's transaction
# hadn't committed yet when we looked. Try again rather than failing the
# caller — a bet shouldn't 500 because of a scheduler tick's timing.
logger.info("round-open conflict with nothing active yet (attempt %s), retrying", attempt + 1)
continue
# Published pre-commit (the caller commits right after) — acceptable: this
# only tells subscribers "go refetch", and by the time an SSE client's
# refetch request actually lands, this in-process commit (microseconds
# away) has essentially always already happened.
broadcaster.publish()
return round_
logger.error("could not open a round after %s attempts", _OPEN_ROUND_ATTEMPTS)
return await get_active_round(session)