Files
plm-lottery/BUGS.md
T
davide 97545ad91f Run SQLite in WAL mode with a busy_timeout (B-39)
create_async_engine had no connect_args and there was no PRAGMA
anywhere in the repo. SQLite's default rollback-journal mode lets a
writer block every reader for the duration of its transaction, and a
second writer arriving while one is already active fails immediately
with "database is locked" rather than waiting at all - realistic given
five concurrent background tasks (scheduler, confirmation poller, RBF
bumper, two reconcilers) plus every HTTP handler share one file, and
nothing previously handled that error.

app/db/base.py now registers a "connect" event on the engine that sets
journal_mode=WAL, synchronous=NORMAL and a 5-second busy_timeout on
every new DBAPI connection - applied only when the dialect is sqlite,
so a future PostgreSQL DATABASE_URL is unaffected. WAL lets readers and
writers proceed without blocking each other, and busy_timeout gives a
second writer a real window to wait instead of failing instantly.

Left out: an explicit application-level retry wrapper for "database is
locked" in the background loops, the other half of the proposed fix -
busy_timeout already gives SQLite itself several seconds to resolve
writer-vs-writer contention before ever raising, and every background
loop already catches and logs an unhandled exception before its next
scheduled tick, which is itself a retry, just not an immediate one.

Suite grows from 211 to 214 tests. BUGS.md moves B-39 to Previously
fixed.
2026-07-27 14:53:22 +02:00

148 lines
8.5 KiB
Markdown

# Known bugs
A second full-codebase audit on 2026-07-27 found **25 further issues** (4 critical, 6 high,
7 medium, 8 low), listed below as B-40 … B-49. B-25 through B-39 are fixed (see "Previously
fixed" below) — no Critical-severity finding remains open; the other 10 are Medium/Low.
The 139-test suite was green at the time of the audit, so none of these were caught by existing
coverage — every fix lands with a regression test (the fifteen fixes so far brought the suite
from 139 to 214).
The recurring pattern across the open findings is worth stating once: the code is rigorous
about the failure modes that have actually been hit, and silent about the ones that have not.
The payout phase is now fully recoverable; the "drawing" phase (waiting on a block) is now
observable (B-36) but still has no equivalent resume-after-restart — see "Known gaps / TODO"
in [CLAUDE.md](CLAUDE.md), which is also where other by-design limitations (single-shared-token
admin auth, single-process assumptions, no user-facing history, etc.) are documented.
---
## Medium
### B-40 — `bump_fee` holds a DB session open across N network calls
`tx/broadcast.py:80` issues one `get_transaction` **per input** (up to 15s each) and then a
`broadcast`, all with the session open. This is precisely the pattern B-18 removed from
`_trigger_payout` via its three-phase structure; it survives here.
Side note in the same function: `_prevout_amount` does `round(value_coins * 100_000_000)` on a
float from the server — acceptable at these magnitudes, but it is floating-point money
arithmetic in a codebase that is otherwise strictly integer-satoshi.
**Proposed fix.** Restructure into the same three phases: read what is needed and close the
session, do the chain work, then reopen to persist. For the float: prefer the raw (non-verbose)
transaction and parse the output value as an integer with `embit`, which is what
`reconcile.py:_release_inputs` already does for inputs.
### B-41 — All confirmation logic depends on `verbose=True`, which is not universally supported
`poll_once`, `reconcile._tx_exists_on_chain` and `bump_fee` all call
`blockchain.transaction.get(txid, True)`. Several Electrum server implementations and versions
reject the verbose flag ("verbose transactions are currently unsupported"). Falling back onto
such a server means **no confirmations, no reconciliation, no bumps** — and the code would read
that as a transport error and stay silent.
Related: `reconcile.py:83` decides whether to **abandon a transaction** by substring-matching
the error text (`"missing"`, `"not found"`, `"no such"`, `"unknown"`). It works against
ElectrumX; it is fragile as the basis for a decision that releases funds.
**Proposed fix.** Use `blockchain.transaction.get_merkle` (or the scripthash history) for
confirmation and existence checks — both are portable and give the confirming height directly.
Probe verbose support once at connect time and record it on the client, so an unsupported
server is detected loudly at session start rather than silently mid-operation.
---
## Low / hygiene
### B-42 — `/docs` exposed in production
FastAPI mounts Swagger by default, so the entire API surface — `/admin` included — is publicly
enumerable. The README advertises it.
**Fix:** `docs_url=None, redoc_url=None, openapi_url=None` in production (env-gated), or place
them behind `require_admin`.
### B-43 — No HTTP security headers
The [Caddyfile](Caddyfile) sets no CSP, no `X-Frame-Options`/`frame-ancestors`, and no HSTS
(Caddy does not add it on its own). The JWT lives in `localStorage`, so any XSS exfiltrates
it, and the page is iframeable.
**Fix:** a `header` block in the Caddyfile with `Strict-Transport-Security`,
`X-Content-Type-Options: nosniff`, `Referrer-Policy` and a CSP tight enough for two static
pages with no external assets (`default-src 'self'`).
### B-44 — README and CLAUDE.md contradict each other
The README says to run `uvicorn --reload` directly and
`docker compose run --rm app python scripts/generate_master_key.py`; CLAUDE.md says explicitly
that neither is supported. Whoever opens the repo reads the README first.
**Fix:** align the README's Quick start with the Docker-only workflow documented in
CLAUDE.md and `docs/setup.md`.
### B-45 — Unvalidated and unpaginated admin list endpoints
`limit: int = 50` on `/admin/rounds` and `/admin/audit-log` has no bounds (`-1` means
"everything" on SQLite), and `/admin/pending-transactions` has no limit at all — it grows
without end.
**Fix:** `Query(default=50, ge=1, le=500)` on both, and the same treatment plus a status filter
on the pending-transaction list.
### B-46 — `secrets.compare_digest` on a `str` raises on non-ASCII input
`api/routes/admin.py:27` raises `TypeError` — a 500 instead of a 403 — when the header contains
non-ASCII characters.
**Fix:** compare the UTF-8 encoded bytes of both sides.
### B-47 — Unbounded `String` columns for large text
`raw_tx_hex` (`db/models.py:146`) and `payload_json` (`:178`) should be `Text`. It works on
SQLite and PostgreSQL and breaks elsewhere.
**Fix:** switch both to `Text` in a migration.
### B-48 — No cap on input count in `select_utxos`
A user with hundreds of small UTXOs builds a huge transaction whose fee — deducted from the bet
amount — materially erodes their contribution to the pool, and it can exceed standardness
limits.
**Fix:** cap the selected inputs (e.g. 50) and fail with a translatable error suggesting a
consolidation, or consolidate the address automatically when the count crosses a threshold.
### B-49 — Rollback paths do not publish an SSE update
`bets/service.py:_release_failed_bet` and `withdrawals/service.py:_release_failed_withdrawal`
restore the balance without calling `broadcaster.publish()`, so dashboards only find out on
their next poll.
**Fix:** one `broadcaster.publish()` at the end of each, as every other state-changing path
already does.
---
## Previously fixed
- **B-25** — the payout had no two-phase write, unlike bets and withdrawals
- **B-26** — a payout failure or a process restart could wedge a round in `paying_out` forever
- **B-27** — an RBF bump reset the reconciler's own abandon clock, so a repeatedly-bumped tx was never abandoned
- **B-28** — a hostile Electrum server (or a MITM) could single-handedly pick the round's winner
- **B-29** — a UTXO absent from one server's `listunspent` was marked spent immediately, irreversibly, on a single unauthenticated reply
- **B-30** — a lost scripthash subscription meant a user's deposits were never credited, with no periodic safety net
- **B-31** — resubscribing on reconnect ran serially before anything else started, freezing the chain tip (and so an in-flight draw) for the whole sweep
- **B-32** — an RBF bump could retry forever below BIP125's relay-mandated minimum fee delta, with no ceiling on the fee rate either
- **B-33** — `POST /auth/login` had no rate limiting, so a password could be brute-forced against an enumerable username list
- **B-34** — password change/reset didn't invalidate already-issued JWTs, so a stolen token survived a change meant to lock it out
- **B-35** — API timestamps round-tripped as naive datetimes, so the frontend parsed them as local time instead of UTC
- **B-36** — a stalled draw wait had no timeout, no log, and no audit trail, so a frozen round showed nothing in `/admin`
- **B-37** — a withdrawal covered by unconfirmed change answered "insufficient balance" instead of distinguishing it from actually having no funds
- **B-38** — the SSE subscriber cap was global, so one client opening enough connections degraded every other user to polling
- **B-39** — SQLite ran without WAL or a `busy_timeout`, so a writer could block every reader and a second writer failed immediately instead of waiting
See git history for the fix-by-fix breakdown (commits `f13f685`, `50a43ae`, `933760e`, and the
B-28/B-29/B-30/B-31/B-32/B-33/B-34/B-35/B-36/B-37/B-38/B-39 fixes). Suite grew from 139 to 214 tests over the fifteen.
A full-codebase audit on 2026-07-26 (commit `d4e0974`) found 24 bugs across every Python
module under `app/`, both static frontends, and the Docker/Caddy deployment — 5 critical,
7 high, 7 medium, 5 low. All 24 were fixed and verified against the current code on
2026-07-27; the fixes are covered by the regression suite (grew from 79 to 139 tests) and
five of them were additionally confirmed against a real mainnet deployment (see git history
between `fb734bb` (documenting the findings) and `845ba98` (recording the audit outcome) for
the fix-by-fix breakdown — each commit message names the bugs it closes and where their
tests live).