We Audited Our Own Monetization Before Store Submission. It Did Not Go Well — At First.
Ten days before submitting NihonGO! to the app stores, we did the thing everyone should do and almost nobody does: threat-model our own money paths. The findings were humbling enough that we wrote them down in full.
The findings
1. Subscriptions could be self-granted. The upgrade endpoint trusted a client-sent tier, and receipt verification returned verified: true for any input. Fix: the server now derives the tier from a provider-verified receipt; the client sends a purchase token, never a tier. Expiration actually runs on a schedule and is exercised in tests.
2. CORS allowed any subdomain containing .supabase.co, with credentials. Substring matching on the origin header is not a security boundary. Fix: explicit allowlist.
3. The most expensive endpoint had no auth. Speech-to-text — our highest per-call cost — accepted unauthenticated requests with no rate limit. Fix: auth required, per-user rate limits, usage logging with alerts on anomalies.
4. A missing Redis error listener could crash the backend. Worse, cached entitlement paths could fail open. Fix: error listeners, circuit-breaker degradation to the database, and integration tests that kill Redis on purpose.
5. The readiness probe returned 503 forever because it queried a table that did not exist (profiles vs user_profiles). The probe had never once passed; nobody noticed because nothing depended on it.
6. An AI SDK minor version renamed maxTokens to maxOutputTokens. The build stayed green and model spend became unbounded. Fix: pinned versions, explicit wrapper-level limits, and a per-request budget alert.
7. Subscriptions never expired because the cron job was never wired. Every upgrade was effectively a lifetime license.
8. Streaks computed in UTC reset at 9 AM for Japanese learners. Fix: compute in the user timezone at write time and store the result, not the rule.
9. A contrast helper looked like accessibility but its pow() was this * this, so the WCAG check was meaningless. Fix: correct relative-luminance implementation plus tests over known contrast pairs.
10. CI was green for the wrong reasons. Skipped tests, broken mocks, and one suite returning 401 for every case due to a missing auth mock. Unmasking surfaced 141 analyzer issues and 32 real failures.
The order matters
We audited in this order: money → auth → reliability → honesty of the suite. Almost all the serious findings were in the first category. If your time is limited, start where the revenue is — then work backwards through what protects it.
What changing each finding looked like
Every fix shipped with a test that fails if the bug returns. Not a comment, not a checklist item — an assertion. A few examples:
// before: client sent the tier
await upgrade({ tier: "pro" })
// after: client sends a receipt, server derives the tier
await upgrade({ purchaseToken })
The fail-open cache decision was the one worth arguing about. The naive fix for a crashing Redis client is to swallow errors — which converts an availability bug into an authorization bug. Degrading to database reads is slower and correct.
The numbers after 10 days
- Backend tests: 263 passing (from 25 failing)
- Flutter tests: 186 passing (from 7 failing)
- Analyzer issues: 0 (from 141)
- Revenue paths: 4/4 server-verified and regression-tested
A checklist you can steal
- Can a client grant itself paid access? (Test it with curl, not the UI.)
- Which endpoint costs the most per call? Does it require auth and a rate limit?
- Kill your cache and your database — does anything fail open?
- When did your test suite last fail? If you can't remember, turn skips off and find out.
- Does your readiness probe actually pass?
- Does anything user-visible depend on UTC assumptions?
- Is any accessibility or security claim enforced by a test, or just by a comment?
Green CI is a hypothesis, not evidence. The audit does not end when the tests pass — it ends when you can explain why every red one was red.