Multi-Tenant Isolation Across Two Production Apps by Joshua BrownMulti-Tenant Isolation Across Two Production Apps by Joshua Brown

Multi-Tenant Isolation Across Two Production Apps

Joshua Brown

Joshua Brown

The problem
Two production apps, an office manager and a field app, each holding several contractors' data, joined by a shared event spine across two separate databases. Every table had a tenant id column. That is not isolation. That is a convention you are hoping every query remembers forever, including the ones you write at eleven at night.
What I found before writing a single policy
Both apps were connecting to PostgreSQL as the superuser, with every table owned by that same role. Postgres silently ignores row-level security for a table's owner and for any superuser. So every policy anybody wrote would have been decoration. Nothing would have errored. Nothing would have warned. The table would have reported RLS enabled and the data would have been wide open to every tenant.
That is the shape of this whole job. The dangerous version of a security bug is not the one that crashes.
How I did it
Roles first. Both apps moved off the superuser onto dedicated NOSUPERUSER, NOBYPASSRLS roles that own the schema, so boot-time migrations still work and policies actually bind.
Tenant context second. An AsyncLocalStorage scope carries the tenant through the request and sets it on the Postgres session, so a policy has something real to read.
Instrument before enforcing. Before arming a single policy, every query that ran with no tenant context got logged. Then I went and fixed those, which turned out to be crons, public token routes, and webhook handlers. Enforcement came last, with FORCE plus both USING and WITH CHECK clauses across 45 tables.
The bug that proved the method
The first time I armed the policies, everything broke, and the cause was not the apps. The connection pool wrapper had a branch that amounted to, if this is not a SQL string, pass it straight through. Drizzle hands the pool a config object, not a string. So every single ORM query took the pass-through branch and skipped the tenant transaction entirely. The tenant never reached Postgres at all.
Here is the part worth keeping. Postgres answers a policy evaluated against a null tenant with zero rows, not an error. Reads came back empty and writes threw, but nothing anywhere said why. Silence, not a stack trace.
The fix was the wrapper. The lesson was the test. An empty audit log is not evidence that every query is scoped, because it is equally consistent with the audit never having run. So the health check now asks Postgres directly what tenant it currently sees and reports the answer. A verdict, not an absence.
Stack
PostgreSQL row-level security, Node.js, Drizzle ORM, AsyncLocalStorage, multi-tenant SaaS architecture, Railway.
What this means if you are hiring
If you run multi-tenant SaaS and your isolation is a WHERE clause every developer is trusted to remember, you do not have isolation, you have a good track record. Those are different things and only one of them survives a new hire. I can walk your schema and your query paths and tell you exactly where the gaps are before a customer finds one for you.
Like this project

Posted Sep 24, 2026

Row-level security across two multi-tenant apps and 45 tables. Instrument first, enforce second, and never trust an empty audit log as evidence.