# Security model

What is protected, how, and what is deliberately left to the deployment.

## Threats this design takes seriously

1. **One customer reading another's data.** The product holds several customers'
   market data and their provider credentials. This is the failure that would end
   the product.
2. **A leaked registry file.** If the file holding provider credentials is read,
   those credentials must still be useless.
3. **An error message that maps the service.** Stack traces and paths are the
   cheapest reconnaissance there is.
4. **One customer exhausting a shared resource.** Quota, memory, disk.

## Tenant isolation

Every customer's processing runs in its own workspace:

```
var/tenants/<client_id>/data/      dataset, state, manifest, locks, reports
var/tenants/<client_id>/live/      live payload window and snapshot
```

The mechanism, not the intention:

- **Identity is derived, never accepted.** The API resolves the tenant from the
  bearer token. A `client_id` in a query string, a body or a header is ignored.
  The function that builds a tenant context takes the authenticated client
  *record*, not an id, so a caller cannot pass one that came off the wire.
- **The id is narrow by construction.** Only `[A-Za-z0-9_-]`, maximum 64
  characters. A dot, a slash or a null byte is refused before any path is built,
  and the resolved workspace is checked to be inside the tenant root.
- **The engine's data root is a configuration value.** A per-customer cycle runs
  with its own root and its own working directory. A regression gate proves that
  the same input produces byte-identical derived output in any root — moving the
  root is isolation, not a change to market semantics.
- **Live processing runs in its own process.** The live pipeline keys state by
  the provider's fixture id and holds it in module memory, so two customers
  watching the same real match would otherwise share a series. A child process
  per customer, with its own working directory, makes isolation a property of the
  operating system rather than of a convention.
- **Locks are per tenant.** One customer's cycle never blocks another's. A lock
  whose holder is dead, or which has outlived its ceiling, is taken over — a
  crash is recoverable rather than permanent.
- **Only the owner releases a lock**, so a slow run cannot delete the lock of the
  run that took over from it.

Tested in both directions, every time: that A sees A's data, and that A cannot
see B's. Including the hardest case — a fixture **both** customers hold, where
the identifier is identical and only the workspace can distinguish them.

A fixture belonging to another account returns `FIXTURE_NOT_FOUND`, the same as
one that does not exist. No response confirms what another account holds.

## Credentials at rest

**Provider credentials** — AES-256-GCM, per-record nonce, authenticated, bound to
their own connection record by additional authenticated data. The master key
lives in the environment or a secret manager; `key_version` travels with each
record so a new key can be installed alongside the old one.

Not a platform keystore: development runs on Windows and production on Linux, and
a credential that cannot be restored on the production host is a credential we
have lost.

**Client tokens** — never stored. A salted scrypt hash is stored; comparison is
constant-time. The plaintext is returned exactly once, at creation or rotation.

**The master key** — never written to any file the product owns. A gate walks the
whole data directory after a run and fails if it finds it.

## Secret leakage

The leakage gate does not grep for the word `token` and declare failure — a
module that handles credentials must mention them, and a gate that cannot tell
storage from exposure is a gate that gets switched off.

Instead it creates a real credential and a real client token, drives every
surface that could carry them, and searches the output: every response body,
every log line, every stored file, the usage store, health payloads, and
exception messages. It then asserts the inverse: that the sealed credential
appears in the connection store **and nowhere else**, and the token hash in the
client registry **and nowhere else**.

Everything that reaches a log line passes through redaction first — by key name
and by pattern, recursively, including inside arrays.

Log lines carry the **route template**, never the real path: a path carries
identifiers, and an identifier plus a timestamp is a behavioural record of a
customer we have no reason to keep.

## Error responses

`INTERNAL_ERROR` and a request id. No stack trace, no path, no provider URL, no
upstream error body.

An unknown token and a token that never existed get the same answer. A suspended
or revoked account gets a specific one — the holder owns that token and needs to
know why it stopped working.

## Limits

| | |
|---|---|
| Request body | 16 KB public, 64 KB admin |
| Response | 8 MB ceiling; a route that would exceed it fails rather than sending |
| Page size | 50 default, 200 maximum |
| Rate | 120 requests/minute per account; 30/minute unauthenticated per address |
| Request timeout | 30 s |
| Provider timeout | 20 s |

All declared in one module. A gate fails if a route writes its own.

Rate limits are keyed by **account**: one of your servers cannot throttle
another, and adding addresses does not raise your ceiling.

`X-Forwarded-For` is honoured **only** when a trusted proxy is configured.
Trusting it unconditionally would let any caller choose their own rate-limit
bucket, which is the same as having no rate limit.

## Write durability

Every mutable store writes through a distinct temporary file, `fsync`, then
`rename`. On failure the temporary file is removed — otherwise a directory fills
with one file per failed flush, which is how a disk-space incident starts.

A crash leaves the previous file whole. It never leaves a half-written credential
store, which would read as "no connection" and silently stop a customer's
processing.

A read failure on a file that **exists** stops the cycle. A **missing** file is a
legal empty state. These are different: a file that cannot be read is not
evidence that it contains nothing.

## Startup integrity

The service refuses to start in production when:

- no provider master key is installed
- a provider adapter fails its contract
- no administrator password is set

In development these are warnings. The check reports severities, not one
boolean, so a warning cannot quietly become a blocker or the reverse.

## Deployment assumptions

**The application does not terminate TLS.** It binds loopback and expects a
reverse proxy in front of it.

Required in production:

- HTTPS, with HTTP redirected
- the proxy to set `X-Forwarded-For`, and `HX_TRUST_PROXY=1` so it is honoured
- HSTS, and the usual secure headers at the edge
- the master key supplied by the environment or a secret manager, not a file in
  the deployment
- a process supervisor, so a restart is automatic
- backups of the client registry and the provider connection store

No self-signed certificate is generated to satisfy a check. A gate that is
satisfied by a fake certificate teaches a deployment to ship with one.

## Graceful shutdown

Stop accepting, let in-flight requests finish, flush buffered counters, exit.
Without the flush, the last interval of usage counters is lost on every deploy —
small, and systematically in one direction.

Processes are stopped by signalling a known process id. Nothing in this system
kills by image name; that would take an operator's unrelated services with it.

## Things deliberately not done in v1

- No streaming transport in production. The change feed is the contract a
  transport would carry.
- No multiple simultaneous providers, no fallback, no blending.
- No automated billing. Usage counters are telemetry, buffered and flushed; they
  are not billing records and must never be invoiced from.

See also: [Authentication](authentication.md) ·
[Provider setup](provider-betsapi.md) · [Errors](errors.md)
