One person, near-zero cost, no data broker
Why this tool cannot become a data broker — the cost curve of local-first software, the actual free-tier bill, and the four structural reasons there is no revenue machinery pointed at your archive.
The question behind this title is not really about money. It is: if I put my archive
into a product, what is the machine on the other side of the transaction, and what
does that machine want from me? Every familiar answer is some version of a data
broker — your attention sold to advertisers, your data sweetened into a free tier,
your usage feeding a model. Those answers exist because the usual economics force
them: ingesting, processing and storing a person's data costs the operator real
money per user, and someone always has to pay that bill.
This product's answer is that the bill mostly does not exist. Analysis runs on your
device; the default deployment has no server processing path at all; the optional
account layer fits inside free tiers that its own planning documents itemise; and
the things a broker would need — collected data, a sales channel, an incentive —
are each structurally absent rather than merely promised away. This post lays the
economics out so the ethics can be checked instead of asserted.
Local-first inverted the cost curve: your device does the multi-gigabyte work, so
each additional user costs the operator approximately nothing — the default
deployment is a static app on a free hosting tier, and the optional claim/sync layer
itemises to Vercel Hobby, Neon Free, Cloudflare R2 Free and a deferred email
provider. A data broker needs three things this product does not have: data worth
selling (the default mode collects none; the opt-in mode stores lean encrypted text,
media never), a channel to sell it through (no ad SDKs, no affiliate links, no
marketing scripts — see the refusals list), and a revenue motive (no growth metric,
no investor deck; the license is noncommercial). Add instant erase with a 30-day
grace and the brokerage case is not unethical here — it is unbuildable.
- 2minthe cost inversion: whose CPU, whose bill.
- 4minthe actual free-tier line items, and the three absences that kill the
- broker model.
- ongoing — what "near-zero" honestly hides, and how to verify all of it.
The cost curve that makes the rest possible
Traditional cloud software has one unavoidable fact: every user's data transits the
operator's servers and is processed on the operator's CPUs. At one user that is
pennies; at a million users it is a data centre — and a data centre whose marginal
cost scales with user count is a data centre that eventually needs revenue per user,
which is where "the user is the product" stops being a proverb and starts being an
accounting identity.
Local-first software deletes the line item. Your archive is unzipped, parsed,
indexed, compared and asked questions on your machine — the analysis engine has
zero network calls, deliberately.1 The operator's marginal cost per user is
therefore not "compute plus storage plus bandwidth"; it is: serving a static
application bundle, roughly the same bytes for the millionth user as for the first.
Storage scales with the user's own device. Bandwidth moves the archive exactly zero
times in the default mode.
This is not charity — it is the product's architecture paying its own way. And it is
the precondition for everything else on this page: a product that never receives
your data has a structurally different set of temptations than one that bathes in it.
The actual bill
Concrete line items, taken from this project's own planning documents rather than
vibes:2
| Component | Tier used | Headroom as documented |
|---|---|---|
| App hosting (Next.js, static + claim routes) | Vercel Hobby | Functions/workflows for the claim layer's small routes |
| Account metadata (claim/sync rows) | Neon Free | 0.5 GB — "hundreds of thousands of rows"; sparse traffic ≈ a few compute-hours a month |
| Opt-in profile blobs | Cloudflare R2 Free | 10 GB-month, 1M class-A writes, 10M class-B reads, **zero egress fees** — "hundreds of lean profiles" before leaving free tier |
| Transactional email (confirmation/reset) | Deferred | Only relevant to signed-up users; no marketing mail exists to send |
For a deployment with the claim layer unconfigured — the default — the table
collapses to its first row: a static site. The economics of that state are the
economics of a blog. The claim/sync tier was designed backwards from its free-tier
limits: schema kept to lean rows, blobs kept to one encrypted pack per person,
media excluded entirely (which is as much an economics decision as a security one —
see the refusals post).
Note what is not on the bill: no media-delivery charges (none is stored), no
per-row storage tier for archives (none is received), no inference compute for
generated answers (no model runs), no data pipeline to enrich anything (nothing is
collected). The missing line items are the interesting ones — each is a service
another product needs precisely because it holds your data, and this one's bills
are absent for the same reason its brokerage case is.
Why the broker model does not fit
A data broker is a machine with three inputs. Here is each one, checked:
1. Data worth selling. In the default mode, the operator never receives your
archive — there is no upload endpoint in the processing path, so the dataset that
would be sold does not exist on the operator's side of the wire.3 In the
opt-in modes, what the operator holds is an encrypted email address, credential
hashes, session hashes, and (if you explicitly enable sync) a lean encrypted pack of
your own aggregates — no media, no raw messages without a separate explicit opt-in,
and no other person's data, ever.4 Compare that with a broker's inventory
wishlist: this is closer to a password manager's holdings than a social platform's.
2. A channel to sell it through. The refusal list documents the absence: no
advertising SDKs, no affiliate links, no marketing scripts, no behavioural tracking —
one redacted page-view beacon is the entire outbound telemetry of the
product.5 A broker without a buyer-side integration is just a person with
a spreadsheet.
3. An incentive — revenue machinery pointing at user data. There is none. No
growth target depends on engagement, no ad auction improves with more data, and the
license under which the product is published is noncommercial: use, audit and running
are permitted; selling the software or derivative services is not.6 The
same license therefore constrains anyone who forks the idea — the refusal travels
with the source. And on the retention side the posture points the same way: erase
is instant (profile gone at once, account hard-deleted after a 30-day grace), local
data clears with a single tap, and a broker's core asset — retained history — is
actively destroyed on request rather than quietly archived.7
The ethics conclusion is boring on purpose: selling data would require having it,
reaching buyers, and wanting to — and all three fail as engineering facts, not moral
resolutions. That is the only durable kind of "we would never."
What "near-zero" is honestly hiding
The title says near-zero, not zero, and the difference is where the honesty lives:
- Free tiers change. Neon, Vercel and R2 can raise limits or prices; the
planning docs' headroom math is a snapshot, not a contract. The mitigations are
ordinary: lean schemas (headroom measured in hundreds of thousands of rows),
declared providers (any change is a documented change), and a default mode that
costs nothing regardless of what hosting does. - Time is the real cost. One person's evenings are the actual budget of this
product; there is no staff to pay because there is no staff. That cuts both ways —
it is why nothing costs anything, and why support, cadence and incident response
are what a single maintainer's calendar allows, no more. - "One person" means the operation is individuals, not a company. No funding
round to service, no board asking about data-monetisation surfaces, no acquisition
pressure where a user base's value is appraised in exactly one dimension: whether
the archive inside it can be sold. - The domain and the platform account are real money — the kind of money that
fits in a personal budget line. Near-zero is a description of the shape of the
costs, rather than variable and per-user), which is the part that matters for the
broker question: per-user cost is what forces per-user monetisation, and it is the
line this architecture deleted.
The inversion compounds on itself: an archive that never transits a server is an
archive that cannot be breached on one, subpoenaed from one or mined in one — the
cost saving and the custody guarantee are the same fact viewed from two sides, which
is why "local-first" is one word and not two.
If it is free and local-first, what is the catch?
There is no usage-based catch: the default mode has no per-user cost to recover.
The genuine costs are time (one maintainer's), the small fixed bill for the domain
and platform accounts, and free-tier limits on the optional claim layer — all
documented above. The product is noncommercial by license, so "the catch" cannot
quietly become a paid data tier either.
How do I verify the app never sends my archive anywhere?
Load your archive with networking disabled and use the app — processing completes
with zero outbound requests for archive contents, which is the local-first post's
documented test. Independently: the analysis engine contains no fetch or socket
calls at all, and the only outbound requests in the application source are
claim-layer emails.
What exactly would someone buy if they acquired this product?
On the default deployment: nothing — no archives are received. On an
enabled claim/sync deployment: encrypted credentials, an encrypted email, and lean
encrypted aggregate packs — media never, raw messages never without separate
opt-in — all subject to instant erase, and all constrained by a noncommercial
license that restricts what a buyer could do with the code itself. The security post
has the full inventory.
Does 'no data broker' mean no third parties at all?
It means no third parties receive your data in default mode; the deployment
itself still runs on named infrastructure (Vercel; plus Neon/R2/email only when the
claim and sync layers are enabled) — every provider is named in the security
document rather than discovered later. Page views pass through the redacted beacon
described in the refusals post. There is no data-buying or data-selling party in
that list because there is no such role in the system.
Why is the license noncommercial — is that about privacy?
It is about keeping the refusal list durable. A permissive license would let a
derivative product add the monetisation this one refuses; the noncommercial terms
keep the code — and by extension the posture — out of exactly the market this post
is about. Auditing, running and reading the source remain permitted; that is the
part that makes the claims on this page checkable.
Questions this comes up
The security post for custody, crypto and erase in detail; the local-first post for
the network-off verification; the refusals post for the seven features that would
have been a broker's toolkit; the off-by-default post for exactly what enabling the
optional layer creates.
1: packages/shared/src — the analysis engine contains no fetch,XMLHttpRequest or socket usage; parsing, indexing (including the ask-anything
BM25 worker) and all statistics execute in the browser.
2: docs/PROFILE-SYNC-PLAN.md — the provider/free-tier table (Vercel Hobby,
Neon Free with its 0.5 GB and compute-hour notes, Cloudflare R2 Free with zero egress,
email deferred) and its headroom math, quoted as documented.
3: The claim/sync routes are the product's only server code; with the layer
unconfigured they answer 503 and no processing path exists — see the off-by-default
post for the route-by-route behaviour.
4: db/schema.ts and SECURITY.md claim/sync sections — encrypted email,
argon2id hashes, opaque session tokens, encrypted profile packs of owner-side
aggregates; media excluded from the sync tier by design.
5: The refusals post — dependency list, redacted single beacon, absence of
ad/marketing SDKs, with file references for each.
6: Root LICENSE — PolyForm Strict 1.0.0: noncommercial use, audit and run
permitted; distribution, commercial use, modification and forking are not; every
package manifest carries SEE LICENSE IN <file>.
7: SECURITY.md data-lifecycle section — instant profile erase, 30-day
account grace then hard delete, single-tap local clear; the sync tier mirrors the
same erase-plus-grace.
Footnotes
- compute
- tiers
- default
- held
- refusals
- license
- erase