How LMKFR counts what is countable — and what it refuses to infer
Medians that return null instead of zero, people counted through a normalised name rule, streaks defined as a seven-day gap — the exact counting rules this engine uses, and the list of conclusions it will not draw.
Every number this product shows you was produced by arithmetic you could, in
principle, check — and every number it does not show you was left out on purpose.
Both halves of that sentence are design decisions, and neither is obvious in an era
where "insights" usually means a black box with a percentage on top. This post is the
methodology page: the three rules the counting engine runs on, the exact definitions
behind the metrics people quote at each other ("streak", "best time", "top
interactions"), and — the part that actually takes discipline — a plain list of the
conclusions it refuses to draw.
Nothing here is aspirational. Every rule is cited to the file that implements it, and
every refusal is a behaviour you can observe in the app: if a screen cannot show you a
thing, it is usually because the thing failed a rule below, not because the feature
ran out of time.
The engine counts only what the export recorded, normalises names before counting
people, and writes every time window down as a literal rule — a streak is "consecutive
posts no more than seven days apart", a gap is measured in whole days, best times are
the five busiest day-hour buckets. Medians and averages return null on empty input
rather than a fake 0. What it refuses: causation, composite 0–100 scores, prediction,
watch-history counted as action, fuzzy identity guessing, and any claim about what
another person felt or did. The ask-anything engine is deterministic retrieval with
no model behind it. Deterministic does not mean complete — it means the limits are in
the files and the rules, not in a mood.
- 2minthe three counting rules, in one breath.
- 5minthe metric definitions (streak, gap, best time, unique person, median).
- ongoing — the refusal list: what will never appear, and why each entry is excluded.
Rule one: nothing is counted that the file did not record
The engine is a pure function over your archive: same ZIP in, same numbers out, on any
machine, forever.1 No network, no model, no "enrichment" from outside sources.
That single property cascades into every number you see — a count can only ever be a
count of rows the export contained, which means the request's date range, format and
folder selection are upstream of every metric in the app. The methodology phrase for
this is measurement, not estimation: if the row is not there, the number does not
move.
The corollary is enforced at the taxonomy level, and it is the first honest rule most
users never notice: consumption and action are different series. Watch history —
what you viewed — is recorded separately from what you did (likes, comments, saves,
searches, answers), and the analytics classification keeps watch history out of
"activity you did" on purpose.2 A tool that blended them could show you a
glowing "engagement" number built mostly from scrolling. This one will show you what
you consumed and what you acted on, labelled apart, every time.
Rule two: a person is a normalised name, not a feeling
Counting people sounds easy until the same account appears as alex, Alex , andaléx across three files. The engine's answer is deliberately unglamorous: a
normalisation function rewrites surface forms into a comparable key before any
tallying, so double-counts die at the door.3 Unique-person counts, "accounts"
totals, and every per-person ranking are built on that key — a string rule, visible in
the source, applied identically everywhere.
What it will not do is guess that two different keys are the same human. No profile
matching, no avatar comparison, no "these probably overlap" heuristics: renaming an
account, owning two handles, or a person being present under similar names in
different lists will not be resolved into one identity by inference. The per-record
attribution the app does perform — best-effort account names for "who you saved from"
cards, say — is documented as best-effort for exactly that reason: it reads the field
the record carries, it does not reconstruct a person from vibes.4 The
conservative failure mode is over-splitting: you might see one human counted as two
closely-spelled accounts. That error is visible, fixable in reading, and far safer
than the alternative — inventing a merge between two accounts that were never the same
one.
Rule three: every window is a written rule
Metrics with hidden windows are how dashboards lie politely. Here are the definitions,
verbatim in spirit from the code that computes them:
- Streak — consecutive posts and reels whose gaps round to ≤7 days each,
counted over sorted timestamps; the run resets on the first longer gap.5
The dashboard's "N-week posting streak" sentence and Compare's "longest streak" row
both read this one function, so they cannot disagree with each other. - Gap metrics — every interval between consecutive posts is measured in whole days
(
/ 86400), withlongestGapthe maximum andaverageGapthe rounded mean of the
intervals. Rounded, because a metric reported to twelve decimals of a day is
precision theatre.5 - Best posting times — the top five day-hour buckets by count, sorted
descending, cut at five. Not "the optimal hour to post": the five busiest cells of
your own history, which is what the data can honestly support.6 - Comparison windows — two exports run through the same function, always. The
delta rows (posts per month, streak, gaps, new follows over the last three months
versus the three before) are differences between two identically-computed values,
never a re-derivation with different rules per side.7 - Year and silence spans — persona timing records the longest streak and the
longest silence in days with start and end dates, so "quietest stretch" is a
dated fact rather than a mood.8
Each definition is short enough to argue with, which is the point: you can disagree
with a seven-day streak threshold — you just cannot be surprised by it.
Illustrative. Posts on 3 and 9 May: six days apart, rounds to 6, inside the rule
— one run. Add a third on 21 May: the twelve-day gap resets it, leaving a streak of
2, longestGap 12, averageGap rounded from (6 + 12) / 2 = 9. Move that third post
to 17 May instead and the first gap becomes 14: now two streaks of 1, longestGap
14, averageGap from (14 + 4) / 2 = 9 again. Identical averages, opposite shapes —
which is exactly why the medians, the individual gaps and the streak ship together
instead of one headline number shipping alone.5
The windows also live in exactly one place each: the consistency function is called
by the dashboard and by Compare, the best-times buckets come from a single
histogram, and no card quietly re-implements a threshold locally. A definition that
exists in one file can be wrong once — visibly, fixably — rather than subtly, in
three cards that disagree with each other.
Medians, averages, and the null
The statistics layer is three functions, and their most important line is the one that
handles nothing: median, average and sum over a number array — each returningnull when the input is empty, not 0.9 A zero is a measurement ("none
happened"); a null is an absence ("there is nothing to measure"). A feed with no
recorded likes and a feed whose likes are missing from the request look identical the
moment either renders a 0, and that conflation has destroyed more honest dashboards
than any other single shortcut.
Median gets pride of place over mean wherever the question is "typical", because one
huge thread should not be able to vote your typical day into the ground — the same
reason household-income questions use medians. The functions are deterministic and
dependency-free: no statistics library, no floating-point surprises you cannot trace
in a ten-line file. If you want to check this engine's arithmetic, start there; it is
designed to be finished reading in one sitting. And when a screen shows nothing, it
is showing you the null rather than a hopeful zero — the two render differently in
every view that respects the distinction, because the whole point of measuring
record-only is that missing data never gets to impersonate a result.
The refusal list
Now the other half of the methodology: the conclusions this engine will not draw.
Each entry is a real omission, made for a stated reason, and the list is the product's
actual character — a tool is its refusals more than its features.
- No causation. Curves show when, never why — the posting-collapsed post is
entirely about holding this line. No screen correlates a decline with a cause,
because the file contains no causal field to read. - No composite scores. There is no 0–100 "health", "virality" or "engagement
score" anywhere in the product — the decision ledger for this project bans them
outright, and what you get instead is counts, medians and side-by-side deltas. If
two raw numbers cannot carry a verdict, no multiplier is allowed to invent one.10 - No prediction. Nothing forecasts followers, reach or next month's performance;
the Compare view states in its own copy that Instagram gives you a snapshot, not a
prediction about where your numbers will be.11 Extrapolating your posting
curve into a projection would violate rule one: the future is not a row in the
file. - No watch-history-as-action. Covered in rule one; it appears again here because
it is the most tempting violation in the whole product — viewing is abundant, and a
tool desperate for bigger numbers would have blended it in years ago. - No feeling attribution. Relationship annotations you set (blocked, muted,
close friends) are recorded as your actions. The engine explicitly does not
claim "you blocked them" from indirect signals, and it will not decode a
someone's silence into an intention.12 A gap between two people is a gap
between two records; what it meant belongs to the humans. - No cross-account identity merging. Rule two, restated as policy: similar names
are never fused by heuristic, and no social graph between accounts on this
platform exists for the product to traverse — the public-profile privacy control
that would need one is left unbuildable on purpose, with the gap stated in the
interface rather than faked.13 - No invented completeness. Absent folders read as absent (with completeness
warnings where a folder that should exist did not arrive), never as empty
activity. The eight-things post's table is the content-side version of this rule;
the parser's presence checks are the code-side version.
Why the ask-anything engine has no model
The one feature that sounds like it needs a language model — asking your archive
questions in plain sentences — is the most deterministic thing in the app. Questions
are handled by a worker-built BM25 index over your own text, assembled with
memoised, injected-clock context: retrieval and ranking over files, not generation.14
The sidebar says it in the product's own voice: no model, nothing uploaded.15
There is no LLM dependency anywhere in the codebase, no completion endpoint, and no
plan that adds one — the answer to "what did I say to this person in March" should be
a search result, and a search result is checkable in a way a generated paragraph never
is.
Are these numbers the same on every machine?
Yes. Pure functions over the same archive produce identical output — no randomness, no
clock reads at computation time, no external services. Two people with the same ZIP
and the same app version will see the same numbers, which is what makes any figure
here auditable rather than merely repeatable.
Why does some tile show a dash instead of zero?
Because the input was empty, and empty is not zero. The statistics functions returnnull when there is nothing to measure; the interface renders that as a dash. A 0
would assert that you measured and found none — a claim the file cannot support when
the folder is simply absent.
Is a streak of five posts really 'a streak'?
It is consecutive posts no more than seven days apart — the definition the code uses
and the label the UI shows. The threshold is a choice, stated openly so you can
discount it if your posting rhythm works in weeks rather than days. Every window rule
in the app is of that kind: declared, visible, arguable.
Why won't it tell me why my followers dropped?
Because no file in your export contains causes — only events. The engine will show you
when the drop happened, what continued around it, and how it compares to the other
export; the narrative is yours to supply, labelled as yours. See the posting-collapsed
post for the full treatment.
Does the ask-anything feature use AI to answer me?
No. It is local retrieval: a BM25 index built in a browser worker over your own
archive, with deterministic context assembly. No model, no network call, nothing
uploaded — the sidebar states it, the dependency list proves it (there is no AI SDK in
the codebase), and the answers are rank-and-retrieve over text you can open yourself.
Could two of these counts be wrong because of how names are counted?
They can over-split: if the same person appears under genuinely different names, no
heuristic will fuse them, because fusing is the riskier error. They will not
over-merge: distinct accounts are never collapsed on similarity. The normalisation
rule handles surface variants only; everything past that is left to you, on purpose.
Questions this comes up
The people-interactions post for how interaction totals are assembled per person; the
compare post for running two archives through the same functions; the posting-collapsed
post for the refusal to infer causes; the activity post for the consumption-versus-
action split in full.
1: packages/shared/src/ — the analysis engine is pure TypeScript with no
network calls; the AMA path additionally injects its clock and memoises context
(packages/shared/src/ama/context.ts) so even time-shaped answers are reproducible.
2: packages/shared/src/analytics/activity-taxonomy.ts — event kinds are
classified into groups (interactions, collected, people, content) with watch history
held apart from actions; people-mix.ts keeps relationship-only events out of
engagement totals.
3: packages/shared/src/names.ts — normName() normalises surface forms for
comparison; user indexing and all per-person counts run through it.
4: packages/shared/src/analytics/activity-taxonomy.ts — per-record
account attribution is documented as best-effort ("Best-effort account name for one
record").
5: packages/shared/src/analytics/content-performance.ts —getPostingConsistency: sorted post/reel timestamps, gap in days via / 86400,
streak continues while gapDays <= 7, returns longestStreak, longestGap,averageGap (rounded), totalPosts.
6: content-performance.ts — day-hour buckets sorted by count descending,
sliced to the top five.
7: packages/shared/src/analytics/compare.ts — both exports rungetPostingConsistency (and the other metric functions) identically; deltas are
subtractions of two like-computed values.
8: packages/shared/src/analytics/you-persona.ts — timing recordslongestStreak, longestSilenceDays, streakFrom, streakTo.
9: packages/shared/src/math.ts — median, average, sum, each returningnumber | null and null for empty input.
10: The project's decision ledger (D5) bans composite 0–100 scores; observable
consequence: Insights, Compare and Features render counts, medians and deltas, with no
score tiles — an earlier score-bearing design was stripped out of the product before
these posts were written.
11: apps/web-next/src/views/Compare.tsx — inline copy stating a snapshot
is not a prediction about where your followers will be in three months.
12: packages/shared/src/analytics/you-persona.ts — the persona module
documents that "you blocked them" is never claimed from indirect signals; relationship
badges are emitted only from list memberships that exist in the file.
13: apps/web-next/src/views/Curator.tsx — the followers-only visibility control
is left unbuilt with the reason shown: it would need a social graph the product does
not have, and a non-working control is worse than a stated gap.
14: apps/web-next/src/lib/amaClient.ts (worker-built BM25 index, synchronous
search provider) + packages/shared/src/ama/context.ts (deterministic context
assembly); no AI SDK exists in any package.json.
15: apps/web-next/src/components/Sidebar.tsx — the ask-anything nav tip
describes it as "no model, nothing uploaded".
Footnotes
- pure
- taxonomy
- norm
- besteffort
- streak
- besttimes
- compare
- persona
- math
- scores
- nopredict
- nofeelings
- gap
- ama
- sidebar