The conversations: who actually got your words
Every message you sent has an addressee — this post reads your Instagram export's distribution side: your share of the volume, the per-contact and per-group ratios, the words behind the messages, the reply-pace medians, and the limits (no read receipts, anonymised senders, media isn't words).
The messages folder is mechanics — sharding, the participant rule, unsent messages,
HTML fallback — and mechanics say what is in the folder. What they cannot say is
where your words went. A message is not a diary entry: it has an addressee, and once
every message you ever typed sits in one place an unavoidable arithmetic appears —
some people got almost all of it. None of it requires inference; it is counting over
rows you own, and this post does the counting.
That arithmetic is the distribution of your words: your share of the total
volume, the ratio in each individual relationship, the actual word counts behind
the message counts, and the reply-pace medians that say who kept the thread
alive. None of it requires inference — it is counting over rows you own. All of
it requires care, because a percentage about people can sound like a verdict
about them, and the honest version of this story is a ledger, not a ranking.
The export hands you every message row with a sender and a timestamp; the
product counts, per conversation, who sent what (youSent / theySent) and rolls
that into three views of the same distribution: your overall share of every
message in the archive (one percentage), your share per relationship (a
ratio bar per contact and per group), and the words behind them — total
words you wrote versus received, averages per message, your longest message ever.
The Messages view carries all of it: an "Overall split" ratio, "You vs Them" on
every conversation card, "Reply pace & balance" with median reply gaps, and a
month-by-month "Message flow - you vs them".1 The persona view
adds the word-level layer — "words you wrote" as a headline number, top words
with fillers counted separately instead of silently dropped, questions, emoji,
caps and terseness rates — all measured only from messages you sent.2
The honest limits are built into the same code: no read receipts or "seen"
state, so received never appears; an anonymised export ("User 1 / User 2")
needs you to tell the app which one is you;3 and "ghosted" or
"left on read" are export-window phrases — silence inside the archive, never a
claim about today.4
- 2minthe three ratios and where each one lives in the app.
- 4minwords vs messages: what is counted, what is separated out, what is never inferred.
- 4minthe honest limits: no receipts, anonymised senders, the export window.
What "your words" means, precisely
A message row in this export is a sender, a timestamp, and usually a text
string (media rows carry the type, not the text). The distribution layer counts
three different things on top of those rows:
- Message counts. Per conversation, each row is attributed to you or to
the other side —
youSentandtheySentaccumulate as the parser walks the
threads, and the totals roll up into the view's "Overall split".5 This
is the coarsest measure and the most robust: it does not care how long the
messages were. - Word counts. Text messages are tokenised; your totals include
wordsYou,avgWordsYou, yourlongestYoumessage, average characters,
question rate, emoji rate, caps rate and single-word ("ok", "haha") rate —
every one of them computed only from rows you sent.2 The received
side gets its ownwordsThem, so the comparison is symmetric. - Time between rows. "Reply pace & balance" takes the gaps where the
sender actually alternates — you, then them, or them, then you — caps the
window at twelve hours so an overnight pause does not masquerade as a
three-day wait, and reports the median for each direction.6 Medians,
not means, because one weekend-long reply skews an average and cannot move a
median.
Everything is computed from your rows at read time. There is no model in the
path, no "sentiment", no contact ever looked up — and the whole distribution
never leaves the device it was computed on.
The three ratios, and where each one lives
Your overall share is one number: of every message the export attributes to
the two sides, what percentage is yours. It appears as the "Overall split" ratio
card on the Messages overview, alongside the DM/group breakdown, and as the
"You vs Them" ratio bar the Features page describes as "the percentage of
messages you sent vs received in every view".7 A share near 50% means
alternation; a share at 70%+ means you were the engine of most of your threads
(and the mirror image — 30% — means you were the one being written to).
Neither is a health score. It is a split, like a pie chart of who typed.
Per-relationship shares are the interesting ones: each contact card carries
its own ratio bar (youSharePct = your share of that pair's two-way volume) and
each group carries a per-group share with the member count.8 This is
where the aggregate dissolves into specific histories: the friend you mostly
listened to, the group chat you carried, the thread where the balance flipped
after a certain year. The Top DM contacts and "Group activity breakdown"
sections are sortable views over exactly these per-relationship numbers.1
The word-level share sits one layer down: wordsYou versus wordsThem,
plus the per-message averages that correct the message-count bias — a conversation
of long letters and a conversation of "ok" can have identical message counts and
completely different word totals. The persona view shows "words you wrote" as a
headline KPI precisely because it, not the message count, is closest to "how
much of your thinking went where".2
Words are counted carefully, not sloppily
Two decisions in the word layer are worth knowing, because they change what the
numbers mean:
- Fillers are separated, not deleted. The top-words list drops stopwords
and words under three characters, but filler words are routed to their own
counter — "counted separately so it is visible instead of silently dropped" —
and the persona surfaces them as "filler words you reach for".9 The
design principle: hiding the boring words would make your vocabulary look
more interesting than it was. It is a counts product; it shows the counts. - Only text counts as words. Photos, reels, stickers and voice notes are
real messages but contain no typed words — the conversation-state classifier
labels these threads "media only: no typed words at all".10 So a
months-long exchange of reels can carry hundreds of messages and zero words
in both directions, and that is not a bug in either number.
Group chats and broadcast channels are counted in the overall totals where the
export lets them be (the sharding/participant rules from the folder post decide
what qualifies as a conversation), and each group gets its own you-vs-them
share — but a group's "they" is several people summed, so a group ratio is not
comparable to a two-person ratio. Read them in their own units.
The per-contact layer: statuses, lists and their fine print
Underneath the three headline ratios sits a per-contact classification that is
worth understanding precisely, because its labels are the ones most likely to
be misread. Every contact in the persona gets one of five statuses, and the
order the rules are checked in is the whole story:11
dead— the thread's last message is a year or more before theexport's reference point; the sorted list puts the quietest first.
active— the thread went silent for less than 30 days before theexport's edge. Inside the window, ordinary pauses count as ordinary.
they-ghosted— your last message has been sitting unanswered for 30+days inside the export (and the thread is not yet
dead).you-ghosted— their last message has been sitting unanswered for 30+days inside the export (and the thread is not yet
dead).silent— neither side's last word clears the rule cleanly.
The code's own comment is the reading guide: "Naming is about WHO went quiet,
not who spoke last."11 The status is always about unanswered rows
inside the archive — never about intent, never about today. Alongside the
statuses run the counters that feed the lists the view actually shows:
unanswered messages on each side, double-texts (two of your rows before one of
theirs), sub-minute replies in each direction, and left-on-read totals — each
one a pure count over rows, each one sorted into a list you can open and
inspect message by message.12
The sub-minute counters deserve one specific note, because they are the
closest this product comes to a compliment and the easiest to inflate: "you
replied under a minute" counts your rows arriving inside sixty seconds of
theirs, per contact, with the median gap reported next to it. One person who
habitually answers instantly while you sleep will dominate your "instant
replies" list without either of you being unusual — the median next to the
count exists precisely so a burst of one good week cannot masquerade as a
personality.
And one guard runs under all of it: contacts with too few rows are excluded
from the distribution-flip lists entirely (a two-message thread at 50/50 is
not a balanced conversation, it is a coincidence), and the export-window
cutoff is applied before any status is assigned — so an archive requested last
month can never label a thread dead merely because the account holder has
been inactive since.4 Every list above is therefore a count over
rows that cleared a minimum, a window and a naming rule — in that order — and
each row in each list opens to the messages that produced it.
The honest read: a ledger, not a verdict
The temptation with distribution numbers is to turn them into judgments — I
give more than I get, they never made an effort. The data cannot support
that leap, and the product refuses to make it, for reasons that are worth
internalising before you look at your own bars:
- Volume is not care. A person who sent you 40 long messages and a person
who sent you 6 short ones may care identically; only the counting differs.
- Direction is not quality. Being written to more than you write is not
neglect, and writing more is not neediness. The ratio is a fact about rows.
- The window is the world. Every percentage here is of your export — the
period the archive covers, ending at its last message. Older threads are
frozen at whatever balance they had when the export ended.4
What the distribution is good for is quieter and more durable: noticing which
relationships were reciprocal, which ones you carried alone, which year the
balance in a specific thread changed, and where your actual thinking — measured
in words, not in posts — was spent. That is a nostalgic fact about your own
attention, and it is the one this post exists to help you read.
The limits, in the order they bite
No read receipts, no "seen", no block list. The persona's own caveat string
says it: "Instagram exports carry no block list, no read receipts and no
'seen' state. 'Left on read' means they never wrote back — the closest honest
signal, never a confirmed block."4 So "got your words" in this post's
title means were addressed to them and counted in the archive — never they
read them. The export cannot know the latter about anyone.
Anonymised exports need one answer from you. When a thread labels its
senders "User 1 / User 2" instead of names, the sender-matching that powers
every ratio needs to know which label is the archive's owner — the app asks you
to point at your own label, and the KPIs resolve from there.3 Until
you do, per-contact shares cannot be trusted; this is disclosed as a caveat
rather than guessed.
"Ghosted" and "left on read" are export-window phrases. Per the built-in
caveats: ghosted means the last word in that chat was followed by 30+ days of
silence inside the export, and recency is measured to the archive's last
message, "not to today, so an older archive never makes everyone look
abandoned".4 These are the difference between "the rows stop here" and
"the relationship stopped here" — the product only ever claims the first.
Media-only and voice-note-heavy threads undercount words by design. Stated
above, repeated here because it distorts the word totals more than anything
else: message counts include every row; word counts include only typed text.
Compare word totals only between text-heavy threads.
Groups sum people. A group's "they" side is everyone else combined; a
40/60 group share says nothing about any single member. Per-person distribution
exists only in two-sided threads, so treat every group bar as one row of
household arithmetic rather than as a person's relationship with you.
Is my message ratio a red flag if it's 70/30?
It is a fact about rows, not an assessment of anyone. Two-sided chats where one
person types more are ordinary — people communicate at different volumes, and
life circumstances (work, exams, time zones) move any ratio over time. Read the
bar, notice whether it surprised you, and resist converting it into a verdict
about you or them; the data cannot tell you effort, care or intent, and no view
in this product scores them.
Why do words and messages disagree — more messages but fewer words?
Because they count different things: message counts include every row (photos,
reels, stickers and voice notes among them), while word counts include only
typed text. A thread of exchanged reels can be thousands of messages and zero
words; a thread of three long letters can be three messages and six hundred
words. Use message counts for activity, word counts for how much you wrote.
My export says User 1 and User 2 everywhere — are my ratios broken?
Not broken — waiting on one answer. Anonymised exports strip sender names, so
the app cannot tell which side you are; it asks you to identify your own label,
and all per-contact shares, your share percentages and the persona word totals
resolve once you do. This is a deliberate non-guess: picking a label at random
would silently swap you with someone else in every ratio in the archive.
Does 'who got your words' mean who read them?
No — and the distinction is structural, not pedantic. Exports carry no read
receipts and no "seen" state, so nothing in the archive can show that any
particular message was ever read. The distribution says where your words were
sent and counted; whether they were received is unknowable here, which is
exactly why the app's strongest phrasing stops at "they never wrote back" and
never says "they saw it".
Why does reply pace use a median and a twelve-hour cap?
Because reply gaps are wildly skewed: most alternations happen within minutes,
a few happen days later, and one vacation would dominate any average. The median
reports the typical wait honestly. The twelve-hour window exists so that
overnight and cross-timezone sleeps are not counted as deliberate silence —
gaps beyond it simply do not participate, in either direction.
Can I see these ratios per person over time?
Yes, in two places: each conversation's own view carries the month-by-month
"Message flow - you vs them", so you can watch a two-sided balance change year
over year inside one thread, and the persona's contact list sorts by your share
and by volume for the cross-thread view. What does not exist — anywhere, in any
product — is a ratio over time for the whole archive by contact, because the
export's own history is all the time there is.
Questions this comes up
The messages-folder post for what counts as a conversation (sharding, groups,
channels, unsent); the people-interactions post for the rest of the picture —
follows, reactions and comments beyond DMs; the eight-things post for what this
distribution cannot ever include (a call, a deleted message's words, the reply
that vanished); the relationships post for the graph those conversations sit
inside; the year-in-review post for where your words landed by year.
1: apps/web-next/src/views/Messages.tsx — "Overall split"
(SentRatio over DM + group totals), "You vs Them" per conversation (line 350),
"Reply pace & balance" (line 690), "When you talk - hour by hour" (line 700),
"Message flow - you vs them" (line 732), "Top DM contacts" (line 377), "Group
activity breakdown" (line 394).
2: packages/shared/src/analytics/you-persona.ts — YouTotals
(sharePct, wordsYou/wordsThem, avgWordsYou/avgWordsThem, longestYou,questionPct, emojiPct, capsPct, tersePct), all incremented only when the
sender matches the owner's self-keys; apps/web-next/src/components/YouPersona.tsx
— "words you wrote" KPI (line 387), "Filler words you reach for" (line 620).
5: packages/shared/src/parsers/messages.ts — the type-stats pass:youSent/theySent accumulate per conversation from sender === self checks
over every message row.
6: apps/web-next/src/views/Messages.tsx — reply-pace computation:
alternating-sender gaps only, g >= 12 * 3600 discarded, per-direction medians
(gapYouToTheir, gapTheirToYou), plus initYouPct (share of days your message
came first).
7: apps/web-next/src/views/Features.tsx — "You vs Them": "A visual
ratio bar showing the percentage of messages you sent vs received in every view."
8: packages/shared/src/analytics/you-persona.ts — YouContact.youSharePct
("your share of the two-way volume, 0-100"), YouGroupStat.youSharePct, and thetopYouTalk/topTheyTalk sorted lists.
9: packages/shared/src/analytics/you-persona.ts — fillers are counted
into a separate map so they are "visible instead of silently dropped"; stopwords
and words under three characters never enter the top-words list.
10: packages/shared/src/analytics/conv-states.ts — the media-only
state: "No typed words at all - photos, reels, stickers, voice notes."
4: packages/shared/src/analytics/you-persona.ts — the CAVEATS
array: no block list/read receipts/"seen"; recency measured to the export's last
message; "Ghosted means the last word in that chat was followed by 30+ days of
silence inside the export."
3: packages/shared/src/analytics/you-persona.ts — youSelfKeys()
and the anonymised-export caveat: "Exports that anonymise everyone as 'User 1 /
User 2' need you to tell LMKFR which one is you."
13: packages/shared/src/ama/tone.ts — the refusal contract for
questions about inner life and other people's states; this product never scores
reciprocity or infers feeling from timing.
11: packages/shared/src/analytics/you-persona.ts — YouChatStatus
(active | they-ghosted | you-ghosted | silent | dead), GHOST_DAYS = 30,DEAD_DAYS = 365, and the assignment block's comment: "Naming is about WHO went
quiet, not who spoke last: your message sitting unanswered for 30+ days means
THEY ghosted you."
12: packages/shared/src/analytics/you-persona.ts — per-contact counters
(unansweredByThem, leftOnReadByYou, youDoubleTexted, theyDoubleTexted,instantYou/instantThem, per-direction replyMinYou/replyMinThem medians)
feeding the theyGhosted/youGhosted/leftOnRead/speedDemons/instantLovers sorted lists; apps/web-next/src/components/YouPersona.tsx —
"They ghosted you" (line 551), "You ghosted them" (line 557), "You reply
instantly" (line 581), "They reply instantly" (line 588), "never answered" KPI
(line 395).
Footnotes
- messages-view
- persona
- selfkeys
- caveats
- parse
- reply
- features
- contacts
- fillers
- media
- status
- lists
- refuse