Measuring yourself without shaming yourself: the quantified-self trap
Why turning your Instagram history into a scoreboard backfires - Goodhart's law, self-concept injury, and how a tool that only counts refuses to grade you.
Tracking is useful until the metric becomes the goal - at which point Goodhart's
law applies to you: the number stops describing your behaviour and starts
reshaping it. That is why this product counts, times and compares, and never
grades: no composite 0-100 scores, no forecasts, no "you're doing badly"
tiles. A 200-post year is a collapse for one account and a debut for another;
the file cannot know which, so the software refuses to guess.
Self-tracking promised self-knowledge. What often arrives instead is a little
judge in a widget: a score going down, a streak broken, a red arrow where your
Tuesday used to be. The quantified-self movement was built to make behaviour
visible; most fitness and social dashboards quietly turned visibility into
evaluation - and evaluation is where the psychology curdles.
Goodhart's law, applied to yourself
Goodhart's original observation was about economics: when a measure becomes a
target, it ceases to be a good measure.1 Streaks were meant to
describe habit; they became the thing users protect, sleeping beside their
phone to keep a flame alive. Follower counts were meant to describe reach;
they became the thing people buy, beg and restructure their faces for.
The mechanism is well documented in self-regulation research: surveillance of
the self by a metric narrows behaviour toward the metric rather than toward
the underlying goal. When you watch a number, you optimise the number - and
the number is only a shadow of what you cared about (connection, expression,
curiosity). This is the trap:
- You measure to understand.
- The measure needs a scale to be readable.
- The scale implies good and bad.
- You start serving the scale.
A 0-100 "engagement score" completes step 3 by design. It has baked into it
someone's opinion of how much you should post, reply, be replied to. The
number looks like a thermometer - neutral, clinical - but thermometers don't
judge you for running a fever.
Self-concept injury: when the dashboard talks back
Why does it matter whether the number has an opinion in it? Because metrics
about social behaviour feed directly into self-concept - your working theory
of who you are. Psychology has long observed that self-discrepancies (the gap
between your actual self and the self you believe you should be) generate
distress, and that the size of the reference point drives the size of the
pain.2 A neutral count creates a small reference point ("I posted
40 times"). A scored grade creates a large one ("I am at 62/100 - a C").
The result is a two-sided failure that shows up in every quantified-self study
worth reading:
- Under-performance shame - the number is low, and because the number is
the target now, you are low. Deltas get read as verdicts, not observations.
- Over-performance capture - the number is high, and you become afraid to
change behaviour in any way that might lower it. The streak ends you, not the
other way around.
Both outcomes are forms of the same error: treating a descriptive statistic as
an evaluative one. And both are designed in whenever a product renders your
life as a single graded scale.
The refusal this product was built on
This is not a philosophical preference floating above the code - it is a
decision in the ledger with observable consequences. The decision D5 bans
composite 0-100 scores outright;3 the practical consequence across
Insights, Compare and Features is counts, medians and deltas only - score tiles
are nowhere to be found. Compare states in its own copy that it shows a
snapshot, not a prediction of where your numbers will be.4 There are
no forecast rows, no "you should post more" nudges, no red/green grades on your
own history.
What replaces the score is context you supply:
- Counts with their units and time windows visible, so "12 messages" can be a
quiet week or a reunion, and the app does not decide which.
- Deltas between two real exports - two measurements, not a projection.
- Medians and distributions instead of averages-as-targets, because a median
describes what happened, it does not imply what should happen next.
- Your own reading of the timeline - the year-in-review and heatmap exist to
jog memory, not to rank it.
Using your archive without becoming its subject
Practical rules that follow from the psychology and from what the code allows:
- Read trends over windows, not spot values. A single count invites
judgment; a span of counts invites observation. The heatmap and timeline
views exist for exactly this. - Never average yourself across people you don't know the context for.
Median reply time across ten conversations tells you about your ten
conversations, not about your character. - Treat surprise as information, not verdict. "I didn't realise I
searched for that in 2022" is the archive doing its job. "I'm at 40/100"
is an artefact this product refuses to produce - by policy, not by omission. - If a metric starts changing your behaviour, shrink its surface. That is
Goodhart working as predicted. The tool gives you filters and windows, not
streaks and badges, precisely so the option to shrink is always available.
Why doesn't this app have an overall engagement or 'archive health' score?
Because the decision ledger bans composite scores. A single number would have
to encode somebody's opinion of how much you should have posted, replied or
been replied to - and that opinion cannot survive contact with real variation
between accounts. You get counts, medians and deltas instead, each with its
window and units attached.
Do any scores exist anywhere in the product?
There is a per-person engagement tier in account views (core, active,
at-risk, casual, dormant) computed from measured message exchange between you
and that person. It is a sort key for a directory - never a composite score of
your account, never a prediction, never summed with anything else into a grade.
Is tracking my activity history bad for me?
Not inherently - visibility supports change when you want it. The documented
harms come from graded metrics that turn description into judgment. Keep the
metrics descriptive (how much, when, compared to what) and the archive stays a
mirror instead of a judge.
The year-in-review feels celebratory - is it manipulating me?
It is a memory aid: counts and spans laid out over a timeline you already
lived, in your browser, with no comparison to anyone else and no target implied.
There is no leader position, no percentile, and nothing to win or lose - the
numbers describe a year that is over.
Related reading
The parts-we-refused-to-build post for all seven refusals with their enforcing
code; the counts post for how every metric is derived and why units and windows
are always shown; the year-in-review post for how the same numbers are staged
as memory rather than as a grade.
1: Goodhart, C. - "Problems of Monetary Management: The U.K.
Experience" (1975), on measures that cease to work once they become targets;
restated as Goodhart's law in the quantified-self literature.
2: Higgins, E. T. - self-discrepancy theory (A Theory of
Cognition, Affect, and the Self, 1987): distress scales with the gap between
actual and reference selves - cited here to explain why a graded self-metric
hurts more than a descriptive one.
3: The project decision ledger D5 bans composite 0-100 scores;
observable consequence documented in content/blogs/2026-10-07-the-parts-we-refused-to-build.md
(section 5, with its scores footnote in that post).
4: Compare view copy, apps/web-next/src/views/Compare.tsx - the
explicit statement that Compare shows a snapshot between two exports rather
than a prediction.
Footnotes
- goodhart
- discrepancy
- ledger
- snapshot