Local-first vs cloud: where should your Instagram export be analyzed?
Analyzing your Instagram export in your browser versus uploading it to a cloud analysis service - what leaves your device in each model, what privacy, retention and deletion each implies, offline capability, cost posture, and how to choose honestly.
Analyzing your Instagram data export locally means the ZIP never leaves your device:
your browser unzips it, parses it and computes every number, and the cloud model
inverts that - the file, or a copy of it, is uploaded to somebody else's server
first, and the analysis happens there. Both models can produce honest numbers from
the same file. They differ in custody: what leaves your device, who else can read
it, how long it is kept, how deletion works, and whether you can verify any of those
answers yourself. For an archive that holds your messages, your logins, your
contacts and a great deal about other people as well as you, that difference is the
actual decision.
This post compares the two models at the category level. It names no product,
praises no brand and criticises none: the questions it asks of a cloud service are
the same ones you can ask of any tool that wants the file, including
this one. Where it describes
local-first behaviour in detail, that behaviour is checkable rather than asserted -
the rest of this site documents exactly what the analysis engine does and does not
send.
Local-first keeps the archive inside your device's memory: nothing is uploaded,
nothing is retained on a server, deletion means deleting local files you can watch
disappear, and the analysis keeps working with the network switched off - which is
also how you verify the claim instead of trusting it. Cloud analysis requires the
export to leave your device first, which structurally introduces a retention window,
a breach surface and a delete process you must trust rather than check. Choose
local-first when the file is sensitive and your device can carry the work, which is
the normal case for a personal archive. Choose a cloud tool when your device
genuinely cannot do the job, or when you actively want the analysis stored somewhere
and are willing to pay for that in money or in trust.
- 1minthe two models, and what physically leaves your device in each.
- 5minprivacy, retention, deletion, offline capability and cost, compared line
- by line.
- 2minwhen a cloud tool is legitimately the better choice, and the limits
- local-first does not pretend to solve.
Two models, defined
Local-first means the analysis runs on the machine holding the file. You request
the export yourself, download the ZIP, and open it in an application whose parsing,
indexing and arithmetic execute in your own browser. The operator's servers, if the
product has any, serve application code - they do not receive your archive. That is
also the definition used by the local-first
post on this site, and it is a property you can
test rather than a slogan: if the archive's bytes never leave, then disabling the
network changes nothing about the analysis.
Cloud analysis means the file, or a copy of it, is transmitted to a third party
first. The transmission arrives in several familiar shapes: a form with an upload
field, a "drop it here to analyse" page, or a request that you connect an account so
the service can fetch data on your behalf. All three place your records on hardware
you do not control, operated by people you do not control, under a retention policy
you have to read rather than observe. The analysis itself may be excellent. The
question this post stays with is the one that precedes it: where is the file while
the excellent analysis is happening?
Nothing about cloud analysis is dishonest by construction, and nothing about
local-first is virtuous by construction. A local tool can still ship telemetry, and
a cloud tool can still delete on schedule and charge you money for the privilege.
The comparison below is about what each architecture makes possible, and what each
one asks you to take on faith.
What leaves your device in each model
The first table is the whole difference in one view. Both columns describe the
category, not any particular product.
| Thing | Local-first (analysis in your browser) | Cloud / third-party service |
|---|---|---|
| The export ZIP | Stays on the device; it is read into memory for parsing | Uploaded - a copy now exists on someone else's hardware |
| Parsed records: messages, contacts, logins, activity | Computed and held in browser memory or local storage | Transmitted as data, then stored at least for the session |
| Derived numbers, charts, timelines | Rendered locally from your own computation | Returned over the network, and must exist somewhere server-side to be returned |
| Login or account credentials for the analysis | None - the file itself is the input | A new account on the service, and in some designs a connection to your Instagram account as well |
| Network traffic during analysis | Zero requests carry archive contents | The archive itself, plus ongoing calls for results |
| Other people's data inside your archive | Never leaves with you | Leaves with you: your DMs are other people's words too |
| Who can answer "where is my data now" | You, by looking at your own disk | The operator, on request, subject to their policy |
Two rows deserve emphasis. The other people's row is the one people forget: an
export is a copy of a conversation archive, so uploading it transmits your
correspondents' messages, shared media references and contact details as well as
yours - the trust post makes
this point at length, because consent to a copy of your own history is not consent
to redistribute theirs. And the derived numbers row is subtler than it looks: a
hosted service that can show you your results again tomorrow has to keep something
tomorrow - your input, your results, or enough state to rebuild them. The retention
question for a cloud tool is therefore never whether, it is how long, on whose
hardware, and what happens when you ask for it to end.
Privacy, retention and deletion: what each model implies
Local-first gets its privacy from the shape of the system rather than from a
policy. Data that was never transmitted cannot be breached from a server, subpoenaed
from a server, or mined on a server - the archive's entire exposure is the machine
it already sat on before any tool existed. Retention is whatever you decide: the ZIP
and the browser's local storage are files you can list, copy or delete. Deletion is
observable, which is the property that matters most - you do not file a request and
wait for confirmation, you delete and then check. The honest cost of that model is
that custody is now entirely yours: the same
storage discipline any export needs
still applies, because a local file left in a synced Desktop folder or an unzipped
archive parked in Downloads for three years is a local problem, not a network one.
Cloud analysis gets its privacy from promises: a written policy, a contract, an
audit, and the operator's continued existence and good behaviour. Those can all be
real. The structural points are that you cannot verify them the way you can verify a
network-off local run, and that every cloud copy multiplies the places your archive
exists: primary storage, backups, logs, and whatever subprocessors the service uses
for compute or storage. Retention becomes a number in a policy document, and
deletion becomes a process with a queue - which is exactly why a service that cannot
state its retention window and its deletion behaviour in plain language has not
answered the question yet. For contrast, this site's own optional account layer is
documented down to that level: instant erase of the profile, a 30-day grace on the
account, then hard deletion, with the same erase-plus-grace mirrored on the sync
tier (security post). That is the
granularity worth demanding of any model, cloud included.
The third question in the trust set - whether derivatives of your upload are used to
train anything - only exists in the cloud column. A local analysis produces its
numbers in your browser and leaves nothing behind for anyone to learn from. Once the
file has been uploaded, "what happens to the derived data" becomes a question with a
policy answer rather than a physical one.
Offline capability: what happens when the network goes away
Local-first treats the network as irrelevant to the work. The analysis engine in
this product contains no fetch, no XMLHttpRequest and no socket usage, so once
the application is loaded, parsing, indexing, timeline construction and every
statistic execute without a single outbound request. That is also the cheapest
verification available to anyone: load your archive with networking disabled and use
the app. If the numbers still appear, the claim "nothing was sent" was never a claim
- it was an observation.
Cloud analysis has the opposite property. No network, no tool: the file is on the
server, your session is on the server, your results are on the server, and an
offline device is a device that cannot reach any of them. There is a second,
slower version of the same dependency worth naming. A local analysis stays runnable
from the ZIP you still hold, indefinitely, including after a product is abandoned,
a policy changes, or a company shuts down. A hosted analysis history lives at the
pleasure of the host: when the service ends, the analysis ends with it, and what you
keep is whatever you thought to export.
Cost posture: who pays for the analysis, and how
Nothing about either model is free; the models just move the bill.
Local-first concentrates the cost on your side: CPU time, memory, disk space and
battery, paid once per analysis by the person who wanted the analysis. On the
operator's side it deletes the per-user line item - an archive that never transits a
server never costs the operator compute or storage per user, which is why the
economics of a local-first product look like the economics of a static website
(the economics post itemises
this from the repository's own planning documents). That architecture is what allows
the analysis itself to be offered without a data business attached to it.
Cloud analysis carries real per-user costs that do not go away: the archive
occupies storage while it is retained, the analysis occupies compute while it runs,
and the results occupy storage for as long as the service promises to show them
again. Those costs are funded somehow, and "somehow" is part of the product: a
subscription is the straightforward version, a free tier with constraints is another,
and a free service that answers none of the three trust questions in writing is
implicitly telling you which column of that sentence it lives in. The honest summary
is not that one model is cheaper. It is that a local tool charges you in resources
you can see, and a cloud tool charges you in money or in trust, and an operator who
says the analysis costs them nothing has not explained who is paying.
When a cloud tool is legitimately better
An honest comparison has to say where its own preferred model loses, so:
- Your device cannot carry the work. Parsing a large ZIP and holding the parsed
records costs memory. An old laptop, a constrained phone or a locked-down machine
that cannot keep the archive in a browser tab is a real constraint, and a server
with more RAM is a legitimate answer to it. If the local attempt crashes the tab
repeatedly, "use local-first" is not advice, it is a dead end. - You want the analysis stored somewhere other than the machine you ran it on.
Local-first means results live where they were computed. If what you actually need
is one workspace reachable from several devices, a hosted service supplies that
shape directly, and building it yourself out of a ZIP and a folder is a hobby
rather than a solution. - You are choosing custody deliberately, not by default. A cloud service that
answers all three questions in writing - processed where, kept how long, derivatives
used for what - charges you a visible price, deletes on a stated schedule, and is
under a jurisdiction you find acceptable is a fair trade. Paying money for a
documented retention policy is a cleaner transaction than receiving a free service
whose funding model is unstated. - The task itself is server-shaped. Some analyses are genuinely awkward on one
device: very large inputs, sustained processing you do not want to babysit, or a
workflow where the export arrives by connector rather than by file. When the
constraint is the work rather than the trust, moving the work is the rational
move.
The common thread: cloud is the better choice when the binding constraint is your
device or your need for hosted state, and the service's answers are written down.
When neither of those is true, cloud is not better - it is simply elsewhere.
The honest limitations of local-first
The local-first column above reads well, and these are the places it does not
resolve:
- It requires a decent device and a modern browser. The work that a server would
do is being done by your hardware, so analysis speed and the size of archive a
browser tab can hold are bounded by the machine in front of you. This is a real
limitation, not a footnote, and it is the first reason a cloud tool can be the
correct choice. - You bring the export. Local-first tools read files; they do not fetch them.
Requesting the archive is still a manual errand inside Meta's own Download Your
Information tool, with its published clock - up to 30 days for the email, then
about 4 days before the link expires1 - and the download, unzipping and
storage of the ZIP are yours. Nothing about choosing local-first shortens that
wait. - Your analysis lives where you ran it. With no account, there is nothing to sign
into on a second device, and no server-side history waiting for you elsewhere. If
you want the same numbers somewhere else, you bring the ZIP and run it again -
which is also the honest answer to "how do I move to a new laptop": re-run, do not
sync. - Diagnostics are limited by design. Because no one receives your archive,
no one can look at it to help you. Support conversations have to work from the
description you choose to give, which is the price of a system where the operator
structurally cannot see your data. - Local custody still includes local risks. Browser storage, the downloads
folder, a synced desktop, a stray git repository - the network attack surface
shrinks to zero and the file-handling discipline stays fully in force. Local-first
moves the risk off the wire; it does not move it off your disk. - The code you run is code you trust. There is no server-side step to audit
after the fact, so the verification has to happen against the application itself -
which is exactly why this product's claims are written next to the file paths and
network tests that back them, and why "check it instead of believing it" is the
recurring instruction on this site.
A one-minute decision procedure
Before handing any export to any tool, the three questions from the trust
post, in order:
- Is the file processed on a server? If the answer is yes, or is not answered
plainly, everything in the cloud column above applies.
- Is it kept afterwards, for how long, and where? A service that cannot state
this in writing has not finished answering.
- Are its derivatives used to train anything? This question only has a physical
answer in the local model; in the cloud model it is policy.
Then the verification step, which settles more than the promises do: watch the
network tab, or go offline and run the analysis again. A local tool shrugs. A cloud
tool stops. Either result tells you which column you are actually in.
Questions this comes up
The local-first post for what local mode does
and does not send; the security post for
the engineering posture behind parsing a ZIP of an entire history; the economics
post for why the cost curve
inverts; the trust post for the
method of judging any tool, including ones this site has never heard of.
Does local-first mean my Instagram export never touches the internet?
No, and the distinction matters. Requesting the export is inherently online: you ask
Meta's Download Your Information tool for it and download the ZIP when the email
arrives. Local-first describes what happens after that. Once the file is on your
device, parsing, indexing and every statistic run in your browser, and in local mode
no byte of the archive is sent to us - which is a claim you can settle yourself by
loading the archive with the network disabled and watching the analysis finish
anyway.
Is a cloud analysis service automatically unsafe?
No. Unsafe is not a category, it is an unanswered set of questions: where the file
is processed, how long it is kept, what happens to derived data, how deletion works,
and who else is in the chain. A service that answers all of those in writing, charges
money rather than harvesting attention, and deletes on a stated schedule is offering
a normal, honest trade. The problem is not cloud analysis; it is cloud analysis that
cannot answer the questions, which is why this post gives the questions rather than
a verdict about businesses it has not inspected.
What do I need to analyze my export locally?
A browser, room for the ZIP plus the space the parsed records occupy, and enough
memory for the archive to be held while it is read. There is nothing to install and
no account to create for local analysis, and the work is done in the tab - which is
also why a very old or memory-constrained device is the one honest reason to look at
the cloud column instead. If the tab survives loading your archive, the machine can
carry the analysis.
Can I start locally and move to a cloud tool later?
Yes, and the order is the safe one. A local analysis leaves the original ZIP in your
possession, so switching models later is a decision you can still make - hand the
file to a hosted tool, read its retention answers first, and delete there if you
stop wanting it. The reverse direction is weaker: once an archive has been uploaded,
your only lever over that copy is a deletion request and whatever the policy says
about backups, which is a materially less certain position than still holding the
file.
1: Meta Help Centre, "Download Your Information" - the official description
of the request process and its two published clocks (up to 30 days for delivery,
about 4 days for the download link once it arrives).
<https://help.instagram.com/581066165581870>
Footnotes
- meta-dyi