Which folders, documents and field names the archive can hold — and exactly how much of it we looked at.
An Instagram export is the archive Meta’s own “Download your information” tool builds and emails to you. It arrives as a ZIP of folders, and those folders hold JSON documents, HTML copies of the same documents, media files and a few plain-text markers. This page is about the shape of that archive: what it is called, which documents can appear, and which field names those documents use. It is not about what any archive says about anyone.
LMKFR is not affiliated with, endorsed by, or associated with Instagram or Meta. Those names appear here only to describe the export Meta’s own tool produces.
Sixteen export archives were inspected at the structure level only. Two things were read: the names of the entries inside each ZIP — folder names and file names — and, inside the JSON documents, the key names at every level of nesting. Key names are format: they say that a field exists, not what it held.
Nothing else was read. No values, no counts of records, no dates or numbers describing anyone’s activity, and no handles, usernames, display names or profile pictures were opened, recorded or published. Where a path or a folder name would identify a person — conversation folders are named after the people in them — it is shown here as a placeholder.
Everything below is the union of what was seen across the whole sample, written as “the format can include…”. There is no per-archive breakdown and no “this many of the sixteen contained X”: a count of exports is a statement about people’s accounts, not about a file format. The sample size is stated only because it describes our own method.
The ZIP opens straight onto named sections — no single wrapper folder in the sample — plus a start_here.html index at the root. The union of top-level entries is:
your_instagram_activity/ — the largest section: comments, likes, media, messages, saved items, story interactions, shopping, monetization, and several categories that carry only a placeholder (below).connections/ — followers_and_following/ plus a contacts/ folder.personal_information/ — personal_information/, information_about_you/, device_information/, and placeholder folders for autofill and meta-account data.security_and_login_information/ — a login_and_profile_creation/ folder of account-history documents.logged_information/ — searches, link history, and past insights.ads_information/ — ad and topic documents under ads_and_topics/ and instagram_ads_and_businesses/.apps_and_websites_off_of_instagram/, preferences/, media/ and files/.Nesting is consistent enough to read by eye: a section, a category folder, then documents. Media files are not mixed into the JSON — they sit in their own folders, such as media/posts/, media/stories/ and messages/inbox/<conversation>/photos/, with some media folders grouped into period-named subfolders.
A document is normally a .json file with an .html twin of the same name beside it, so the same record set can exist in two encodings. Alongside them:
followers_1.json, message_1.json / message_2.json, posts_1.json, post_comments_1.json. A reader that stops at the first shard reads a fraction of the format.no-data.txt instead of a document — the folder exists and says so, rather than being missing.messages/inbox/<conversation>/message_1.json and messages/message_requests/<conversation>/message_1.json. The folder name is the conversation title, so it is masked here; the document name underneath is fixed..jpg, .png, .webp, .mp4) next to the documents that reference them.follow_requests_you’ve_received.json, profiles_you’ve_favorited.json — and one media entry was observed with no file extension at all. Both are things a parser has to survive.The documents reuse a small set of envelopes. Below, the field names are real key names observed in the sample, grouped by what they describe. Representative examples only — this is a map of the vocabulary, not a dump of it.
personal_information.json holds a profile_user array whose string_map_data entries are named after the field itself — Username, Name, Bio, Email, Phone Number, Date of birth, Gender, Private Account, Music on profile — and each one carries value, href and timestamp. The profile photo arrives as media_map_data with a uri and a creation_timestamp. Neighbouring documents such as profile_based_in.json, locations_of_interest.json and possible_phone_numbers.json use the labelled envelope described below.
The security section uses named collections — account_history_login_history, account_history_logout_history, account_history_registration_info, account_history_password_change_history, account_history_account_privacy_history, account_history_account_active_status_changes — whose rows are string_map_data or string_list_data entries. Login and signup rows name fields such as IP Address, User Agent, Cookie Name, Language Code, Port and Time. Profile edits are a separate document with Changed, Previous Value, New Value and Change Date.
Lists — followers, following, likes, searches — are arrays of string_list_data entries carrying href, timestamp and value, usually with a document-level title. The collection keys around them are part of the format too: relationships_following, likes_comment_likes, searches_user, searches_hashtag, searches_keyword. Mutuals, unfollowers and similar views are not documents — they are differences computed between the lists that are.
Content documents carry a media array (or a named variant such as ig_stories or ig_profile_picture) whose entries hold creation_timestamp, title, uri, cross_post_source.source_app and a media_metadata block that can reach camera and photo EXIF keys such as device_id and source_type. The uri is a reference: the bytes are the separate media entries in the folders next to the JSON.
A thread document has a thread header — participants (each with name), title, thread_path, is_still_participant, joinable_mode, magic_words — and a messages array. Each message can carry sender_name, content, timestamp_ms, reactions (with actor, reaction, sticker_reaction), share (link, share_text, profile_share_name, profile_share_username, original_content_owner), plus photos, videos, audio_files and call_duration. Request threads add is_pending. Sharded siblings repeat the same message keys, which is why a thread has to be read across all of its shards.
The most widely reused envelope is an object with fbid, a top-level timestamp, a media array, and a label_values array whose entries carry label, value, title, href, timestamp_value, and nested dict / vec structures. It appears under likes, story interactions, ads and topics, saved items, shopping, settings, searches, and other activity — which is why one parser path covers a large share of the archive.
Past insights use named collections — organic_insights_interactions, organic_insights_audience, organic_insights_reach, organic_insights_posts — whose string_map_data keys name the metric: Impressions, Accounts reached, Followers, Profile visits, Saves, Likes, Date Range, each with value, href and timestamp.
Time is field-named rather than uniform: timestamp, creation_timestamp, timestamp_ms, timestamp_value, and the timestamp nested inside a string_map_data entry. A reader that looks for one name will silently miss the documents that use another.
Because every document has a fixed name and a fixed set of keys, an archive can be read by mapping files to fields — deterministic work that needs no model and no inference about what a sentence meant. That is the whole basis of what the Features page lists: the same input always produces the same figures, and the method can be written down.
The structure also shows, without opening a single value, that an archive is a set of records rather than a feed. Media is referenced by uri and stored as separate files. Sections with nothing to report still exist, carrying a no-data.txt marker rather than an empty document. And views that compare one list against another — who follows back, who stopped following — are not stored anywhere; they are computed from the lists that are.
The format is not fixed. Which documents appear depends on the options chosen when the export is requested — the encoding, the date range, the categories ticked — and on which sections an account has at all. Folder and file names also change over time, so an older archive and a newer one will not line up exactly, and the same record can exist in both a JSON and an HTML document with different field names.
This page therefore describes what the format can include, not what any particular export does contain. Nothing here states how many records a document holds, when anything happened, or which accounts had which sections — those are contents, not format, and they were never read.
For the reading rules and the counting rules behind the app, see the FAQ; for what is stored and what leaves your device, the Security page.
The Blog works through the same archive section by section, in prose. The About page explains what this project is and how the analysis works, and the Changelog dates what has shipped.