Your posts, stories, reels and DM media in your Instagram export

Your own content is spread across several files. Posts, reels, stories, profile photos, recently deleted, watch history and DM media are all parsed separately. This is what each bucket is, how it is found in the archive, and what can look like a missing post when it is not.

Your content is not one folder. Posts, reels, stories, profile photos, recently deleted, watch history and DM media are each collected from different file patterns. The parser scans the tree for those JSON/HTML files and merges them into separate buckets so nothing gets blended incorrectly.

The buckets

The content parser returns six buckets plus DM media: posts, reels, stories, profilePhotos, recentlyDeleted, watchHistory and dmMedia. Each is populated by matching file paths, not by guessing from a single "your_content" file.

Posts

posts*.json files are walked recursively. Any object that contains an array media alongside a timestamp/caption (creation_timestamp, taken_at, title or caption) is treated as a post entry. Captions become both title and caption, timestamps are taken from the container, the first media entry supplies mediaPath and mediaType, location name/address is captured if present, and hashtags are extracted from the caption.

Reels

Reels are collected from files matching /reels?(_\d+)?\.json$/i. They are parsed with the same media-file logic.

Stories

Stories are collected from JSON files whose path contains stories or matches stor(?:y|ies)(?:_\d+)?\.json. The collector deduplicates by uri and prefers per-media timestamps (m.creation_timestamp, m.taken_at) falling back to container timestamps.

Profile photos

profile_photos.json is parsed as a media file and produces profilePhotos.

Recently deleted

Pulled from recently_deleted_content/*.json and recently_deleted_content.json.

Watch history

Reads JSON files matching: watch_history.json, watched_videos.json, viewed_reels.json, or videos_watched.json (case-insensitive). HTML variants with matching names/paths also contribute.

DM media

DM media comes from message_\d+\.json (and .html) under messages/inbox/ or messages/message_requests/. The thread folder is extracted, repaired for mojibake, and used as chat. Each message's photos[0].uri and videos[0].uri produce entries in dmMedia; timestamp prefers media creation_timestamp over message timestamp. HTML parsing extracts image/video links from table rows.

HTML fallback

If JSON parsing yields zero posts, the parser falls back to HTML for posts and watch history; DM media is still collected from message HTML shards.

Enrichment and structure

All media items carry hashtags extracted from captions. MediaItem includes optional href, location, gps (lat/lng/accuracy), and chat (for DM media). Timestamps are normalized to seconds where applicable.

Why is DM media in Content instead of Messages?

DM media is your content shared in chats; surfaced in content view as dmMedia with associated chat.

My posts.json is empty but I see posts in HTML?

If JSON produces zero posts, the parser falls back to HTML posts and still merges DM media.

Are stories deduplicated?

Yes. Stories are deduplicated by uri to avoid counting the same media twice when the JSON structure repeats it.

Does the export include recently deleted items?

If present under recently_deleted_content, they are collected into recently_deleted.

Sources of truth

  • packages/shared/src/types/content.ts - MediaItem, ContentData buckets.
  • packages/shared/src/parsers/content.ts - path-based collection for posts/reels/stories/profile photos/recently deleted/watch history, DM media extraction from message shards with mojibake repair and timestamp preference, HTML fallback.
Your posts, stories, reels and DM media in your Instagram export — LMKFR Blog