Last Known Copy

HISTORY

The bookmark with nothing behind it: what a link record is worth when its target is gone

This address held one saved link, catalogued into a Diigo group and cited from there. The record survives in someone else’s index; the page it pointed at does not, and cannot be recovered from what the record kept.

Published 2026-08-06 · Updated — · 2,763 words · 16 sources · Status of facts checked 2026-08-06

This address never held an article. It held a bookmark: one row in a database, published at a URL that a module generated from the row’s own title. From about 2007 until roughly 2010 the domain ran a social bookmarking community built on Drupal, and records of that kind were the bulk of what it served.

The record is gone, the site is gone, and the page the record pointed at is gone. What is left is a citation to this address from a group at groups.diigo.com, which is the reason anyone arrives here at all. That makes this URL a small and unusually clean specimen of the thing this publication exists to record: a pointer that outlived everything it pointed at, including itself.

The query, and the empty result

The Wayback Machine indexes this domain heavily. Asked for every URL at this host that it holds with an HTTP 200 capture, the CDX Server API returns 53,114 distinct addresses (checked 2026-08-06).1 It holds no capture of this path.

The query run against the CDX Server API on 2026-08-06 was https://web.archive.org/cdx/search/cdx?url=infopirate.org/bm_mind-opening-philosophy-article-reality&output=json. It returned an empty result: no capture, at any timestamp, at any status code. The negative result is the citation here, and it is repeatable.

The same API, asked for every URL at this domain beginning /bm_url=infopirate.org/bm_*&collapse=urlkey&output=json, run 2026-08-06 — returned 18,531 distinct URLs. Of those, 16,872 were captured with HTTP 200, 1,443 with a 404, 204 with a 301, and twelve with other codes. The earliest successful capture is dated 2007-07-13; the last first-time capture of a new record URL is dated 2016-03-13. Grouped by the year each URL was first captured, the 16,872 break down as 19 in 2007, 1,982 in 2008, 6,930 in 2009, 5,957 in 2010, 1,979 in 2011 and five in 2016.a

So the archive took roughly seventeen thousand of this site’s bookmark records and did not take this one. Nothing about that is remarkable. A crawler reaches what it is pointed at, when it is pointed at it, and a link record on a mid-size community site in 2009 had no claim on anyone’s attention. What matters is the cost, and the cost is specific rather than atmospheric: every field that record held is now unrecoverable except one.

What a record of this kind held

The shape of the record is knowable even though its contents are not, because the software is documented and its captured siblings are not.

In Drupal 6 a piece of content is a node, and the node table carries the columns nid, title, uid — the account that owns it — and created and changed as Unix timestamps, alongside type and status flags; the body text sits in the revisions table.4 Free-text tags are taxonomy terms attached to the node, each term a numbered row with a page of its own.5 A record page captured at this domain on 2007-07-16 renders exactly that set: a title, an owner reached through a per-account bookmark listing, a submission timestamp to the minute, a short note, a list of tag links each carrying its own numeric term id in its class attribute, a vote count, and the outbound link to the destination under a class named for that role.2 The site’s navigation offered four vote-ranked windows — today, this week, this month and all time — so the score was a published field and not an internal one.

One detail of the markup is worth recording, because it is the kind of thing a capture preserves and an index does not. The 2007 record pages load a stylesheet from a module directory named urlicon. URL Icon is a Drupal filter that inspects the anchors in a body of text, adds a CSS class reflecting each anchor’s target, and fetches and stores that target’s favicon under the site’s files directory.8 Seven of those stored icons survive in the archive for this domain, captured 2007-07-02, each named for the host it was fetched from.1 That is evidence about what this site linked to. It is not evidence about what any particular record linked to, and it does not name the destination of this one.b

Set against an empty address, that is a rich record. Set against the page it described, it is a label on an empty shelf. Title, tags, note, owner, date and score are all statements about a document. Not one of them is the document.

The path is the last field standing

One field did survive, and it survived by an accident of URL design.

Drupal sites of this vintage generated readable paths with the Pathauto module, which builds an alias from the content’s own text rather than leaving it at a numeric node path.6 Generation is lossy by design. The module’s 6.x-2.x branch defines a default ignore list of twenty-seven words at line 23 of pathauto.module: a, an, as, at, before, but, by, for, from, is, in, into, like, of, off, on, onto, per, since, than, the, this, that, to, up, via, with.7 Those are stripped out. Case is folded and punctuation is replaced by the separator.

That this install used that list is testable rather than assumed. The 18,530 distinct /bm_ paths in the CDX index hold 111,907 hyphen-separated tokens between them; three of those tokens are words on the ignore list. Words one space away from the list survive in bulk in the same corpus — and appears 2,140 times, how 2,004, your 1,198, you 582, or 247 (checked 2026-08-06).1

Run that in reverse on bm_mind-opening-philosophy-article-reality and the result is a boundary, not an answer. The record’s title contained the tokens mind, opening, philosophy, article and reality, in that order. Any of the twenty-seven ignored words could have stood between any two of them and left no trace. Capitalisation, punctuation and any hyphenation in the original are gone. A large number of distinct English titles collapse onto this one string, and this publication does not assert which of them was written.

That the erosion is real rather than theoretical is visible across the captured alias set at this domain. Of the 18,530 distinct /bm_ paths, 992 contain the digits 039 standing where an apostrophe was — the numeric core of the HTML entity ', surviving after the alias generator removed the punctuation around it. A further 396 carry percent-encoded non-ASCII, including sequences that are the UTF-8 bytes of a curly quotation mark encoded a second time. The two artefacts have different causes and neither cause could be verified for this install.

What the alias emphatically does not contain is the destination. The outbound URL was the one field with any power to recover the target, and it was never written into the path, never captured, and is not held by this publication. It could not be identified from any surviving trace.

Vanished, in the archive’s own vocabulary

The Internet Archive published a workable set of terms for this in April 2026. Writing on the archive’s blog, Sawood Alam classifies a URL as preserved when it is alive on the web and also archived, rescued when it is dead but archived, endangered when it is alive but archived nowhere, and vanished when it is dead and unarchived.3

The scale sits in the same piece. Pew Research Center reported in May 2024 that 38% of webpages that existed in 2013 were no longer accessible a decade later, and that a quarter of pages sampled across 2013 to 2023 had gone.9 Checking Pew’s 5.4 million shared URLs against the Wayback Machine, Alam reports 72% archived in total: 56% preserved, 16% rescued from the dead, and 18% alive but not yet archived in the Wayback Machine. About one URL in ten is vanished outright. Alam states that smaller web archives were not counted, so the endangered share is an upper bound.3

This address falls into the last category twice over. The record is vanished, and so, on the evidence the record left, is its target. The endangered figure is the one worth carrying away: close to one URL in five in that sample was alive when checked and held in no Wayback Machine capture, which is one shutdown away from the condition this address is already in.

Four ways a saved link stops working

These failures are usually discussed as one event. They are four, they have different remedies, and they arrive in this order.

What was stored at save time decides what stays answerable

The useful question about a collection is not which tool holds it. It is which of four things each record captured at the moment of saving, because once the target 404s that decision is fixed and cannot be revisited.

What the record storesStill answerable after the target 404sNot answerableWhere this shape is found
URL aloneWhether an archive holds a captureWhat it was called; what it concerned; whether it is even the right linkA plain list of addresses
URL, title, tags, note, dateWhat it was called, roughly what it concerned, when it was saved, and what the saver thought worth writing downWhat the page saidThe record type at this address; the browser interchange format
URL plus extracted readable textWhat the page said, regardless of any archiveLayout, images, embedded media, what else was on the pageRead-later stores that keep article text
URL plus stored snapshotWhat the page said and what it looked like, including assetsThe rest of the site around it; anything behind a login the crawler could not reachArchiving tools that write a snapshot folder for every URL

The second row is where nearly everything sits, and the interchange layer is the reason. The Netscape bookmark file format — the <!DOCTYPE NETSCAPE-Bookmark-file-1> file that browsers export — defines a bookmark as <DT><A HREF="{url}" ADD_DATE="{date}" LAST_VISIT="{date}" LAST_MODIFIED="{date}">{title}</A>.13 A URL, a title and three timestamps. There is no field for the page. A format with no slot for the document cannot carry the document between services, so a migration routed through it lands in row two whatever the source and the destination were separately capable of.

Rows three and four have published costs. Readeck’s documentation puts storage at about 300 kB for each bookmark saved, which is the price of keeping readable text.14 ArchiveBox documents what it writes into a snapshot folder for every URL: an HTML and a JSON index, a SingleFile HTML snapshot, a wget clone with a WARC, a printed PDF, a 1440×900 screenshot, a rendered DOM dump, extracted article text, and a link to any copy on archive.org.15 Those are the two shapes that still answer a question after the source page has gone, and both state their cost in disk rather than in features.

A copy the service holds is not a copy the saver holds

One correction to a common assumption belongs here, and Ma.gnolia supplies it. Backblaze’s account of the failure, published 2009-03-02, records that the service was valued partly for caching the pages its users linked to, and describes the backup arrangement that did not survive contact with a restore: a file sync over FireWire to another machine, with no integrity checking, no versioning, and never once tested.16

Ma.gnolia was therefore a row-four service. It held copies of pages. Its users still lost those copies, because the copies were in the same custody as the index. Record shape and custody are separate variables: a stored copy raises what a record can answer, and changes nothing about who has to be asked.

Stated as a rule rather than a prescription: a collection that must survive the loss of the source page requires a stored copy of that page. A collection that must survive the loss of the service holding it requires that copy to sit in the saver’s own custody, in a format readable without that service. Those are two requirements, and satisfying the first does not satisfy the second. A collection that only needs to be re-findable while the web around it stays intact requires neither.

By that test the record at this address failed at the first requirement. It stored a pointer, and what remains is exactly what was stored: not the page, not the note, not the tags, not the date — five tokens in a path, and a citation from a group listing that has now outlived the entire site it pointed into. The audit that follows from this is at auditing a link collection for rot; the same four-way distinction, applied to searching rather than to saving, is set out at what a bookmark search actually indexes.

Where to find it now

This address, /bm_mind-opening-philosophy-article-reality
Gone No capture exists. The CDX query for this exact URL returned an empty result set on 2026-08-06, and this is the only one of the nine legacy addresses this publication serves for which that is true. No export of the record survives, no export endpoint exists to request, and this publication holds no copy of the old site’s database, its bookmarks or its accounts.
The page this record pointed at
Unverified Its URL could not be identified. The destination was written into the record and into the record’s rendered page, neither of which was captured. The alias preserves five tokens of a title and nothing further.
The other bookmark records at this domain
Gone from the live web, and partly recoverable from the Internet Archive: 18,531 distinct /bm_ URLs sit in the CDX index for this domain, 16,872 of them captured with HTTP 200, first captures running 2007-07-13 to 2016-03-13 (checked 2026-08-06). Captured records carry title, owner, timestamp, tags, note, vote count and destination. Nothing from them is republished here.
The citing group listing at groups.diigo.com
Kept groups.diigo.com answered a request on 2026-08-06 with a 302 to https://groups.diigo.com/index, which returned 200, and diigo.com returned 200 for its annotation and bookmarking product the same day. That listing is the only non-archival pointer to this address this publication located; whether any other exists could not be verified.
Ma.gnolia
Gone Data irretrievable after the failure of 2009-01-30, confirmed by its founder on 2009-02-17; no export period existed. What users recovered afterwards came from third-party caches and feeds rather than from the service. ma.gnolia.com resolves and returns HTTP 200 (checked 2026-08-06), but what stands there is not a bookmarking service; the claim that page makes about itself is one this publication will not repeat, and the reason is recorded on the unverified list.
Furl and Delicious, for contrast on where the data went
Gone Furl closed on 2009-04-17 with bookmarks migrated into Diigo. Read-only Delicious was placed in read-only mode on 2017-06-15 — no new bookmarks and no API, but no deletion and no stated deadline; delicious.com now sits behind Cloudflare, returned HTTP 403 to an automated request on 2026-08-06, and is not an operating bookmarking service. Both outcomes are recorded row by row in the Ledger.

References

  1. Internet Archive. “Wayback CDX Server API.” internetarchive/wayback, GitHub, n.d. https://github.com/internetarchive/wayback/blob/master/wayback-cdx-server/README.md — retrieved 2026-08-06. Queries against infopirate.org run 2026-08-06.
  2. Internet Archive Wayback Machine. Capture of infopirate.org/bm_calculate-adsense-revenue-any-site, 2007-07-16. https://web.archive.org/web/20070716031617/http://infopirate.org/bm_calculate-adsense-revenue-any-site — retrieved 2026-08-06. Cited for record structure only; no text from the page is reproduced.
  3. Alam, Sawood. “Gone but Not Forgotten: Recovering the Dead Web.” Internet Archive Blogs, 2026-04-23. https://blog.archive.org/2026/04/23/gone-but-not-forgotten-recovering-the-dead-web/ — retrieved 2026-08-06.
  4. Drupal. “modules/node/node.install, branch 6.x.” git.drupalcode.org, n.d. https://git.drupalcode.org/project/drupal/-/blob/6.x/modules/node/node.install — retrieved 2026-08-06.
  5. Drupal. “modules/taxonomy/taxonomy.install, branch 6.x.” git.drupalcode.org, n.d. https://git.drupalcode.org/project/drupal/-/blob/6.x/modules/taxonomy/taxonomy.install — retrieved 2026-08-06.
  6. Pathauto. “README.txt, branch 6.x-2.x.” git.drupalcode.org, n.d. https://git.drupalcode.org/project/pathauto/-/blob/6.x-2.x/README.txt — retrieved 2026-08-06.
  7. Pathauto. “pathauto.module, branch 6.x-2.x.” git.drupalcode.org, n.d. https://git.drupalcode.org/project/pathauto/-/blob/6.x-2.x/pathauto.module — retrieved 2026-08-06.
  8. Auditor, Stefan. “urlicon.module and README.txt, branch 6.x-1.x.” git.drupalcode.org, n.d. https://git.drupalcode.org/project/urlicon/-/blob/6.x-1.x/urlicon.module — retrieved 2026-08-06.
  9. Chapekis, Athena, Samuel Bestvater, Emma Remy and Gonzalo Rivero. “When Online Content Disappears.” Pew Research Center, 2024-05-17. https://www.pewresearch.org/data-labs/2024/05/17/when-online-content-disappears/ — retrieved 2026-08-06.
  10. Cegłowski, Maciej. “Pinboard Acquires Delicious.” Pinboard Blog, 2017-06-01. https://blog.pinboard.in/2017/06/pinboard_acquires_delicious/ — retrieved 2026-08-06.
  11. Wikipedia contributors. “Furl.” Wikipedia, n.d.; cited as a secondary source. https://en.wikipedia.org/wiki/Furl — retrieved 2026-08-06.
  12. Hacker News. “Gone but Not Forgotten: Recovering the Dead Web.” Item 48739682, submitted 2026-06-30; comments by firefoxd and afpx both dated 2026-07-04. https://news.ycombinator.com/item?id=48739682 — retrieved 2026-08-06.
  13. Microsoft. “Netscape Bookmark File Format.” Microsoft Learn, 2017-08-15. https://learn.microsoft.com/en-us/previous-versions/windows/internet-explorer/ie-developer/platform-apis/aa753582(v=vs.85) — retrieved 2026-08-06.
  14. Readeck. “Readeck Documentation.” readeck.org, n.d. https://readeck.org/en/docs/ — retrieved 2026-08-06.
  15. ArchiveBox. “Output Formats: What ArchiveBox saves for each URL.” ArchiveBox/ArchiveBox, GitHub, n.d. https://github.com/ArchiveBox/ArchiveBox#output-formats — retrieved 2026-08-06.
  16. Budman, Gleb. “Ma.gnolia Wilts with No Backup.” Backblaze Blog, 2009-03-02. https://www.backblaze.com/blog/magnolia-wilts-with-no-backup/ — retrieved 2026-08-06.

Notes

  1. The collapse=urlkey parameter returns one row for each distinct URL rather than one row for each capture, and the row retained is the earliest, so the yearly figures count when each record URL was first archived and not how often it was revisited afterwards. They are counts of records, not of crawls. The 18,531 rows resolve to 18,530 distinct paths, which is the figure used for the token counts.
  2. Module identification rests on the asset paths loaded by the captured pages, which name the module directories, and on the stored favicon files the archive holds under this domain’s files directory. Which release of Drupal and of each module this install ran could not be verified; the 6.x branches are cited because the captured markup matches that generation. The one alias prefix that is a reliable marker is bm_ for bookmark records; captured tag links point at bare title-slug paths, some with a leading underscore and some without, and no rule governing that difference could be verified.

From about 2007 until roughly 2010 this address held a single bookmark record in a Drupal social bookmarking community that ran at this domain: a submitted title, an outbound link, a note, tags, an owner and a vote count, none of which is republished here. There is no Wayback capture of this path at any date, which makes it the only one of the nine legacy addresses this publication serves with no archival record at all; captures of the domain’s other bookmark records run from 2007-07-13 to 2016-03-13, and the site itself degrades after 2010. A group listing at groups.diigo.com still points here (checked 2026-08-06) — a record of the record, outliving both the bookmark and the page it named.