HISTORY
The bookmark with nothing behind it: what a link record is worth when its target is gone
This address held one saved link, catalogued into a Diigo group and cited from there. The record survives in someone else’s index; the page it pointed at does not, and cannot be recovered from what the record kept.
Published 2026-08-06 · Updated — · 2,763 words · 16 sources · Status of facts checked 2026-08-06
This address never held an article. It held a bookmark: one row in a database, published at a URL that a module generated from the row’s own title. From about 2007 until roughly 2010 the domain ran a social bookmarking community built on Drupal, and records of that kind were the bulk of what it served.
The record is gone, the site is gone, and the page the record pointed at is gone. What is left is a citation to this address from a group at groups.diigo.com, which is the reason anyone arrives here at all. That makes this URL a small and unusually clean specimen of the thing this publication exists to record: a pointer that outlived everything it pointed at, including itself.
The query, and the empty result
The Wayback Machine indexes this domain heavily. Asked for every URL at this host that it holds with an HTTP 200 capture, the CDX Server API returns 53,114 distinct addresses (checked 2026-08-06).1 It holds no capture of this path.
The query run against the CDX Server API on 2026-08-06 was https://web.archive.org/cdx/search/cdx?url=infopirate.org/bm_mind-opening-philosophy-article-reality&output=json. It returned an empty result: no capture, at any timestamp, at any status code. The negative result is the citation here, and it is repeatable.
The same API, asked for every URL at this domain beginning /bm_ — url=infopirate.org/bm_*&collapse=urlkey&output=json, run 2026-08-06 — returned 18,531 distinct URLs. Of those, 16,872 were captured with HTTP 200, 1,443 with a 404, 204 with a 301, and twelve with other codes. The earliest successful capture is dated 2007-07-13; the last first-time capture of a new record URL is dated 2016-03-13. Grouped by the year each URL was first captured, the 16,872 break down as 19 in 2007, 1,982 in 2008, 6,930 in 2009, 5,957 in 2010, 1,979 in 2011 and five in 2016.a
So the archive took roughly seventeen thousand of this site’s bookmark records and did not take this one. Nothing about that is remarkable. A crawler reaches what it is pointed at, when it is pointed at it, and a link record on a mid-size community site in 2009 had no claim on anyone’s attention. What matters is the cost, and the cost is specific rather than atmospheric: every field that record held is now unrecoverable except one.
What a record of this kind held
The shape of the record is knowable even though its contents are not, because the software is documented and its captured siblings are not.
In Drupal 6 a piece of content is a node, and the node table carries the columns nid, title, uid — the account that owns it — and created and changed as Unix timestamps, alongside type and status flags; the body text sits in the revisions table.4 Free-text tags are taxonomy terms attached to the node, each term a numbered row with a page of its own.5 A record page captured at this domain on 2007-07-16 renders exactly that set: a title, an owner reached through a per-account bookmark listing, a submission timestamp to the minute, a short note, a list of tag links each carrying its own numeric term id in its class attribute, a vote count, and the outbound link to the destination under a class named for that role.2 The site’s navigation offered four vote-ranked windows — today, this week, this month and all time — so the score was a published field and not an internal one.
One detail of the markup is worth recording, because it is the kind of thing a capture preserves and an index does not. The 2007 record pages load a stylesheet from a module directory named urlicon. URL Icon is a Drupal filter that inspects the anchors in a body of text, adds a CSS class reflecting each anchor’s target, and fetches and stores that target’s favicon under the site’s files directory.8 Seven of those stored icons survive in the archive for this domain, captured 2007-07-02, each named for the host it was fetched from.1 That is evidence about what this site linked to. It is not evidence about what any particular record linked to, and it does not name the destination of this one.b
Set against an empty address, that is a rich record. Set against the page it described, it is a label on an empty shelf. Title, tags, note, owner, date and score are all statements about a document. Not one of them is the document.
The path is the last field standing
One field did survive, and it survived by an accident of URL design.
Drupal sites of this vintage generated readable paths with the Pathauto module, which builds an alias from the content’s own text rather than leaving it at a numeric node path.6 Generation is lossy by design. The module’s 6.x-2.x branch defines a default ignore list of twenty-seven words at line 23 of pathauto.module: a, an, as, at, before, but, by, for, from, is, in, into, like, of, off, on, onto, per, since, than, the, this, that, to, up, via, with.7 Those are stripped out. Case is folded and punctuation is replaced by the separator.
That this install used that list is testable rather than assumed. The 18,530 distinct /bm_ paths in the CDX index hold 111,907 hyphen-separated tokens between them; three of those tokens are words on the ignore list. Words one space away from the list survive in bulk in the same corpus — and appears 2,140 times, how 2,004, your 1,198, you 582, or 247 (checked 2026-08-06).1
Run that in reverse on bm_mind-opening-philosophy-article-reality and the result is a boundary, not an answer. The record’s title contained the tokens mind, opening, philosophy, article and reality, in that order. Any of the twenty-seven ignored words could have stood between any two of them and left no trace. Capitalisation, punctuation and any hyphenation in the original are gone. A large number of distinct English titles collapse onto this one string, and this publication does not assert which of them was written.
That the erosion is real rather than theoretical is visible across the captured alias set at this domain. Of the 18,530 distinct /bm_ paths, 992 contain the digits 039 standing where an apostrophe was — the numeric core of the HTML entity ', surviving after the alias generator removed the punctuation around it. A further 396 carry percent-encoded non-ASCII, including sequences that are the UTF-8 bytes of a curly quotation mark encoded a second time. The two artefacts have different causes and neither cause could be verified for this install.
What the alias emphatically does not contain is the destination. The outbound URL was the one field with any power to recover the target, and it was never written into the path, never captured, and is not held by this publication. It could not be identified from any surviving trace.
Vanished, in the archive’s own vocabulary
The Internet Archive published a workable set of terms for this in April 2026. Writing on the archive’s blog, Sawood Alam classifies a URL as preserved when it is alive on the web and also archived, rescued when it is dead but archived, endangered when it is alive but archived nowhere, and vanished when it is dead and unarchived.3
The scale sits in the same piece. Pew Research Center reported in May 2024 that 38% of webpages that existed in 2013 were no longer accessible a decade later, and that a quarter of pages sampled across 2013 to 2023 had gone.9 Checking Pew’s 5.4 million shared URLs against the Wayback Machine, Alam reports 72% archived in total: 56% preserved, 16% rescued from the dead, and 18% alive but not yet archived in the Wayback Machine. About one URL in ten is vanished outright. Alam states that smaller web archives were not counted, so the endangered share is an upper bound.3
This address falls into the last category twice over. The record is vanished, and so, on the evidence the record left, is its target. The endangered figure is the one worth carrying away: close to one URL in five in that sample was alive when checked and held in no Wayback Machine capture, which is one shutdown away from the condition this address is already in.
Four ways a saved link stops working
These failures are usually discussed as one event. They are four, they have different remedies, and they arrive in this order.
- The target dies. The saved URL returns an error, a parked page or unrelated content. The record is intact and useless.
- The service dies. The record itself becomes unreadable. Ma.gnolia is the documented case with no window at all: a database and filesystem failure on 2009-01-30, followed by founder Larry Halff’s statement of 2009-02-17 that the user data was irretrievable. There was no export period, because there was no warning.
- The service survives but the record does not travel. Read-only, export-free or login-walled custody keeps the row and withholds the file. Delicious was put into read-only mode on 2017-06-15 with no new bookmarks and no API, and users were told there was no time pressure to migrate.10 Furl went the other way on 2009-04-17: the service closed and the bookmarks were migrated into Diigo.11
- The archive had it and no longer serves it. This is the case most fallback advice omits. In the Hacker News thread on Alam’s post, commenter firefoxd, writing 2026-07-04, describes a page that 404ed, was substituted with a Wayback link, and was then removed from the archive as well; commenter afpx, the same day, reports that many old bookmarked URLs have been removed from the Internet Archive, and that only some versions are affected.12 Both accounts are attributed rather than endorsed — from the index as it stands today, a withdrawn capture and a URL that was never crawled look identical.
What was stored at save time decides what stays answerable
The useful question about a collection is not which tool holds it. It is which of four things each record captured at the moment of saving, because once the target 404s that decision is fixed and cannot be revisited.
| What the record stores | Still answerable after the target 404s | Not answerable | Where this shape is found |
|---|---|---|---|
| URL alone | Whether an archive holds a capture | What it was called; what it concerned; whether it is even the right link | A plain list of addresses |
| URL, title, tags, note, date | What it was called, roughly what it concerned, when it was saved, and what the saver thought worth writing down | What the page said | The record type at this address; the browser interchange format |
| URL plus extracted readable text | What the page said, regardless of any archive | Layout, images, embedded media, what else was on the page | Read-later stores that keep article text |
| URL plus stored snapshot | What the page said and what it looked like, including assets | The rest of the site around it; anything behind a login the crawler could not reach | Archiving tools that write a snapshot folder for every URL |
The second row is where nearly everything sits, and the interchange layer is the reason. The Netscape bookmark file format — the <!DOCTYPE NETSCAPE-Bookmark-file-1> file that browsers export — defines a bookmark as <DT><A HREF="{url}" ADD_DATE="{date}" LAST_VISIT="{date}" LAST_MODIFIED="{date}">{title}</A>.13 A URL, a title and three timestamps. There is no field for the page. A format with no slot for the document cannot carry the document between services, so a migration routed through it lands in row two whatever the source and the destination were separately capable of.
Rows three and four have published costs. Readeck’s documentation puts storage at about 300 kB for each bookmark saved, which is the price of keeping readable text.14 ArchiveBox documents what it writes into a snapshot folder for every URL: an HTML and a JSON index, a SingleFile HTML snapshot, a wget clone with a WARC, a printed PDF, a 1440×900 screenshot, a rendered DOM dump, extracted article text, and a link to any copy on archive.org.15 Those are the two shapes that still answer a question after the source page has gone, and both state their cost in disk rather than in features.
A copy the service holds is not a copy the saver holds
One correction to a common assumption belongs here, and Ma.gnolia supplies it. Backblaze’s account of the failure, published 2009-03-02, records that the service was valued partly for caching the pages its users linked to, and describes the backup arrangement that did not survive contact with a restore: a file sync over FireWire to another machine, with no integrity checking, no versioning, and never once tested.16
Ma.gnolia was therefore a row-four service. It held copies of pages. Its users still lost those copies, because the copies were in the same custody as the index. Record shape and custody are separate variables: a stored copy raises what a record can answer, and changes nothing about who has to be asked.
Stated as a rule rather than a prescription: a collection that must survive the loss of the source page requires a stored copy of that page. A collection that must survive the loss of the service holding it requires that copy to sit in the saver’s own custody, in a format readable without that service. Those are two requirements, and satisfying the first does not satisfy the second. A collection that only needs to be re-findable while the web around it stays intact requires neither.
By that test the record at this address failed at the first requirement. It stored a pointer, and what remains is exactly what was stored: not the page, not the note, not the tags, not the date — five tokens in a path, and a citation from a group listing that has now outlived the entire site it pointed into. The audit that follows from this is at auditing a link collection for rot; the same four-way distinction, applied to searching rather than to saving, is set out at what a bookmark search actually indexes.
Where to find it now
- This address,
/bm_mind-opening-philosophy-article-reality - Gone No capture exists. The CDX query for this exact URL returned an empty result set on 2026-08-06, and this is the only one of the nine legacy addresses this publication serves for which that is true. No export of the record survives, no export endpoint exists to request, and this publication holds no copy of the old site’s database, its bookmarks or its accounts.
- The page this record pointed at
- Unverified Its URL could not be identified. The destination was written into the record and into the record’s rendered page, neither of which was captured. The alias preserves five tokens of a title and nothing further.
- The other bookmark records at this domain
- Gone from the live web, and partly recoverable from the Internet Archive: 18,531 distinct
/bm_URLs sit in the CDX index for this domain, 16,872 of them captured with HTTP 200, first captures running 2007-07-13 to 2016-03-13 (checked 2026-08-06). Captured records carry title, owner, timestamp, tags, note, vote count and destination. Nothing from them is republished here. - The citing group listing at groups.diigo.com
- Kept groups.diigo.com answered a request on 2026-08-06 with a 302 to
https://groups.diigo.com/index, which returned 200, and diigo.com returned 200 for its annotation and bookmarking product the same day. That listing is the only non-archival pointer to this address this publication located; whether any other exists could not be verified. - Ma.gnolia
- Gone Data irretrievable after the failure of 2009-01-30, confirmed by its founder on 2009-02-17; no export period existed. What users recovered afterwards came from third-party caches and feeds rather than from the service. ma.gnolia.com resolves and returns HTTP 200 (checked 2026-08-06), but what stands there is not a bookmarking service; the claim that page makes about itself is one this publication will not repeat, and the reason is recorded on the unverified list.
- Furl and Delicious, for contrast on where the data went
- Gone Furl closed on 2009-04-17 with bookmarks migrated into Diigo. Read-only Delicious was placed in read-only mode on 2017-06-15 — no new bookmarks and no API, but no deletion and no stated deadline; delicious.com now sits behind Cloudflare, returned HTTP 403 to an automated request on 2026-08-06, and is not an operating bookmarking service. Both outcomes are recorded row by row in the Ledger.
References
- Internet Archive. “Wayback CDX Server API.” internetarchive/wayback, GitHub, n.d. https://github.com/internetarchive/wayback/blob/master/wayback-cdx-server/README.md — retrieved 2026-08-06. Queries against
infopirate.orgrun 2026-08-06. ↩ - Internet Archive Wayback Machine. Capture of
infopirate.org/bm_calculate-adsense-revenue-any-site, 2007-07-16. https://web.archive.org/web/20070716031617/http://infopirate.org/bm_calculate-adsense-revenue-any-site — retrieved 2026-08-06. Cited for record structure only; no text from the page is reproduced. ↩ - Alam, Sawood. “Gone but Not Forgotten: Recovering the Dead Web.” Internet Archive Blogs, 2026-04-23. https://blog.archive.org/2026/04/23/gone-but-not-forgotten-recovering-the-dead-web/ — retrieved 2026-08-06. ↩
- Drupal. “modules/node/node.install, branch 6.x.” git.drupalcode.org, n.d. https://git.drupalcode.org/project/drupal/-/blob/6.x/modules/node/node.install — retrieved 2026-08-06. ↩
- Drupal. “modules/taxonomy/taxonomy.install, branch 6.x.” git.drupalcode.org, n.d. https://git.drupalcode.org/project/drupal/-/blob/6.x/modules/taxonomy/taxonomy.install — retrieved 2026-08-06. ↩
- Pathauto. “README.txt, branch 6.x-2.x.” git.drupalcode.org, n.d. https://git.drupalcode.org/project/pathauto/-/blob/6.x-2.x/README.txt — retrieved 2026-08-06. ↩
- Pathauto. “pathauto.module, branch 6.x-2.x.” git.drupalcode.org, n.d. https://git.drupalcode.org/project/pathauto/-/blob/6.x-2.x/pathauto.module — retrieved 2026-08-06. ↩
- Auditor, Stefan. “urlicon.module and README.txt, branch 6.x-1.x.” git.drupalcode.org, n.d. https://git.drupalcode.org/project/urlicon/-/blob/6.x-1.x/urlicon.module — retrieved 2026-08-06. ↩
- Chapekis, Athena, Samuel Bestvater, Emma Remy and Gonzalo Rivero. “When Online Content Disappears.” Pew Research Center, 2024-05-17. https://www.pewresearch.org/data-labs/2024/05/17/when-online-content-disappears/ — retrieved 2026-08-06. ↩
- Cegłowski, Maciej. “Pinboard Acquires Delicious.” Pinboard Blog, 2017-06-01. https://blog.pinboard.in/2017/06/pinboard_acquires_delicious/ — retrieved 2026-08-06. ↩
- Wikipedia contributors. “Furl.” Wikipedia, n.d.; cited as a secondary source. https://en.wikipedia.org/wiki/Furl — retrieved 2026-08-06. ↩
- Hacker News. “Gone but Not Forgotten: Recovering the Dead Web.” Item 48739682, submitted 2026-06-30; comments by firefoxd and afpx both dated 2026-07-04. https://news.ycombinator.com/item?id=48739682 — retrieved 2026-08-06. ↩
- Microsoft. “Netscape Bookmark File Format.” Microsoft Learn, 2017-08-15. https://learn.microsoft.com/en-us/previous-versions/windows/internet-explorer/ie-developer/platform-apis/aa753582(v=vs.85) — retrieved 2026-08-06. ↩
- Readeck. “Readeck Documentation.” readeck.org, n.d. https://readeck.org/en/docs/ — retrieved 2026-08-06. ↩
- ArchiveBox. “Output Formats: What ArchiveBox saves for each URL.” ArchiveBox/ArchiveBox, GitHub, n.d. https://github.com/ArchiveBox/ArchiveBox#output-formats — retrieved 2026-08-06. ↩
- Budman, Gleb. “Ma.gnolia Wilts with No Backup.” Backblaze Blog, 2009-03-02. https://www.backblaze.com/blog/magnolia-wilts-with-no-backup/ — retrieved 2026-08-06. ↩
Notes
- The
collapse=urlkeyparameter returns one row for each distinct URL rather than one row for each capture, and the row retained is the earliest, so the yearly figures count when each record URL was first archived and not how often it was revisited afterwards. They are counts of records, not of crawls. The 18,531 rows resolve to 18,530 distinct paths, which is the figure used for the token counts. ↩ - Module identification rests on the asset paths loaded by the captured pages, which name the module directories, and on the stored favicon files the archive holds under this domain’s files directory. Which release of Drupal and of each module this install ran could not be verified; the 6.x branches are cited because the captured markup matches that generation. The one alias prefix that is a reliable marker is
bm_for bookmark records; captured tag links point at bare title-slug paths, some with a leading underscore and some without, and no rule governing that difference could be verified. ↩
From about 2007 until roughly 2010 this address held a single bookmark record in a Drupal social bookmarking community that ran at this domain: a submitted title, an outbound link, a note, tags, an owner and a vote count, none of which is republished here. There is no Wayback capture of this path at any date, which makes it the only one of the nine legacy addresses this publication serves with no archival record at all; captures of the domain’s other bookmark records run from 2007-07-13 to 2016-03-13, and the site itself degrades after 2010. A group listing at groups.diigo.com still points here (checked 2026-08-06) — a record of the record, outliving both the bookmark and the page it named.