AI Skill Report Card

Extracting Archived URLs from Wayback Machine

A88·Oct 8, 2026·Source: Extension-page
15 / 15

Get the full list of archived URLs for a domain via the CDX API:

http://web.archive.org/cdx/search/cdx?url=example.com*&output=txt

Get the closest archived snapshot for a single URL:

http://archive.org/wayback/available?url=example.com/old-post

In Google Sheets (with the ImportJSON add-on installed):

=IMPORTJSON("http://archive.org/wayback/available?url="&A1,"/archived_snapshots/closest/url")
Recommendation▾
Add a brief note on rate limiting/throttling considerations when hitting the archive.org API at scale for many URLs.
14 / 15

Progress:

  • Step 1: Pull the raw URL list from the CDX API
  • Step 2: Clean and dedupe the list
  • Step 3: Crawl current status codes and prioritize by value
  • Step 4: Resolve snapshot URLs for suspicious/unclear redirects
  • Step 5: Extract old content from snapshots if needed
  • Step 6: Reinstate or redirect content accordingly

Step 1: Pull raw data from CDX API

Use http://web.archive.org/cdx/search/cdx?url=DOMAIN*&output=txt. Always consider a limit parameter — without it, large/old domains return huge datasets. Full filtering options are in the CDX API GitHub docs (date ranges, status filters, collapsing).

Step 2: Clean and filter

  1. Import the .txt file into Google Sheets (File > Import).
  2. If columns aren't auto-split, select column A → Data > Split text to columns → separator = space.
  3. Filter column D (mimetype) to keep only text/html — drop css/js/images/etc.
  4. Delete all columns except the URL column (C).
  5. Find & replace :80 out of URLs.
  6. Select the URL column → Data > Remove duplicates.

For recurring/bulk work (many client domains), script steps 1–2 instead of doing them manually each time.

Step 3: Audit and prioritize

Crawl the cleaned URL list with Screaming Frog (or similar) to get current status codes (expect a mix of 200, 301, 404). Don't try to manually review everything on large sites — prioritize URLs worth migrating by merging with:

  • Organic traffic (Analytics)
  • Backlink profile (Ahrefs/SEMrush)
  • Conversions / pageviews (customer journey relevance)

Only content answering "yes" to generating traffic, backlinks, conversions, or being part of the customer journey needs migration/redirect verification.

Step 4: Verify ambiguous redirects

When a redirect maps an opaque old URL (e.g. /post/123456) to a new slug (e.g. /post/how-to-tie-a-tie), you can't confirm correctness from the URL alone — pull the old snapshot:

  • Google Sheets: install the JSON add-on (Add-ons > Get Add-ons > search "JSON"), activate it (Add-ons > Import JSON > Activate), then use IMPORTJSON() as shown in Quick Start to get the snapshot URL for each old URL in bulk.
  • Python equivalent:
Python
import requests api_endpoint = 'http://archive.org/wayback/available' out = [] for url in urls: r = requests.get(api_endpoint, params={'url': url}).json() extraction_url = r['archived_snapshots']['closest']['url'] out.append(extraction_url)

Open the resulting snapshot URLs and compare content against the migrated destination to confirm the redirect target is correct.

Step 5: Extract old content at scale

If valuable content wasn't migrated at all:

  1. Identify the CSS class/id or XPath that wraps the main content in the site's template (inspect a live or archived page — archive.org preserves the original markup/classes unchanged).
  2. In Google Sheets, use =IMPORTXML(snapshot_url, "xpath_expression") to pull the content directly from the snapshot.
  3. In Python, use BeautifulSoup with the equivalent CSS selector/XPath against the snapshot's HTML.
  4. Refine the selector to isolate only the main content (exclude author bios, comments, sidebars) — this takes iteration but pays off at scale.

Step 6: Reinstate content

Copy the recovered content back onto the client's site (it's their own original content) or use it to validate/fix redirect mappings.

Recommendation▾
Include an example showing a 'bad' outcome (e.g., over-broad XPath pulling in sidebar content) to contrast with the good extraction example.
17 / 20

Example 1: Input: Domain minderest.com, need full historical URL list. Output: http://web.archive.org/cdx/search/cdx?url=minderest.com*&output=txt → imported into Sheets, split by space, filtered to text/html, deduped → clean list of unique historical URLs ready for crawling.

Example 2: Input: Screaming Frog shows /post/123456 301-redirects to /post/how-to-tie-a-tie, but the old URL gives no clue about topic. Output: IMPORTJSON("http://archive.org/wayback/available?url=example.com/post/123456","/archived_snapshots/closest/url") returns the last snapshot URL; opening it confirms the archived page was indeed about tying a tie, validating the redirect.

Example 3: Input: A blog section was dropped during migration and never redirected or recreated. Output: For each old URL, resolve the snapshot URL, then IMPORTXML(snapshot_url, "//div[@class='content']/div[2]") extracts the original post body text for republishing.

Recommendation▾
Clarify CDX API output format columns upfront (timestamp, original, mimetype, statuscode, digest, length) since Step 2 references column letters without defining the schema.
  • Always cap/limit CDX API queries on large or long-lived domains to avoid oversized responses.
  • Prioritize ruthlessly: not every old URL deserves a redirect — use traffic, backlinks, and conversion data to decide.
  • Script the CDX fetch + cleanup steps if auditing multiple domains regularly.
  • Spend time finding a precise, generic XPath/CSS selector — it saves far more time than it costs when extracting many pages.
  • Cross-reference archive.org data with Analytics/Ahrefs/SEMrush before concluding a URL matters.
  • Forgetting the limit parameter and pulling an unmanageably large CDX dataset.
  • Treating every CDX row as a unique URL — remember rows are snapshots, so duplicates are expected and must be removed.
  • Assuming small/niche websites will have good archive.org coverage — they often don't.
  • Trying to migrate every old URL instead of filtering by actual value (traffic/links/conversions).
  • Using an XPath/selector that's too broad, pulling in comments, author bios, or navigation along with the real content.
0
Grade AAI Skill Framework
Scorecard
Criteria Breakdown
Quick Start
15/15
Workflow
14/15
Examples
17/20
Completeness
18/20
Format
14/15
Conciseness
14/15