Extracting Archived URLs from Wayback Machine
Get the full list of archived URLs for a domain via the CDX API:
http://web.archive.org/cdx/search/cdx?url=example.com*&output=txt
Get the closest archived snapshot for a single URL:
http://archive.org/wayback/available?url=example.com/old-post
In Google Sheets (with the ImportJSON add-on installed):
=IMPORTJSON("http://archive.org/wayback/available?url="&A1,"/archived_snapshots/closest/url")
Progress:
- Step 1: Pull the raw URL list from the CDX API
- Step 2: Clean and dedupe the list
- Step 3: Crawl current status codes and prioritize by value
- Step 4: Resolve snapshot URLs for suspicious/unclear redirects
- Step 5: Extract old content from snapshots if needed
- Step 6: Reinstate or redirect content accordingly
Step 1: Pull raw data from CDX API
Use http://web.archive.org/cdx/search/cdx?url=DOMAIN*&output=txt. Always consider a limit parameter — without it, large/old domains return huge datasets. Full filtering options are in the CDX API GitHub docs (date ranges, status filters, collapsing).
Step 2: Clean and filter
- Import the
.txtfile into Google Sheets (File > Import). - If columns aren't auto-split, select column A → Data > Split text to columns → separator = space.
- Filter column D (mimetype) to keep only
text/html— drop css/js/images/etc. - Delete all columns except the URL column (C).
- Find & replace
:80out of URLs. - Select the URL column → Data > Remove duplicates.
For recurring/bulk work (many client domains), script steps 1–2 instead of doing them manually each time.
Step 3: Audit and prioritize
Crawl the cleaned URL list with Screaming Frog (or similar) to get current status codes (expect a mix of 200, 301, 404). Don't try to manually review everything on large sites — prioritize URLs worth migrating by merging with:
- Organic traffic (Analytics)
- Backlink profile (Ahrefs/SEMrush)
- Conversions / pageviews (customer journey relevance)
Only content answering "yes" to generating traffic, backlinks, conversions, or being part of the customer journey needs migration/redirect verification.
Step 4: Verify ambiguous redirects
When a redirect maps an opaque old URL (e.g. /post/123456) to a new slug (e.g. /post/how-to-tie-a-tie), you can't confirm correctness from the URL alone — pull the old snapshot:
- Google Sheets: install the JSON add-on (Add-ons > Get Add-ons > search "JSON"), activate it (Add-ons > Import JSON > Activate), then use
IMPORTJSON()as shown in Quick Start to get the snapshot URL for each old URL in bulk. - Python equivalent:
Pythonimport requests api_endpoint = 'http://archive.org/wayback/available' out = [] for url in urls: r = requests.get(api_endpoint, params={'url': url}).json() extraction_url = r['archived_snapshots']['closest']['url'] out.append(extraction_url)
Open the resulting snapshot URLs and compare content against the migrated destination to confirm the redirect target is correct.
Step 5: Extract old content at scale
If valuable content wasn't migrated at all:
- Identify the CSS class/id or XPath that wraps the main content in the site's template (inspect a live or archived page — archive.org preserves the original markup/classes unchanged).
- In Google Sheets, use
=IMPORTXML(snapshot_url, "xpath_expression")to pull the content directly from the snapshot. - In Python, use BeautifulSoup with the equivalent CSS selector/XPath against the snapshot's HTML.
- Refine the selector to isolate only the main content (exclude author bios, comments, sidebars) — this takes iteration but pays off at scale.
Step 6: Reinstate content
Copy the recovered content back onto the client's site (it's their own original content) or use it to validate/fix redirect mappings.
Example 1:
Input: Domain minderest.com, need full historical URL list.
Output: http://web.archive.org/cdx/search/cdx?url=minderest.com*&output=txt → imported into Sheets, split by space, filtered to text/html, deduped → clean list of unique historical URLs ready for crawling.
Example 2:
Input: Screaming Frog shows /post/123456 301-redirects to /post/how-to-tie-a-tie, but the old URL gives no clue about topic.
Output: IMPORTJSON("http://archive.org/wayback/available?url=example.com/post/123456","/archived_snapshots/closest/url") returns the last snapshot URL; opening it confirms the archived page was indeed about tying a tie, validating the redirect.
Example 3:
Input: A blog section was dropped during migration and never redirected or recreated.
Output: For each old URL, resolve the snapshot URL, then IMPORTXML(snapshot_url, "//div[@class='content']/div[2]") extracts the original post body text for republishing.
- Always cap/limit CDX API queries on large or long-lived domains to avoid oversized responses.
- Prioritize ruthlessly: not every old URL deserves a redirect — use traffic, backlinks, and conversion data to decide.
- Script the CDX fetch + cleanup steps if auditing multiple domains regularly.
- Spend time finding a precise, generic XPath/CSS selector — it saves far more time than it costs when extracting many pages.
- Cross-reference archive.org data with Analytics/Ahrefs/SEMrush before concluding a URL matters.
- Forgetting the
limitparameter and pulling an unmanageably large CDX dataset. - Treating every CDX row as a unique URL — remember rows are snapshots, so duplicates are expected and must be removed.
- Assuming small/niche websites will have good archive.org coverage — they often don't.
- Trying to migrate every old URL instead of filtering by actual value (traffic/links/conversions).
- Using an XPath/selector that's too broad, pulling in comments, author bios, or navigation along with the real content.