Internet archiving plays an important role in preserving websites, online publications, historical pages, documents, images, videos, and other digital resources that can disappear or change over time. Websites may be archived for research, journalism, legal documentation, historical preservation, academic work, cybersecurity investigations, or simply to maintain access to information that is no longer available on the live web.
The following list presents 100 well-known internet website archiving platforms and related digital preservation websites. They vary considerably: some crawl and preserve large portions of the public web, while others concentrate on national domains, academic resources, government information, social media, digital collections, or particular forms of online content.
Archiving tip: A website archive is not necessarily an exact copy of a live website. Dynamic content, databases, JavaScript applications, login-protected pages, robots restrictions, images, videos, and externally hosted resources may not always be preserved.
1. Internet Archive / Wayback Machine
The Internet Archive is one of the world’s best-known digital preservation organizations. Its Wayback Machine allows users to search historical snapshots of websites and examine how web pages appeared at different points in time.
2. Archive-It
Archive-It is an Internet Archive service used by libraries, universities, governments, museums, and other organizations to create curated web archives. Institutions can build collections around particular subjects, events, communities, or geographic areas.
3. Common Crawl
Common Crawl maintains a large open repository of web crawl data. Researchers, developers, and organizations can use its datasets for web research, large-scale analysis, search-related projects, and machine-learning applications.
4. Memento Web Time Travel
Memento provides technology and services for accessing archived versions of web resources across different web archives. Its approach makes it possible to locate historical versions of pages preserved by multiple archives.
5. Arquivo.pt
Arquivo.pt is Portugal’s web archive. It preserves Portuguese websites and other online information and provides search and access tools for historical web content.
6. UK Web Archive
The UK Web Archive preserves websites published in the United Kingdom. It is a major resource for researchers investigating British websites, organizations, events, and online publications over time.
7. National Library of Australia Web Archive
The National Library of Australia preserves Australian online content through its web-archiving activities. Its collections support research into Australia’s digital history.
8. Pandora Archive
PANDORA is an Australian web-archiving initiative associated with the National Library of Australia. It has preserved selected Australian online publications and websites for long-term access.
9. Trove
Trove is an Australian digital research platform operated by the National Library of Australia. It provides access to a broad range of Australian cultural heritage materials, including archived web resources.
10. Library of Congress Web Archives
The Library of Congress maintains web archives covering selected U.S. government, political, cultural, and historical web content.
11. Chronicling America
Chronicling America provides access to historical American newspapers and associated digitized materials. It is particularly useful for researching historical news rather than ordinary modern websites.
12. UK Parliament Web Archive
The UK Parliament’s web-preservation activities provide historical access to selected parliamentary online materials and websites.
13. European Web Archive
European web-archiving initiatives preserve online materials associated with European countries, institutions, culture, and public life.
14. Internet Memory Foundation Archives
The Internet Memory Foundation was involved in large-scale European web preservation and research initiatives. Its work contributed to the development of web-archiving technologies and collections.
15. WebCite
WebCite was designed to create permanent references to online documents and web pages, particularly for academic citations. It became known for helping researchers preserve cited web content.
16. Perma.cc
Perma.cc provides persistent archived links, particularly for academic, legal, government, and research citations. It is designed to help prevent link rot.
17. Archive.today
Archive.today, also known through related archive domains, allows users to capture snapshots of web pages and access previously preserved versions.
18. Archive.ph
Archive.ph is part of the Archive.today family of web-archiving services. It is frequently used to preserve individual web pages for later reference.
19. Archive.is
Archive.is is another domain associated with the same web-preservation service family and provides archived snapshots of web pages.
20. Archive.fo
Archive.fo is another entry point associated with the Archive.today archiving ecosystem.
21. UK National Archives Web Archive
The UK National Archives preserves selected online government and public-sector information, making it useful for historical research into British public administration.
22. Government of Canada Web Archive
Canada’s government web-archiving initiatives preserve selected federal government websites and online information.
23. Library and Archives Canada
Library and Archives Canada maintains extensive digital collections, including preserved online resources and government web content.
24. Government Web Archive
Government web archives preserve official websites, publications, announcements, policy documents, and other public information that may change or disappear.
25. EU Web Archive
European Union institutions maintain digital preservation resources covering websites and online publications associated with EU institutions and activities.
26. UK Government Web Archive
The UK Government Web Archive preserves historical versions of government websites and information published by British public authorities.
27. National Archives of Singapore Web Archive
Singapore’s national archival activities include preservation of online resources documenting the country’s government, society, culture, and digital development.
28. Hong Kong Web Archive
Hong Kong’s archival initiatives preserve selected websites and online resources documenting the territory’s cultural, governmental, and social history.
29. National Library Board Singapore Digital Collections
Singapore’s National Library Board maintains digital collections containing historical publications and other digital resources relevant to Singapore.
30. National Diet Library Web Archives
Japan’s National Diet Library operates web-archiving initiatives preserving selected Japanese websites and online publications.
31. WARP
WARP, the Web Archiving Project, is operated by Japan’s National Diet Library. It preserves Japanese government and institutional websites and other selected online resources.
32. National Library of Korea Web Archive
South Korea’s national library preservation activities include digital and web-based resources documenting Korean online culture and publications.
33. Pandora Australia
Pandora remains an important historical reference in Australia’s web-preservation ecosystem, particularly for selected Australian online publications.
34. New Zealand Web Archive
New Zealand’s National Library operates web-archiving activities that preserve selected New Zealand websites and online materials.
35. DigitalNZ
DigitalNZ provides access to digital collections from libraries, museums, archives, and cultural organizations across New Zealand.
36. Bibliothèque nationale de France Web Archives
France’s national library preserves selected French websites as part of its digital heritage programs.
37. BnF Gallica
Gallica is the digital library of the Bibliothèque nationale de France. It contains digitized books, newspapers, manuscripts, photographs, maps, and other historical materials.
38. ArchiveWeb
ArchiveWeb is a term associated with various web-archiving projects and tools designed to preserve or access historical online content.
39. Arquivo.pt Text Search
Arquivo.pt provides specialized search capabilities for finding historical Portuguese web content preserved through its archive.
40. Icelandic Web Archive
Icelandic web-preservation initiatives have archived websites published under Icelandic domains and other resources documenting Iceland’s online history.
41. Danish Web Archive
Denmark’s web archive preserves selected Danish websites and online publications as part of the country’s national digital heritage.
42. Swedish Web Archive
Swedish web-archiving projects preserve selected Swedish websites and online materials for historical and research purposes.
43. Finnish Web Archive
Finland’s national digital preservation activities include archiving websites and other online resources produced in Finland.
44. Norwegian Web Archive
Norway’s national web-archiving activities preserve selected Norwegian websites and online publications.
45. Estonian Web Archive
Estonia maintains web-preservation initiatives documenting websites and digital resources associated with the country.
46. Latvian Web Archive
Latvia’s national digital preservation activities include selected websites and other online publications.
47. Lithuanian Web Archive
Lithuanian web-archiving projects preserve selected Lithuanian websites and digital publications.
48. Swiss Web Archive
The Swiss National Library operates web-archiving activities covering selected Swiss websites and online publications.
49. Austrian Web Archive
The Austrian National Library preserves selected Austrian web content and digital publications.
50. German Web Archive
German libraries and cultural institutions operate web-archiving programs preserving selected German online resources.
Research tip: National web archives can be particularly valuable when investigating old government pages, local organizations, national news publications, historical events, and websites that were never widely indexed by international search engines.
51. Bibliotheca Alexandrina Internet Archive
The Bibliotheca Alexandrina has participated in digital preservation initiatives and maintains extensive digital resources related to cultural and historical materials.
52. Digital Public Library of America
The Digital Public Library of America aggregates digital collections from libraries, archives, and museums across the United States.
53. Europeana
Europeana provides access to millions of digitized cultural heritage objects from European libraries, archives, museums, and galleries.
54. HathiTrust Digital Library
HathiTrust is a major collaborative digital library containing digitized books, journals, and other scholarly materials contributed by participating institutions.
55. Google Books
Google Books provides searchable access to enormous quantities of digitized books and publications. It can be useful for finding historical references to websites, organizations, people, and events.
56. Google News Archive
Google News Archive provides historical newspaper material and can be useful when researching older online or offline news coverage.
57. NewspaperArchive
NewspaperArchive provides access to historical newspaper collections and is useful for researching older events and organizations.
58. Newspapers.com
Newspapers.com contains digitized historical newspapers. It can complement web archives when investigating the history of websites, companies, people, and events.
59. GenealogyBank
GenealogyBank provides historical newspapers and other genealogical records that can help establish historical context for organizations and individuals.
60. JSTOR
JSTOR is a major digital research platform containing scholarly journals, books, and primary sources. Although it is not primarily a web archive, it is an important digital preservation resource.
61. Internet Archive Texts
Internet Archive’s text collections contain millions of digitized books, documents, magazines, and other publications.
62. Internet Archive Movies
The Internet Archive also preserves video collections, including public-domain and freely reusable material.
63. Internet Archive Audio
Its audio collections contain music, recordings, podcasts, speeches, radio programs, and other sound material.
64. Internet Archive Software Collection
The Internet Archive preserves software and historical computer programs, making it valuable for research into computing history.
65. Internet Archive Console Living Room
Console Living Room preserves selected historical video-game software and related computing experiences.
66. OldWeb.today
OldWeb.today allows users to interact with historical versions of websites using archived web content and old browser environments.
67. Webrecorder
Webrecorder develops tools and technologies for capturing interactive web content, including websites that are difficult to preserve using conventional crawlers.
68. Conifer
Conifer is a web-archiving platform based on Webrecorder technology. It allows users and institutions to capture and preserve web resources.
69. Browsertrix
Browsertrix provides browser-based web-archiving technology designed to capture modern websites and complex web experiences.
70. ArchiveWeb.page
ArchiveWeb.page is a browser-based tool for capturing web pages and creating web archives using modern web-archiving technology.
71. PyWB
PyWB is an open-source web-archiving replay system. It can be used to replay WARC-based web archives.
72. Webrecorder Player
Webrecorder’s playback technology allows archived web captures to be viewed in ways that preserve more of the original interactive browsing experience.
73. Heritrix
Heritrix is an open-source web crawler developed for web archiving. It has been used by major web-preservation organizations.
74. OpenWayback
OpenWayback is an open-source system for replaying web archives and viewing historical captures.
75. SolrWayback
SolrWayback is a web-archive search and replay interface built around Solr and Wayback technologies.
76. WARCnet
WARCnet supports research involving web archives, particularly through tools and datasets designed for analyzing archived web content.
77. Archives Unleashed
Archives Unleashed provides tools and services for analyzing large web archives and extracting research value from historical web data.
78. ArchiveSpark
ArchiveSpark is a framework designed to process large web archives and extract structured information from WARC files.
79. Web Archives for Historians
Web-archiving research projects for historians provide tools and methods for studying changes in websites, online communities, and digital culture.
80. IIPC Web Archives
The International Internet Preservation Consortium coordinates cooperation between organizations involved in web preservation around the world.
Preservation tip: For serious research, it is often useful to search more than one archive. A page missing from one archive may exist in another collection or national web archive.
81. ArchiveTeam
ArchiveTeam is a volunteer-driven digital preservation community that works to rescue online content threatened by deletion or shutdown.
82. ArchiveTeam Warrior
ArchiveTeam Warrior is a software environment that helps volunteers contribute computing resources to web-archiving projects.
83. UK Web Archive Search
The UK Web Archive’s search system enables researchers to locate preserved British websites and archived online resources.
84. European Parliament Web Archive
The European Parliament has preserved historical online materials documenting parliamentary activity and institutional information.
85. World Wide Web Foundation Archives
Digital-history and web-preservation projects associated with the history of the World Wide Web provide resources for researching the development of online publishing.
86. Stanford Web Archive
Stanford University has participated in web archiving and maintains collections of preserved web content for research and historical purposes.
87. Harvard Library Web Archive
Harvard Library operates web-archiving initiatives that preserve selected websites and online materials relevant to research and institutional history.
88. Columbia University Web Archive
Columbia University libraries have developed web-archiving collections supporting research and preservation of online resources.
89. Yale University Web Archive
Yale libraries preserve selected websites and digital resources as part of their broader digital preservation programs.
90. University of California Web Archives
University libraries within the University of California system have participated in web archiving and digital preservation initiatives.
91. DigitalNZ Collections
DigitalNZ connects users to digitized materials from participating New Zealand cultural institutions and provides an important gateway to preserved digital heritage.
92. Smithsonian Digital Collections
Smithsonian digital collections provide access to digitized cultural, scientific, historical, and artistic materials.
93. NASA Image and Video Library
NASA’s digital archive preserves extensive historical photographs, videos, mission materials, and other resources related to space exploration.
94. Wikimedia Commons
Wikimedia Commons is a large repository of freely licensed and public-domain images, audio, video, and other media. It is useful for researching historical media and digital content.
95. Wikisource
Wikisource is a multilingual digital library containing transcribed and digitized historical texts and source documents.
96. Wikiquote Archives
Wikiquote provides a structured repository of quotations and historical references. Its revision history can also be useful when researching how information on the site changed over time.
97. Wikipedia Page History
Wikipedia’s revision-history system provides an extensive record of changes made to individual encyclopedia pages. It is not a conventional web crawler archive, but it can preserve historical versions of encyclopedic content.
98. GitHub
GitHub repositories can provide historical versions of websites, source code, documentation, and digital projects through version-control histories, releases, forks, and commits.
99. GitLab
GitLab provides repository version histories that can preserve earlier versions of websites, applications, documentation, and digital projects.
100. Software Heritage
Software Heritage is a major digital preservation initiative focused on preserving source code. It creates a long-term archive of software source code from numerous publicly available sources.
Categories of Internet Archiving Platforms
The 100 resources above can broadly be divided into several categories.
1. General Web Archives
These preserve or provide access to historical versions of websites.
Examples include:
- Internet Archive / Wayback Machine
- Archive.today
- Archive.ph
- Perma.cc
- Common Crawl
- Memento
2. National Web Archives
National libraries and government organizations preserve websites associated with particular countries.
Examples include:
- UK Web Archive
- Arquivo.pt
- National Library of Australia Web Archive
- WARP
- Swiss Web Archive
- Danish Web Archive
- Finnish Web Archive
- New Zealand Web Archive
3. Institutional Web Archives
Universities, libraries, museums, and research institutions maintain collections of websites relevant to their missions.
Examples include:
- Harvard Library Web Archive
- Stanford Web Archive
- Yale University Web Archive
- Library of Congress Web Archives
- Columbia University Web Archive
4. Digital Libraries
Some services are broader digital libraries rather than traditional website crawlers.
Examples include:
- Europeana
- Digital Public Library of America
- HathiTrust
- Google Books
- Gallica
- Internet Archive Texts
5. Web-Archiving Technology
Some platforms provide the technical infrastructure used to capture, process, search, and replay archived websites.
Examples include:
- Heritrix
- PyWB
- OpenWayback
- Webrecorder
- Conifer
- Browsertrix
- ArchiveWeb.page
- ArchiveSpark
6. Newspaper and Historical Media Archives
These platforms preserve older newspapers, books, audio, video, and other historical media.
Examples include:
- Newspapers.com
- NewspaperArchive
- Google News Archive
- Chronicling America
- Internet Archive Movies
- Internet Archive Audio
7. Software and Code Archives
Version-control and software-preservation services can preserve historical versions of websites and applications.
Examples include:
- GitHub
- GitLab
- Software Heritage
- Internet Archive Software Collection
Why Website Archiving Matters
Website archiving is important because online information can disappear very quickly. Companies close, organizations redesign their websites, governments replace old pages, news organizations remove older stories, domain names expire, and social-media posts can be deleted.
Archived websites can therefore be useful for:
- Historical research
- Academic research
- Journalism
- Digital humanities
- Legal research
- Investigative research
- Business history
- Competitive research
- Brand-history research
- Fact checking
- Citation preservation
- Researching discontinued services
- Studying changes to government websites
- Recovering information from discontinued websites
- Studying historical internet culture
Important: An archived page should be treated as a historical record rather than automatically as proof that every statement appearing on that page was accurate. Researchers should evaluate the original source, date, authorship, context, and reliability of the information.
How to Search Archived Websites
When researching an old website, start with the website’s original domain rather than only searching for individual page titles.
A useful workflow is:
- Identify the original domain.
- Search the domain in a major web archive.
- Check the available capture dates.
- Open several snapshots rather than relying on one.
- Look for pages surrounding the relevant date.
- Check whether images and downloadable documents were preserved.
- Search national or institutional archives.
- Compare archived information with contemporary sources.
- Record the capture date and archive used.
- Preserve the archive reference when citing the material.
Website Archiving and SEO Research
Archived websites can also be valuable for search-engine optimization research. Historical snapshots can reveal previous site structures, URLs, navigation systems, content strategies, title tags, page layouts, internal links, and changes to a company’s online presence.
For example, an SEO researcher can use historical captures to investigate:
- Previous website structures
- Old landing pages
- Deleted service pages
- Historical keyword targeting
- Previous business descriptions
- Old contact information
- Former product pages
- Discontinued services
- Historical internal-link structures
- Domain migrations
- Website redesigns
However, archived content should be cross-checked before being treated as current information.
Website Archiving for Digital Preservation
Digital preservation is broader than simply taking screenshots of web pages. A reliable preservation system may need to retain HTML, CSS, JavaScript, images, documents, metadata, timestamps, HTTP responses, redirects, and other technical information.
Modern web applications can be particularly difficult to archive because they may depend on:
- APIs
- Authentication
- Databases
- JavaScript applications
- Streaming services
- Dynamic content
- Geolocation
- Cookies
- Third-party services
- Embedded social-media content
Consequently, no single archiving service can guarantee that every component of every website will remain available indefinitely.
The internet contains an enormous amount of historical information, but online content is inherently vulnerable to modification and disappearance. Website archiving platforms help preserve parts of that digital history for researchers, journalists, businesses, governments, academics, librarians, historians, and the general public.
The Internet Archive and Wayback Machine remain among the most recognizable resources for examining historical web pages, while services such as Common Crawl, Perma.cc, Archive.today, Webrecorder, national web archives, university archives, and specialist digital libraries provide complementary capabilities.
For comprehensive research, it is often better to combine several archival sources rather than depend on one platform. Different archives have different crawl schedules, geographic coverage, collection policies, technical capabilities, and retention practices. Using multiple sources can therefore provide a more complete picture of how a website or online resource existed at a particular point in time.