Internet Archive is a US non-profit digital library founded in 1996 by engineer Brewster Kahle with the goal of providing universal access to published knowledge. Its best-known service is the Wayback Machine, a historical database that continuously crawls and archives public web pages; by the organization's own count it holds more than one trillion page snapshots (as of 2026-08-30). Beyond web archiving, it runs a set of free projects covering digital book lending, media collections, and historical software that runs in the browser.

At a glance

  • URL: https://archive.org
  • Type: Non-profit digital library (501(c)(3))
  • Cost: Free to browse, download, and borrow; Archive-It is a subscription service for institutions; donations accepted
  • Registration: Not required for browsing or downloads; a free account is needed to borrow ebooks, use Save Page Now, or upload material
  • Interface language: Primarily English

Background

Internet Archive was founded in San Francisco in 1996 and began by crawling and preserving web pages. According to its official help center, the organization's mission is "universal access to all knowledge," and its collections gradually expanded into texts, audio, video, software, and images. The main site archive.org serves as the unified entry point, while individual features live in sub-projects such as Open Library, Live Music Archive, and TV News Archive.

Collections and services

  • Wayback Machine: Look up historical snapshots of any public web page by URL and compare versions over time. The accompanying Save Page Now tool lets you manually archive a live page. There are also iOS and Android apps, browser extensions, and Archive-It, a subscription crawling service for libraries, governments, and other institutions.

    The Wayback Machine's snapshot search interface, where entering a URL lists historical captures of that page

  • Texts and ebooks: Digital lending through Open Library, including books scanned in partnership with US libraries and a large body of public-domain works, readable online or borrowed.

  • Audio: The Live Music Archive holds recordings of tradeable live performances, alongside digitized 78 rpm discs and LibriVox audiobooks.

  • Video: The TV News Archive collects television news clips with closed-caption search, useful for fact-checking.

  • Software: Collections such as MS-DOS games, the Internet Arcade, and Console Living Room run directly in the browser without a local emulator.

  • Images: Material collections including NASA imagery and open images from the Metropolitan Museum of Art.

Accounts and open capabilities

A free account, registered with an email address, unlocks borrowing, favoriting, uploading, and page archiving. The site exposes metadata and full-text search APIs for programmatic access to collections. Uploading is open to registered users, who can contribute their own publications, podcasts, and other material.

Copyright and content policy

The collections consist mainly of public-domain material, institution-licensed content, and user uploads. Ebook lending follows a "controlled digital lending" model: loanable digital copies are capped to match physical holdings. In 2020, four publishers including Hachette sued Internet Archive over this model; the district court ruled for the publishers in March 2023, and the US Court of Appeals for the Second Circuit affirmed on September 4, 2024. According to the organization's official blog, it has removed more than 500,000 publisher books from lending as a result and has sought Supreme Court review. The case continues to shape what is available.

When it is useful

  • Recovering old versions of pages that were redesigned, taken offline, or deleted
  • Borrowing or searching for out-of-print and public-domain books and audiobooks
  • Searching TV news clips for fact-checking and media research
  • Running historical software and arcade games in the browser
  • Snapshotting public pages, or citing a stable link to a page's history

Limitations

  • Because of the publishers' lawsuit, most in-copyright books are no longer lendable; some remain available only as 1-hour loans or restricted previews.
  • The web archive is not a full mirror: crawl frequency varies widely by site, and many pages were never captured.
  • Features are spread across sub-projects with inconsistent search experiences; the interface is English-only.
  • Borrowing requires an account, and popular titles can involve waitlists.

References