arXiv (arxiv.org) is a free distribution service and open-access archive for scholarly preprints, covering eight subject areas: physics, mathematics, computer science, quantitative biology, quantitative finance, statistics, electrical engineering and systems science, and economics. Researchers can make papers public before formal peer review, and anyone can read and download the full texts without registration. Its About page states that it hosts more than three million scholarly articles (verified 2026-09-01). Unlike journal platforms, arXiv makes it explicit that materials on the site are not peer-reviewed by arXiv; its value lies in distributing new research with minimal barriers and minimal delay, which has made it the de facto first-stop venue in fast-moving fields such as high-energy physics, mathematics and machine learning.
At a Glance
- URL: https://arxiv.org/
- Type: Open-access scholarly preprint archive
- Cost: Free to read and download; article submission is also free
- Registration: None required for reading; submitting requires a registered account and passing the endorsement process
- Interface language: English; since February 2026 all submissions must include a full English-language version
Background
arXiv was founded in 1991 by physicist Paul Ginsparg as an email-based service running on a NeXTstation under his desk at Los Alamos National Laboratory, receiving about ten submissions a month; today it averages around 24,000 new submissions per month, with recent months frequently topping 30,000, and takes nearly 30 staff to operate (spin-out FAQ). According to the Cornell Chronicle, arXiv passed three million articles in April 2026 — just four years after reaching two million, compared to the 23 and a half years it took to reach the first million.
Institutionally, arXiv was hosted by Cornell University for decades (under the University Library, then the Faculty of Computing and Information Science, and finally Cornell Tech). On July 1, 2026, it spun out from Cornell to become an independent nonprofit, arXiv, Inc., incorporated in Delaware and granted IRS 501(c)(3) tax-exempt status, with the Simons Foundation and Cornell University as founding members. It remains headquartered on the Cornell Tech campus for the near term, and its inaugural CEO, Dr. Penelope Lewis, took office in August 2026 (spin-out FAQ, arXiv blog). Operations are funded by Simons Foundation International, a membership program for institutions, and individual donors; the Cornell Chronicle also reports that $10 million in gifts from the Simons Foundation and Schmidt Sciences is supporting the completion of its cloud migration and code modernization.
A few well-known cases illustrate arXiv's place in scholarly communication: Juan Maldacena's 1997 paper introducing the AdS/CFT correspondence and Grigori Perelman's 2002 proof of the Poincaré conjecture were both first released on arXiv (Cornell Chronicle).
Browsing and Reading
The homepage offers site-wide search plus a hierarchical browse structure: the eight subject areas break down into hundreds of subcategories (for example, cs.AI for artificial intelligence and cs.LG for machine learning under computer science), each with new, recent and search entry points.

Every paper has an abstract page listing its title, authors, abstract, subject classifications, submission history (with timestamps and sizes for each version), journal reference and DOI. A sidebar offers PDF and TeX source downloads, license information, BibTeX citation export, and links to external citation tools such as INSPIRE HEP, NASA ADS, Google Scholar and Semantic Scholar.

New papers appear on a fixed schedule: submissions received by 14:00 US Eastern on a working day are generally announced at 20:00 the same day, Sunday through Thursday, with no announcements on Friday or Saturday (announcement schedule). Each category's "recent" listing groups the latest entries by day and is a common way for researchers to track their field.

Submission and Moderation
Submission is free, but several requirements are worth knowing:
- Endorsement. First-time submitters to a category need endorsement for that area. Since January 21, 2026, an institutional email address alone is no longer sufficient: new authors must either combine an institutional email with authorship of an already-accepted paper in the target endorsement domain, or obtain personal endorsement from an established arXiv author in the field (endorsement policy update, endorsement help).
- Moderation. All submissions are screened by volunteer moderators — subject experts with terminal degrees — who check topicality and scholarly value and may reclassify or decline a submission. This is not peer review, and moderators give no feedback (moderation policy). Quality-assurance checks typically take one to four days before announcement.
- Formats. (La)TeX source is preferred, followed by PDF and HTML; scanned documents are not accepted (submission guidelines).
- English requirement. Since February 11, 2026, every new submission must include a full English-language version, either as the original or as a translation; machine translation is explicitly permitted if faithful to the original. Previously only an English abstract was required (policy announcement).
- Tightened rules against generative-AI abuse. Since late October 2025, review articles and position papers in the computer science category must first be accepted by a journal or conference and have completed peer review (CS category announcement). The moderation policy also requires authors to report significant use of generative AI tools in the paper itself, forbids listing AI tools as authors, and holds named authors fully responsible for content.
Once announced, a paper becomes part of the permanent scholarly record: it can be marked as withdrawn but not deleted, and the license chosen at submission is irrevocable.
Copyright and Licensing
Submitters grant arXiv distribution rights under one of six licenses: CC BY 4.0, CC BY-SA 4.0, CC BY-NC-SA 4.0, CC BY-NC-ND 4.0, the arXiv.org perpetual non-exclusive license 1.0, and CC Zero. Except for CC0, the original copyright holder retains ownership after posting (license information). Note that most papers carry the default arXiv non-exclusive license, which lets arXiv distribute the work but does not let arXiv grant reuse rights to others — anyone building indexes or tools on the full texts must link back to arXiv for downloads (bulk data access). All article metadata is released under CC0.
APIs and Bulk Data
arXiv offers several layers of programmatic access, all governed by its API terms of use:
- arXiv API: an HTTP query interface returning metadata and search results in Atom 1.0, with combined filters by category, author, keywords and more (API documentation).
- OAI-PMH: a metadata harvesting protocol updated daily, the preferred way to bulk-download or synchronize metadata.
- Bulk full text: the complete machine-readable dataset is hosted on Kaggle, and PDFs plus source files for all articles are available from Amazon S3.
- RSS: daily feeds of new submissions per category.
Programmatic access is directed to the dedicated export.arxiv.org site, with a suggested rate of bursts of four requests per second followed by a one-second pause; programmatically downloading the entire corpus is explicitly discouraged (bulk data access). Third-party projects must not use the arXiv name or branding in ways that imply endorsement.
Good For
- Publishing results months or years before formal publication to establish priority, or keeping up with the latest work in a field.
- Reading full texts at no cost, especially for readers without journal subscriptions.
- Bibliometrics, information retrieval and NLP research using the API, OAI-PMH feeds or the Kaggle/S3 datasets.
- Monitoring daily new submissions in specific subcategories as a literature-alert channel.
Limitations
- No peer review: moderation only screens for topicality and basic scholarly value. Correctness is the authors' responsibility, readers must judge quality themselves, and citations should acknowledge the preprint status.
- Limited scope: only eight subject areas are covered; work outside the classification scheme cannot be submitted.
- English-centric: the interface is English-only, and since February 2026 every submission must include a full English version.
- Endorsement barrier for new authors: first-time submitters without both an institutional email and existing papers, and without academic contacts who can vouch for them, may find endorsement slow.
- Fixed announcement rhythm: submissions go through one to four days of checks, and nothing new is announced on Fridays or Saturdays — publication is not instant.
- Permanent record: announced papers and chosen licenses cannot be revoked; papers can only be marked withdrawn.
Alternatives
- bioRxiv and medRxiv: preprint servers for biology and health sciences operated by Cold Spring Harbor Laboratory, complementing arXiv's subject coverage.
- HAL: France's national multidisciplinary open archive, likewise supporting author self-archiving.
References
- arXiv.org e-Print archive (homepage) (checked 2026-09-01)
- About arXiv (checked 2026-09-01)
- arXiv is now an independent nonprofit (spin-out FAQ) (checked 2026-09-01)
- Digital research repository arXiv to start new chapter as nonprofit - Cornell Chronicle (checked 2026-09-01)
- arXiv Submission Guidelines (checked 2026-09-01)
- The arXiv endorsement system (checked 2026-09-01)
- Attention Authors: updated endorsement policy - arXiv blog (checked 2026-09-01)
- arXiv moderation (checked 2026-09-01)
- arXiv License Information (checked 2026-09-01)
- Upcoming policy change to non-English language paper submissions - arXiv blog (checked 2026-09-01)
- Updated Practice for Review Articles and Position Papers in arXiv CS Category - arXiv blog (checked 2026-09-01)
- Submission availability / announcement schedule (checked 2026-09-01)
- arXiv API Basics (checked 2026-09-01)
- Bulk Data Access to arXiv (checked 2026-09-01)






