Semantic Scholar (semanticscholar.org) is a free academic search engine operated by the nonprofit Allen Institute for AI (Ai2), built for researchers who need to screen literature quickly and work with citation data. It runs natural-language processing models over the papers it indexes: paper pages and search results carry a one-sentence TLDR summary, each citation is classified by intent as background, methodology or result extension, and a machine-learning model flags highly influential citations. Compared with general academic search engines such as Google Scholar, its verifiable differences are that it publishes its corpus size and data sources (the homepage search box showed 237,612,676 papers as of 2026-08-31) and offers a free API with downloadable datasets, on top of which tools like Connected Papers and Litmaps have been built (API overview).

Site at a glance

  • URL: https://www.semanticscholar.org/
  • Type: Academic search engine and scholarly-graph data platform
  • Cost: Free; API keys are requested through a free form (API overview)
  • Registration: Searching, browsing paper pages and downloading datasets require no account; a free account adds the library, email alerts, research feeds and author-page claiming (FAQ)
  • Interface language: English; the corpus also focuses on publications primarily in English (FAQ)

Background

Ai2 is a nonprofit research institute founded in 2014 by Microsoft co-founder Paul Allen, and Semantic Scholar launched in 2015 as one of its research engineering projects (About). The team develops its paper-processing and classification models in house, releases open code and datasets, and publishes research in natural language processing, machine learning, human-computer interaction and information retrieval; the About page lists promoting equal access to science among the institute's values.

For content, the site states that it indexes more than 200 million papers sourced from publisher partnerships, data providers and web crawls (About). Its direct partners number more than 50 publishers, data providers and aggregators, covering 500+ academic journals, university presses and scholarly societies (Publishers); sources named in the FAQ include PubMed, arXiv and Springer Nature (FAQ). The API page puts the graph at 214 million papers, 2.49 billion citations and 79 million authors (as of 2026-08-31, API overview).

The Semantic Scholar homepage, with the search box showing the live count of indexed papers

AI-driven features

TLDR summaries. The TLDR shown on paper pages and in search results is a model-generated single sentence covering a paper's objective and findings, intended for triage before reading the abstract. The feature is in beta and covers nearly 60 million papers in computer science, biology and medicine (TLDR feature page); the underlying research is "TLDR: Extreme Summarization of Scientific Documents". The tldrs subset of the datasets currently holds about 58 million records (as of 2026-08-31, datasets release endpoint).

The Semantic Scholar TLDR feature page, with one-sentence summary examples from computer science and biomedicine

Citation analysis. Citation lists label each citation's intent: background citations give historical context or justify importance, methodology citations reuse established procedures, and result citations build on earlier findings (FAQ). Highly influential citations are flagged by a machine-learning model that weighs citation counts and surrounding context to find cases where the cited work materially shaped the citing paper; the site notes the determination depends on access to the citing paper's full text and may miss cases when it is unavailable (FAQ). On counts, the site states that its corpus focuses on published articles and preprints — book coverage is very limited and patents are excluded — so counts can differ from other sites (FAQ).

Search and ranking. Results can be sorted by relevance, citation count, most influential papers or recency; Boolean operators and wildcards are not supported, while quoted phrases are (FAQ). The search reranker is open-sourced in the allenai/s2search repository (GitHub). Fields of study are assigned by a machine-learning classifier over each paper's title and abstract, up to three per paper, and only for English-language papers (FAQ).

Author profiles and disambiguation. Author pages are generated automatically from public sources and publisher partners, with same-name authors separated by the S2AND disambiguation model introduced in June 2021; researchers can claim their page to add or remove papers and update affiliation, ORCID and homepage details (FAQ).

Generative AI features and accuracy caveats. Ask This Paper queries OpenAI's gpt-3.5-turbo-16k with the paper's content to answer questions, and term definitions in the reader and Topic pages are also generated by large language models. The FAQ carries an explicit warning for these features: generated text can contain factual errors that are hard to detect, awkward phrasing or irrelevant information, and users should verify accuracy where possible (FAQ). Topic pages currently cover computer science only, and the augmented Semantic Reader is a beta feature limited to arXiv-hosted papers, with 250k+ papers processed according to the site (FAQ).

Accounts and open data

Accounts. Free accounts can be created with an email address, Google or an institutional identity, with institutional sign-in provided through OpenAthens, eduGAIN and InCommon; once signed in, GetFTR and LibKey integrations link directly to paywalled articles an institution subscribes to (FAQ). Account features include a library with shareable folders, author/paper/topic email alerts, Research Feeds recommendations and author-page claiming, plus Chrome and Firefox browser extensions (FAQ).

API and datasets. The Academic Graph API serves papers, authors, citations, venues and SPECTER2 embeddings, alongside a recommendations API and a datasets API (API overview, API documentation). Most endpoints work without authentication, with unauthenticated requests sharing a pool of 1000 requests per second across all anonymous users and subject to further throttling under heavy load; an individual API key starts at an introductory rate of 1 request per second, with higher limits granted after review (API overview).

The full corpus is released periodically as S2AG datasets covering papers, abstracts, citations, authors, TLDRs, S2ORC full text and SPECTER embeddings (datasets documentation); the latest release was 2026-08-25 as of 2026-08-31, with releases rolling roughly weekly (release list). S2ORC is a corpus designed for NLP and text-mining research, providing structured parses of open-access full text as JSON (About, GitHub). The site lists these datasets as free and open resources, asks that published research cite them, and governs API use under a separate license agreement (API overview, API License Agreement).

The Semantic Scholar API overview page, with corpus statistics and API key rate-limit documentation

Third-party literature tools including Connected Papers, Litmaps and Sourcely describe their services as built on this API and its datasets; one developer notes that metadata such as PDF links, abstracts and automatic summaries is not easily accessible from other academic reference providers like Google Scholar (API overview).

Use cases

  • Deciding quickly whether a paper deserves a close read: TLDRs, citation intent and influence labels offer signals before the abstract.
  • Tracing how a method or finding propagated: intent labels separate "reused the method" from "extended the result".
  • Building literature-driven applications or studies: the free API and downloadable datasets cover papers, authors, citations and TLDRs as structured data.
  • Maintaining a scholarly profile: claiming an author page lets researchers fix misattributions from automatic disambiguation.
  • Reaching paywalled full text through institutional sign-in where a subscription exists.

Limitations

  • Uneven disciplinary and language coverage: TLDRs cover computer science and biomedicine only (TLDR feature page), Topic pages cover computer science only, field-of-study classification is English-only, and the corpus focuses on English publications (FAQ).
  • Counting differs from other platforms: book coverage is minimal and patents are excluded, so citation counts and h-index values may not match other sites; the site explicitly advises against using h-index for comparative research assessment (FAQ).
  • Influential-citation detection depends on full text of the citing paper and can miss cases when it is unavailable (FAQ).
  • Search is limited: no Boolean operators or wildcards (FAQ).
  • AI-generated content can be wrong; the site's own warning asks users to verify generative features (FAQ).
  • The interface is English-only, there is no mobile app, and access is via the responsive site and browser extensions (FAQ).
  • Paywalled full text still requires an institutional subscription or purchase from the publisher; PDFs are provided only where an open-access version exists (FAQ).

Similar services

  • Google Scholar: the closest general-purpose academic search engine; it publishes no corpus size and offers no comparable public data interface.
  • Scopus / Web of Science: commercial citation databases that require institutional subscriptions and are commonly used for formal citation reporting and assessment.

References