The Chinese Text Project (CTP, 中國哲學書電子化計劃) is an open-access digital library of premodern Chinese texts: anyone can search and read transmitted Chinese literature from the pre-Qin era onward in a browser, without an account. It was created in 2006 by Donald Sturgeon, now an Assistant Professor in the Department of Computer Science at Durham University, who still designs, edits, and maintains the site himself. Its homepage states that it holds "over thirty thousand titles and more than five billion characters," making it — by its own account — the largest database of premodern Chinese texts in existence (as of 2026-09-02). Unlike platforms built around page scans, CTP is built on structured, searchable full text, surrounded by research tools: paragraph-aligned Chinese-English parallel texts, a dictionary wired into the corpus, a text-reuse ("parallel passage") database, and side-by-side scanned source images.

Chinese Text Project homepage: a categorized table of contents on the left, and announcements about newly added AI translations in the main column

The screenshot above shows the ctext.org homepage. The left column organizes pre-Qin and Han texts by school of thought (Confucianism, Mohism, Daoism, Legalism, School of Names, School of the Military, Mathematics, Miscellaneous Schools, Histories, Ancient Classics, Etymology, Chinese Medicine, and Excavated texts), post-Han texts by dynasty, and links to the Dictionary, Discussion forum, Library, Wiki, and Data Wiki. The announcements column notes that AI-generated English translations have been added at scale since March 2026.

At a glance

  • URL: https://ctext.org/ (the FAQ lists ctext.cn as an alternative address for users who hit DNS resolution problems)
  • Type: digital library / full-text database of premodern Chinese texts
  • Cost: free and open access; academic institutions can subscribe for bulk download, API keys, and early features; donations accepted
  • Registration: not required for reading and searching; a free account unlocks plugins, Wiki editing, and the forum
  • Interface languages: English and Chinese, with one-click switching between traditional and simplified characters (all data is stored in traditional characters and converted automatically)
  • Launched: 2006 (source: Wikipedia)
  • Maintainer: Donald Sturgeon (德龍), Durham University

Background

Donald Sturgeon founded CTP in 2006. With a background in classical Chinese philosophy and computing, he held postdoctoral positions at Harvard University before joining Durham University's Department of Computer Science, where he works on natural language processing for premodern Chinese and digital libraries (sources: Fairbank Center profile, World Conference on Sinology bio). According to the official FAQ, the site is "designed, edited, and maintained" by him alone.

In October 2016, the Harvard-Yenching Library contributed over five million scanned pages from its Chinese collections — including the Chinese Rare Books Collection — to CTP's Library section. CTP generated approximate transcriptions with its own OCR pipeline and added them to the Wiki, making the images full-text searchable (homepage announcement).

For academic citation, the site asks to be cited as Sturgeon's 2019 paper "Chinese Text Project: a dynamic digital library of premodern Chinese" in Digital Scholarship in the Humanities (FAQ citation guide).

Core features

Structured full text and parallel translations

Texts are organized as book → chapter → paragraph, with stable paragraph numbers that can be linked directly. Many core pre-Qin and Han texts carry English translations — notably copyright-expired versions such as James Legge's — aligned paragraph by paragraph with the original; some texts also have modern Chinese translations.

The Analects, "Xue Er": each paragraph of the Chinese original is followed by James Legge's English translation; the left column lists the book's chapters

The screenshot shows the Analects "Xue Er" page: every paragraph of the original is followed by Legge's translation, and the paragraph number on the left expands into actions such as "Jump to dictionary," "Show parallel passages," and "Show commentary." The chapter list in the sidebar covers the whole Analects alongside other Confucian texts, and the translator is credited at the top of the page, where translations can also be switched off.

Since March 2026, CTP has added AI-generated English translations for pre-Qin and Han texts that previously had none, for the twenty-five dynastic histories, and for hundreds of other historical, literary, and philosophical works. All AI translations are sentence-aligned with the source so errors can be spotted and corrected through the dictionary function. The same release added editable Chinese and English summaries of historical entities (people, works, office titles) in the Data Wiki — over 100,000 so far (homepage announcement).

A dictionary wired into the corpus

The built-in dictionary connects headwords to the entire corpus. Looking up a character returns quotations from the Guangyun and Kangxi dictionaries, Unihan data, CC-CEDICT, and the Revised Mandarin Chinese Dictionary, but the CTP dictionary itself goes further: under each sense it lists real example sentences from the corpus in which the character carries that meaning, with their English translations. Entries include seal-script, bronze, and oracle-bone glyph images, fanqie spellings, a Tang-era phonological reconstruction, Cantonese readings, and page references to works such as the Hanyu Da Zidian and Grammata Serica Recensa (source: dictionary entry for 學).

Dictionary entry for 學: ancient glyph images at the top, Guangyun and Kangxi quotations in the middle, and corpus example sentences under each CTP dictionary sense at the bottom

Text reuse and concordance tools

The tools list includes a "parallel passage database" that algorithmically identifies text reuse across early Chinese literature — clicking "Show parallel passages" on any paragraph reveals where else the same wording appears — and concordance/index tools that locate passages by standard reference numbers or vice versa. These collation-oriented features are what separates CTP from a plain reading site.

Scanned images and the Library section

The Library section hosts photographic reproductions of preserved editions (uploads are accepted only in PDF or DjVu). Page images are linked line by line to transcriptions, and text pages can be checked against the "digital base text" scan. The site is explicit that naming a base text states an editorial intention, not a guarantee: the digital edition may diverge from it, and readers should compare the base-text image before quoting (FAQ).

Library page for the Qianlong Imperial Siku Quanshui Huiyao edition of the Records of the Three Kingdoms, Wu section: metadata and chapter links above, scanned page images below

The screenshot shows the Library page for a Siku Quanshu Huiyao edition of the Sanguozhi Wu chapters: metadata credits Zhejiang University Library and CADAL as the scan source, a table offers download sources and "jump to chapter" links into the Wiki transcriptions, and the scanned pages appear below.

Accounts, subscriptions, and the API

  • Free account: unlocks plugins (such as "Plain text," which lets you copy or download any chapter), direct correction of OCR errors in the Wiki, and forum participation. The FAQ notes that accounts registered with institutional email addresses (.edu, .ac.uk, etc.) start with higher trust and see fewer verification challenges, and that constructive Wiki edits raise trust further.
  • Institutional subscription: aimed at university libraries. Benefits include extended API functions (programmatic access to text structure, machine-readable download of entire works), API keys for research and teaching, and early access to new features. As of 2026-09-02, the subscription page lists about twenty subscribing institutions, including Harvard, Oxford, Stanford, and Academia Sinica.
  • API and plugins: a JSON API at api.ctext.org addresses texts by URN (e.g. ctp:analects/xue-er). Anonymous calls are rate-limited, logged-in users get more, and subscribers more still; an official Python client library and tutorials are available. A plugin system lets third parties attach external tools to the CTP interface with an XML descriptor (API documentation).

Copyright and content policy

CTP is not a simple mirror of public-domain texts, and its layered terms deserve attention (FAQ, copyright section):

  • The website and its content are protected under international copyright law and may not be republished without written permission. Reasonable use is encouraged, but automated bulk downloading of pages is explicitly prohibited.
  • The ancient source texts are mostly public domain, but translations remain under the copyright of their translators (except expired-copyright works such as Legge's), and translators must be credited. Some site content is public domain; determining what may be reproduced is the user's responsibility.
  • Uploading to the Library requires affirming that the material is public domain or authorized by the rights holder, and grants CTP a perpetual, irrevocable license to publish it and create derivative works. Contributing to the Wiki requires agreeing to transfer the copyright of the contribution to CTP irrevocably. User-submitted content is not editorially reviewed, and responsibility rests with the submitter.

When it is useful

  • Tracing a quotation in pre-Qin and Han literature to a stable, paragraph-level link
  • Reading the original with sentence-level English translations and the integrated dictionary, or doing Chinese-English collation
  • Studying textual transmission and intertextuality with the parallel-passage data
  • Pulling corpus data through the API or Python library for digital humanities research
  • Classroom comparison of woodblock scans against typeset transcriptions (the site asks that regular teaching users encourage their university library to subscribe)

Limitations

  • Accuracy is the reader's responsibility: naming a digital base text does not guarantee the transcription always follows it, and the site asks readers to check quotations against the scan. Wiki texts come from OCR and crowdsourced transcription without editorial review.
  • AI translations are aids, not authorities: the site's own announcement warns they "will inevitably contain mistakes" and should be checked sentence by sentence.
  • Heavy anonymous use is throttled: verification images or login prompts appear after sustained automated-looking traffic; bulk download and full API access require a subscription.
  • Redistribution is restricted: translations belong to their translators, and bulk scraping or republishing is prohibited — a different model from freely licensed projects like Wikisource.
  • Dated but functional design: the desktop layout is plain (a mobile layout can be toggled at the top left), and rare characters in Unicode CJK Extensions A–D require a large-coverage font such as the free Hanazono typeface to display correctly (homepage note).
  • Coverage centers on pre-Qin and Han: post-Han material is organized by dynasty but its depth depends largely on user contributions to the Wiki and Library.

Comparable sites

  • Chinese Wikisource: a volunteer-built library of public-domain texts under a free license that permits lawful reuse — a deliberate contrast to CTP's all-rights-reserved model.
  • Shidian Guji (识典古籍): a free Chinese platform for reading ancient texts online.
  • Kanseki Repository (漢籍リポジトリ): an open project hosting full texts of Chinese classics in GitHub repositories, classified into the six traditional divisions (classics, histories, masters, collections, Daoist, Buddhist), suited to bulk access to machine-readable text.

References

  • ctext.org homepage (scale claims, AI translation and entity summary announcements, Harvard-Yenching collaboration, font requirements; checked 2026-09-02)
  • Official FAQ (maintainer, citation format, translation sources, digital base texts, copyright and Library/Wiki terms, alternative domain; checked 2026-09-02)
  • Tools list (text database, parallel passage database, dictionary, concordance and index tools, plugins; checked 2026-09-02)
  • CTP API documentation (JSON API, URNs, rate limits, Python client; checked 2026-09-02)
  • Institutional subscription (benefits and subscriber list; checked 2026-09-02)
  • Donald Sturgeon - Fairbank Center (Assistant Professor of Computer Science at Durham University; checked 2026-09-02)
  • Chinese Text Project - Wikipedia (2006 launch, registration required to contribute; checked 2026-09-02)