Shidian Guji (识典古籍, shidianguji.com) is a free digital reading and editing platform for guji — pre-modern Chinese books written mostly in Classical Chinese, from Confucian classics and dynastic histories to Buddhist and Daoist canons. It is developed and operated by Douyin Group (the ByteDance affiliate behind the Chinese version of TikTok) in partnership with Peking University, and launched in beta in October 2022. Rather than hosting page scans alone, the platform runs an AI pipeline — OCR, automatic punctuation and named-entity recognition — to turn those scans into searchable, punctuated digital text, which readers can view side by side with images of the original woodblock editions. As of 2026-09-01, the homepage states 73,602 titles are available; a Beijing Daily report (August 2026), citing Peking University, puts monthly readership above 2.4 million. Its verifiable distinctions are the industrial-scale AI editing pipeline, a public two-tier text-quality label ("carefully proofread" vs. "machine-processed"), a companion editing platform opened free of charge to public-interest projects, and a large crowdsourced proofreading program.

At a Glance
- URL: https://www.shidianguji.com/
- Type: Digital reading, full-text search and editing platform for pre-modern Chinese books
- Cost: Entirely free, no advertising; the official account describes it as "free forever, with no ads" (in Chinese)
- Registration: Browsing, reading and full-text search need no account; bookshelf, notes, text copying, the "deep research" AI feature and the editing platform require sign-in (a ByteDance account, Chinese phone number recommended)
- Interface language: Primarily Simplified Chinese; an English interface exists and some catalog entries carry English titles; machine conversion between simplified and traditional characters is built in
- Mobile: Standalone iOS and Android apps launched in 2025; the web version keeps the page-image comparison view
Note for non-Chinese readers: although parts of the interface are available in English, the content itself is almost entirely Classical and modern Chinese, and account registration is oriented toward mainland-Chinese users. The platform is most useful if you can read Chinese or work with Chinese source texts.
Background
Classical Chinese texts have long faced a digitization gap. According to a 2022 ByteDance Philanthropy announcement (in Chinese), of roughly 200,000 surviving pre-modern Chinese works, about 80,000 had been scanned as images but only 30,000–40,000 existed as machine-readable text — and scanned images cannot be searched full-text. Per the Peking University Digital Humanities Center project page (in Chinese), Douyin Group donated to the Peking University Education Foundation in March 2022, and the two sides built the platform through a joint "Digital Humanities Open Laboratory": the university supplies design guidance, source page images and expert proofreading, while ByteDance funds, develops and operates the platform. The site's About page notes that copyright in the source materials provided by Peking University remains with the university.
Reported collection size has grown quickly, and figures depend on their date:
- October 11, 2022 (beta launch): 390 titles, mostly from the Sibu Congkan series (Xinhua, in Chinese)
- Late 2023: over 2,200 titles (Peking University News, in Chinese)
- End of 2024: 10,000 titles; first half of 2025: over 20,000 (ByteDance Philanthropy, in Chinese)
- August 2026: over 70,000 titles free to read (Beijing Daily, in Chinese); as of 2026-09-01 the homepage reports 73,602 titles
Catalog and Text-Quality Labels
The library is organized by the traditional Chinese bibliographic classification — Classics (经), Histories (史), Masters (子) and Literary Collections (集) — extended with sections for Daoist and Buddhist texts, excavated documents and archival records. Each entry shows the title, number of volumes, author, dynasty, the physical edition it derives from (for example, a Song-era imprint reproduced in Sibu Congkan) and the text source, such as the "Peking University–ByteDance open text corpus" or the Ru Zang (Confucian Canon) project.

Every text carries a quality badge. "精校" (carefully proofread) means characters, punctuation and entity markup were all reviewed by humans; "粗校" (roughly proofread) or "AI整理" (AI-processed) means the text is machine output, and the page reminds readers that it "should be read together with the original images." For citation purposes this label is the first thing to check.
Full-Text Search
Full-text search works without an account, across either the whole corpus ("search text") or book metadata ("search books"). Results can be filtered by classification and dynasty, sorted by relevance or period, and refined with options like "body text only," "original characters only" and fuzzy search; variant character forms are merged in the display. Each hit cites the book, chapter, edition and page number, so a passage can be traced back to the physical source. A corpus-analysis entry can also turn search results into a table for further analysis.

Reading and Reference Tools
The reader presents horizontally typeset text with modern punctuation, a chapter outline on the left, and an optional side-by-side view of the original page images — traceability to a physical base edition is a stated design principle. Built-in tools include toggling commentary layers, machine conversion between traditional and simplified characters (labeled "machine simplified, for reference only"), hover glosses for rare words, clickable encyclopedia entries for person and place names marked in the text, an entity-relationship graph and historical maps (the site notes some of this data is machine-generated and may contain errors).

On the AI side, the platform says it has machine-translated 10,000 titles into modern vernacular Chinese (Beijing Daily, August 2026); pages state the translation is "machine translation checked by humans, for reference only." An AI assistant based on ByteDance's Doubao model answers questions with a visible "AI-generated, may contain errors" notice; its "deep research" mode requires sign-in.
Editing Platform and Crowdsourced Proofreading
The companion editing platform was originally an internal system and opened in 2024 to individuals and teams with public-interest editing needs, free of charge. Per the official manual (in Chinese), it integrates OCR (covering 98,000 characters), automatic punctuation, structural markup, entity extraction and multi-edition collation, and finished works can be published back to the reading platform. A volunteer program ("I proofread ancient books with AI") organizes public participation in rough OCR correction and careful punctuation review; Peking University News (January 2026) reports nearly 38,000 participants had processed over 20,000 titles — 1.5 billion characters roughly proofread and 100 million carefully proofread. The national Ru Zang compilation project has used the platform since 2025. Special collections include a high-resolution image database of the Yongle Encyclopedia (40 volumes first published in 2023) and a science-and-technology classics database launched in August 2026 with 1,982 titles.
Accounts and Access
No account is needed to browse, read or search; the terms of service (in Chinese) confirm that "you may start using Shidian Guji without registering, though some features may be limited." Signing in (ByteDance account system: phone code, password, or QR code via the Toutiao app) unlocks the bookshelf, notes, curated book lists, text copying and citation export (both capped by quotas, per on-page notices), deep research, the editing platform and proofreading tasks.
There is no public API or bulk dataset download; academics typically cite the corpus as the "Peking University–ByteDance open text corpus." Content endpoints are protected by strict anti-scraping controls, which limits programmatic reuse.
Copyright and Content Policy
The underlying classical texts are in the public domain, but the terms of service restrict reuse of the platform's content: commercial use requires written permission, re-editing and republishing the content outside the site is prohibited, and scraping, hotlinking and robotic downloading are banned (clauses 2.4, 4.3, 4.4 and 7.5). Copyright in source materials provided by Peking University remains with the university. The published contact for infringement complaints and institutional cooperation is guji@bytedance.com.
When It's Useful
- Tracing a quotation to its source in the transmitted canon, with book, chapter, edition and page references.
- Reading punctuated editions of major classics while checking doubtful characters against the original page images.
- Exploratory corpus queries, entity graphs and historical-geography browsing for digital humanities work.
- Public-interest digitization projects that need a free OCR and collation pipeline.
Limitations
- Machine text has a real error rate: the platform's 2022 figure of 96–97% OCR accuracy implies roughly 3–4 wrong characters per hundred; roughly-proofread texts have not been human-reviewed, so formal citation requires checking the page image.
- Vernacular translation, character-set conversion and AI answers are all machine-generated and labeled "for reference only."
- Text copying and citation export require sign-in and are quota-capped; no public API or dataset download exists, and the terms plus anti-scraping controls restrict bulk reuse.
- Coverage centers on Chinese-language books from China — transmitted classics first, with Buddhist, Daoist, excavated-text and local-gazetteer material alongside; whether page images exist for a given title depends on its base edition.
- The interface is primarily Chinese with limited English coverage, and sign-in is oriented toward mainland-Chinese accounts.
Comparable Sites
- Chinese Text Project (ctext.org): an academically oriented full-text library of early Chinese texts with paragraph-level text-image alignment; advanced features require a paid subscription.
- Shuge (shuge.org): a public-domain repository of high-resolution scans of old Chinese books, focused on images rather than searchable text.
- Zhonghua Ancient Books Database (中华经典古籍库): a commercial classics text database run by Zhonghua Book Company, sold by institutional subscription.
References
- Shidian Guji homepage (checked 2026-09-01)
- About - Shidian Guji (checked 2026-09-01, in Chinese)
- Terms of Service - Shidian Guji (checked 2026-08-30, in Chinese)
- Editing platform manual - Shidian Guji (checked 2026-08-30, in Chinese)
- Project page - PKU Digital Humanities Center (checked 2026-08-30, in Chinese)
- Peking University and ByteDance launch a digitization platform for ancient books - Xinhua (checked 2026-08-30, in Chinese)
- Over 2,200 titles including the Yongle Encyclopedia - Peking University News (checked 2026-08-30, in Chinese)
- Three years, 20,000 ancient books - ByteDance Philanthropy (checked 2026-08-30, in Chinese)
- "I proofread ancient books with AI" 2025 review conference - Peking University News (checked 2026-08-30, in Chinese)
- Over 70,000 ancient books free to read online - Beijing Daily (checked 2026-08-30, in Chinese)
- The Shidian Guji app is here - official WeChat account (reposted by Anqing Normal University Library) (checked 2026-08-30, in Chinese)




