When building Guji, a semantic search engine and scholar chatbot over classical Chinese literature, I wanted to answer two practical questions. First, can modern dense embedding models reliably retrieve specific passages when scaled to tens of thousands of classical texts? Second, if modern LLMs have already seen these famous texts during pre-training, do they actually need retrieval, or can persona prompts alone produce accurate historical dialogue?
To find out, I built an evaluation suite over a 54,000-chunk corpus covering the Four Early Histories (Shiji, Han Shu, Hou Han Shu, Sanguozhi), the Zizhi Tongjian, philosophical canons, and Song-dynasty annals, benchmarked embedding models, and ran a three-mode ablation across eight LLMs.
Dense retrieval at scale: where embeddings collapse
I evaluated retrieval against a benchmark of 35 historical queries, ranging from famous quotes and battle strategies to conceptual descriptions of political events. A hit was scored if the retrieved top passages contained the expected target text or chapter.
I tested Google's gemini-embedding-001, the older
text-multilingual-embedding-002, and a character-bigram BM25 lexical
baseline.
| Model / Approach | Success Rate (Top-5) | Search Accuracy (MRR) | Primary failure mode observed |
|---|---|---|---|
| gemini-embedding-001 | 33 / 35 (94%) | 0.876 | Struggles on rare characters and short verbatim proverbs |
| text-multilingual-embedding-002 | 26 / 35 (74%) | 0.644 | Collapses as corpus scales; semantic drift buries targets |
| BM25 (char-bigram) | 22 / 35 (63%) | 0.582 | Fails completely on conceptual or paraphrased queries |
On a small pilot corpus of a few thousand chunks, multilingual-002
appeared competitive. But as the corpus grew past 30,000 chunks, its
performance dropped steeply. Dense vectors map semantically related
concepts into tight neighborhoods; in a massive classical corpus where
hundreds of passages discuss court politics, war, and imperial edicts
in similar vocabulary, semantic similarity alone buried the exact chapter
I was looking for.
Furthermore, pure vector search exhibited a blind spot: verbatim famous proverbs (such as 「王侯將相寧有種乎」) and rare ancient terms (like 觬是). Because embeddings compress sentences into abstract conceptual spaces, exact rare characters often lose their uniqueness. The solution was a hybrid retrieval pipeline: matching quoted terms and rare bigrams directly into the candidate pool (+0.25 flat score boost), followed by reranking on cosine similarity plus n-gram coverage. That hybrid combination resolved the remaining retrieval misses.
The three-mode ablation: RAG vs. Persona vs. Raw
Once retrieval was working, I wanted to know how much retrieved context actually matters to the final answer. I tested four historical test questions across eight different LLMs using three distinct modes:
- RAG + Persona: Top-5 retrieved corpus passages provided alongside a Song-dynasty scholar persona prompt (沈昭遠, an imperial scholar of 1088).
- Persona-only: The exact same scholar persona prompt, but zero retrieved passages (forcing the model to rely on pre-trained memory).
- Raw (No context): The bare question with no persona prompt and no retrieval.
The test questions probed both specific factual recall (such as asking what Book 60 of the Former Han Shu refers to) and era-appropriate viewpoints on historical social customs.
Old-generation models (2024–2025)
| Model | RAG + Persona | Persona-only | Raw (No context) |
|---|---|---|---|
| Qwen-2.5-72B | Correctly identified the Du Zhou biography; exact citations. | Fabricated a fake Chunyu Kun quote in the Huaji biography. | Modern academic monograph style. |
| DeepSeek-V3 | Correct identification; compact; real historical quotes. | Invented fake academic footnote brackets [1]-[4]; costume was fake. |
Contemporary formal administrative prose. |
| Gemini-2.5-Flash | Correct; verbatim quotes. | Safely admitted lack of source ("未見原文不敢妄議"). | Modern balanced analytical essay. |
| Gemini-2.5-Pro | Correct; resolved edition discrepancy between Book 60 and 90. | Relied on memory but misattributed quotes to Hou Han Shu instead of Han Shu. | Long contemporary discursive essay. |
In older models, running persona-only proved dangerous. Because the persona prompt instructed the model to speak as an erudite Song scholar, the model felt compelled to sound authoritative. Without real text in its context window, it fabricated citations out of thin air—inventing quotes and generating fake scholarly footnote brackets that looked authentic at a glance.
New-generation models (2026)
| Model | RAG + Persona | Persona-only | Raw (No context) |
|---|---|---|---|
| Gemini-3.8-Flash | Verbatim blockquotes; noted cross-book discrepancies; biography-aware. | Refuses without source: "案頭既無此編…不敢憑空懸想"; integrates era persona cleanly. | Contemporary institutional policy essay. |
| Gemini-3.1-Pro | Sentence-level citations; genuine inference from retrieved catalog. | Refuses without text; draws appropriately on general historical context. | Modern analytical essay style. |
| DeepSeek-V4-Pro | Verbatim quotes; best historical voice and period stage directions. | Answers accurately from memory; quotes are real; fake footnotes fixed. | Contemporary formal administrative prose. |
| Qwen-3.8-Max | Sharpest analytical critique; era-authentic moral vocabulary. | Refuses to opine without the book; correct from memory when it does answer. | Neutral modern academic paper. |
The difference in newer models was striking. Where older models confabulated sources to satisfy the persona, newer models have internalized source discipline: in persona-only mode, they frequently refuse to speculate if the referenced work is not present on their "desk."
However, raw mode across all models revealed an underlying linguistic reality: without persona constraints and grounding passages, every model immediately defaults to 21st-century registers. Depending on each model's underlying corpus balance, raw answers read like contemporary policy papers, institutional briefs, or modern academic textbooks. None of them preserve an era-appropriate voice or classical vocabulary unless explicitly anchored by retrieved historical passages.
What the results show
These benchmarks clarified three things for me:
- Pre-training memory is not a substitute for retrieval: Even when an LLM has memorized an ancient text, asking it to speak without the source forces it either to invent evidence (in older models) or to hedge and refuse to answer (in newer models).
- Dense retrieval needs lexical grounding: Embedding models alone suffer from semantic crowding as historical datasets grow. Pairing dense similarity with n-gram verification is what makes retrieval reliable.
- Retrieval provides the evidentiary boundary: RAG is not just a way to fetch facts—it is what keeps a persona honest. The retrieved passages give the model permission to be specific, without giving it license to hallucinate history.