⚠️ Ship-vs-build guardrail — read this before the menu
The pattern to name out loud: the engine's cheapness makes it tempting to spawn five verticals instead of shipping the one. Don't. The next move is DISTRIBUTING CiteBible — the single vertical about to go live — not opening new fronts. This whole document is what comes next, once CiteBible proves out: a stranger arrives, reads a genuinely helpful cited answer, and (some of them) pay. Until that loop closes once, a second corpus is architecture-as-procrastination.
What makes the menu safe to keep on the shelf: adding a vertical is ingest a corpus + reskin a front door — days, not a build (the engine, grounding gate, generation, and citation loop are already written and shared — see the build-map). So the leverage is real and the option stays cheap. Pick from this only when CiteBible has earned the right to a sibling.
The two things that are true for every corpus
Before ranking anyone, two mechanics apply identically to all of them — and they're the actual reason this is a distribution strategy, not just a product menu.
1 · The distribution mechanic — the engine is a content-SEO machine, not just an app
For any corpus, the product can auto-generate and index a page per query: "What does [X] say about [situation]?" — "…the Bible say about grief", "…Stoicism say about a toxic coworker", "…OpenStax say about photosynthesis". Each page is a grounded, cited answer built by the same engine that powers the live box. That is programmatic long-tail SEO, and long-tail queries are ~70% of all search and skew informational/question-shaped (Directive) — exactly the shape the engine answers natively.
Why this matters specifically for us: the Strategy doc and the platform research both land on the same finding — for a strong-engineering / weak-distribution solo builder, free-tool SEO is the #1 channel and paid ads are last, because it's a self-serve, compounding channel you build with code, not a sales team or an ad budget. Each indexed page is a permanent, free front door that keeps pulling traffic; the corpus's size becomes the number of doors. The engine doesn't just serve the app — it manufactures the distribution.
2 · The brand — "grounded, not generated"
Every answer is cited to a real source, or it says nothing. The grounding gate abstains on a retrieval miss without calling the model, so a fabricated citation is structurally impossible (verified in code — build-map, step 4). In an AI world drowning in confident hallucination, that honesty is the marketing — it's the one promise the big chat assistants structurally cannot make, and it's the same promise every candidate corpus inherits for free.
It also gives the family a name: the "Cite—" line — CiteBible first, and whichever siblings earn their place (CiteStoic, CiteLaw…). One recognizable promise across every corpus: grounded to a real passage, every time, or it abstains. The competitor to beat isn't another app — it's ungrounded ChatGPT, and the wedge is trust.
The ranked menu — by distribution/marketing leverage
The gate is license (must be public-domain or openly licensed to ingest freely — verified per row). The rank is distribution leverage: audience we can actually reach + how well "cite = the whole value" fits + how cleanly the SEO play works + how little liability drags it down. All figures web-verified; sources at the foot.
| # | Candidate | Corpus + license the gate | Audience / demand | Cite-trust fit | Risk | SEO play | Verdict |
|---|---|---|---|---|---|---|---|
| 1 | Stoicism "CiteStoic" |
Clean PD. Marcus (Long, Gutenberg #2680), Epictetus (#10661), Seneca (Gummere, Wikisource). Strip the PG boilerplate; text is unrestricted. | Large, proven, paying. Daily Stoic ~900k–1.5M subscribers, r/Stoicism ~600k, the Stoic app 3M+ downloads; ~600% search rise/5yr. | very strong The web is full of fake Marcus quotes — "real, locatable passage" is the differentiator. | low Secular, no living rights. Life-advice edge → cite-not-therapy posture (same as CiteBible). | Strong template; Daily Stoic already owns generic listicles → win on cited + situation-specific. | SHIP as #2 |
| 2 | Open textbooks "CiteStudy" |
CC BY, per-title filter. OpenStax 80+ titles (exclude the NC ones like Calculus); LibreTexts adds breadth. | Largest, most validated. OpenStax: $3B saved, 43M learners, 72% of US colleges; Chegg/Course Hero/Quizlet at 60–100M MAU. | strong Students literally need "where in the book?" — cited page = the anti-Chegg, anti-hallucination proof. | low A wrong photosynthesis answer isn't tort. Only trap = "cheating tool" framing (cited-explainer avoids it). | Strongest of all. One page per section maps 1:1 to "explain [concept]"; incumbent (Chegg) is vacating the niche. | biggest SEO ceiling |
| 3 | Dhammapada / Buddhist "CiteZen" |
Cleanest license here. Bhikkhu Sujato (CC0) — a modern, readable, complete translation you can modify. (Others lock their good modern renderings.) | Small faith base, big crossover. r/Buddhism ~739k; overlaps the large secular mindfulness / mental-health market. | high Suttas are citation-addressed (Dhp 1, MN 10); CC0 lets you quote and paraphrase. | lowest No ruling-authority / blasphemy structure; "what the Buddha said about anger" already normalized. | Solid, adjacent to the hot meditation niche; less incumbent clutter than Christianity/Islam. | cleanest sacred sibling |
| 4 | Torah + commentary "CiteTorah" |
Excellent + uniquely rich. Sefaria API: JPS 1917 & Community Translation are CC0, plus openly-licensed Rashi/Talmud/Midrash cross-links. Filter to CC0/PD/CC-BY. | Smallest population (~15.8M), best infra. Sefaria already trained a digitally-sophisticated study audience + hands you the corpus + citation graph. | highest structural fit Jewish study is citation + cross-reference. Maps perfectly. | moderate-high Halachic "what should I do" edges toward psak (ruling) — frame as study, not ruling. | Narrow but ownable; low competition beyond Sefaria itself; low absolute volume. | best corpus, capped reach |
| 5 | Public-domain law "CiteLaw" |
Genuinely PD. US-gov works (17 USC §105), opinions uncopyrightable (Georgia v. PRO); ingest CourtListener bulk. | Highest WTP, B2B. Legal-tech ~$27B; Casetext $650M, Harvey $11B, vLex $1B exits prove it. | best fit of all A fake cite is sanctionable (Mata v. Avianca). Cite-everything is the entire value. | UPL Unauthorized practice of law; the FTC/DoNotPay order shows even the claims are risky. | weak/B2B Sales-led, low-virality; consumer tail is UPL-exposed. | best business, wrong shape |
| 6 | Bhagavad Gita | PD available, popular one locked. Arnold "Song Celestial" (Gutenberg #2388), Telang (SBE). ISKCON "As It Is" is copyrighted → you're stuck with archaic prose. | Large base (~1B), weak infra. Big secular/Western readership, but no Quran.com/Sefaria-scale open platform. | moderate Citable, but lay practice is less citation-reflexive; readers attach to the (copyrighted) Prabhupada voice. | moderate Hindu-text politics are live; archaic-only translations feel academic. | Good, less-saturated ("what does the Gita say about duty/fear"); translation-quality gap. | later, if ever |
| 7 | Quran | Encumbered. Best modern rendering (Sahih Intl via Tanzil) is verbatim-only, no-modification, attribution-locked; PD options (Yusuf Ali/Pickthall) are dated / US-murky. | Largest prize (~2B). Muslim Pro 190M+; highest daily-devotional demand of any. | very high Muslims cite surah:ayah reflexively — but provenance expectations are exacting. | highest Reads as issuing fatwa; mistranslation/blasphemy/backlash a solo builder can't absorb. | Strong demand but saturated + trust-gated toward Islamic institutions. | great market, wrong first move |
| 8 | Health / medical | Most constrained. CDC PD + PMC CC-BY tier + NLM-authored MedlinePlus only; the useful A.D.A.M./drug content is licensed-out; WHO is NC (excluded). | Enormous. 35% of US adults self-diagnose online; "Dr. Google" ~70k queries/min. | strong YMYL/E-E-A-T rewards cited authority — cites manufacture the trust signal Google demands. | headline risk Unauthorized practice of medicine + FDA SaMD; the Tessa pull is the cautionary tale. | hardest terrain YMYL concentrates traffic in Mayo/WebMD/.gov; a new domain climbs slowest. | avoid — liability > leverage |
B2B sibling — "VaultChat"
Same engine over a customer's OWN corpus (handbook / wiki / filings). High willingness-to-pay, but the go-to-market is direct sales, not marketing — no public corpus, no SEO play. A different motion; noted here as the platform's B2B twin, out of this consumer ranking. See the platform research.
Cut — PD literature at large
Gutenberg's ~78k works are a corpus reservoir, not a launch vertical: no single searcher intent to own, weak cite-fit for fiction/poetry, and a red-ocean SEO fight vs SparkNotes/Goodreads/Wikipedia. Its one strong sub-segment is Stoicism (#1). Expand into it later, don't lead with it.
The one non-negotiable
License is the gate, not the tiebreaker. Every ranked row clears "free to ingest" — verified. The Quran and Gita show the trap: the text can be ancient/PD while the good modern translation is copyrighted. Always ingest the openly-licensed rendering, or don't add the vertical.
The recommendation — propose-only
If and when CiteBible earns a sibling, the evidence points one way for the consumer line and a second way for the SEO-volume bet. They're different shapes; pick the shape, not just the corpus.
#1 next vertical: Stoicism ("CiteStoic"). It is the nearest neighbor to CiteBible — the identical product shape (type the situation you're carrying → get cited wisdom back), same emotional hook, so the reskin is the cheapest possible. And it clears every axis at once:
- License is clean and free — Marcus/Epictetus/Seneca are all public-domain (strip the Gutenberg boilerplate). No liability, no living rights.
- The audience is large, proven, and already pays — Daily Stoic (~900k–1.5M), r/Stoicism (~600k), the Stoic app (3M+ downloads). This is a monetizing crowd, not a hoped-for one.
- Citation is uniquely the value — the Stoicism internet is polluted with fabricated Marcus Aurelius quotes; "grounded to a real, locatable passage — or it says nothing" is a direct answer to that exact pain. The brand promise lands harder here than almost anywhere.
- Low sensitivity — secular self-improvement, same cite-not-therapy posture already built into CiteBible.
The honest caveat: Daily Stoic already owns the generic content space. The wedge is not more listicles — it's the cited, situation-specific answer to a messy real sentence, which a static article can't do.
The runner-up is a different bet, not a lesser one: Open textbooks (OpenStax)
If the goal is maximum SEO distribution rather than nearest-neighbor product fit, open textbooks win outright: the biggest proven audience, the strongest 1:1 programmatic-SEO surface ("explain [concept]" over 80+ CC-BY titles), the lowest liability, and an incumbent (Chegg, down to ~3.2M subs from 7.8M) actively vacating the exact cited-vs-hallucinated gap this fills. It's a different shape from CiteBible (study help, not "what are you carrying"), so it's more of a reskin than Stoicism — but its distribution ceiling is the highest on the board. Stoicism = cheapest, warmest, nearest. Textbooks = biggest engine. Sam's call which lever to pull; both are strong, and neither happens until CiteBible ships.
Ship-vs-build, honestly (the close)
This is a menu, not a mandate. The strong-engineering / weak-distribution pattern predicts exactly the failure mode this document could cause: starting a second (and third) vertical because the engine makes it easy, instead of walking the one that's about to be live the last mile to a paying stranger. So the ranking above is deliberately shelved behind one event — CiteBible earning its first dollar. When that happens, the cheapness is a gift: a sibling is days of ingest + reskin, and the evidence says point it at Stoicism (nearest, warmest) or OpenStax (biggest SEO engine). Until then, the single next action is unchanged and lives in Strategy and the build-map: get CiteBible to a live URL and in front of people.
Sibling reads: the engine build-map (how a vertical maps onto existing code) and the Strategy (the business why + free-tool-SEO finding). Propose-only · distribution/marketing lens · 2026-07-06 · every license claim was checked against its source before it was written down. Web-scan figures move — revisit before acting.