User request: one big list at repo root mirroring isapzurich.com/en/library/reading-list — every book with links to ALL downloaded files, INCLUDING cross-section books (link always resolves to the canonical section's file, tagged →secNN; per user: link every time the book is mentioned, even if downloaded in an earlier section). - 483 positions (12 sections), 566 files, 1234 links, 0 broken (verified) - sources: sections/NN/INDEX.md (items) + downloads/NN/MANIFEST.md (files) + cards (RU title) - resolves: per-item blocks, xref rows (secNN #MM), legacy 5-col rows (| EN | file | …), legacy prose «Файлы: `path`» rows, glob xref rows (04-cw9-archetypes-*), card-name drift (sec07 INDEX) by item-number prefix - manifest fixes found while building: sec08/10/11 'files in sec??' headings → '0 files' (23 blocks, paid/absent files), sec12 #03 xref row gained #19, sec11 #13 → 0 files - AGENTS.md: layout line + total corrected (483 positions, 566 files) CANON question answered: the new manifest canon (09-12 style) is exactly what makes this generator possible; sec01-08 were migrated to it in the previous commit.
70 KiB
70 KiB
Jung Reading List → Russian Editions: Research Notes
Goal
For every book in the ISAP Zurich reading list (https://isapzurich.com/en/library/reading-list), find the Russian edition(s): correct Russian title, publisher, year, translator, ISBN, libgen download page (if exists), and РГБ (RSL) bibliographic record (if exists).
We work section by section. Status (детали — в SUMMARY.md/MANIFEST.md секции + git log):
| sec | name | items | RU ✅/🔶/❌ | files (EN+RU) | state |
|---|---|---|---|---|---|
| 01 | fundamentals | 55 | 38/2/15 | 113 | FINAL 2026-09-25 |
| 02 | dreams | 57 | 23/5/29 | 37 | FINAL 2026-09-25 |
| 03 | myths & fairy tales | 68 | 38/0/30 | 129 | FINAL 2026-09-25 |
| 04 | pictures | 39 | 12/1/26 | 28 | FINAL 2026-09-25 |
| 05 | ethnology | 29 | 17/0/12 | 38 | FINAL 2026-09-25 |
| 06 | religion | 28 | 19/1/8 | 38 | FINAL 2026-09-25 |
| 07 | complexes | 27 | 4/0/23 | 20 (EN only) | FINAL 2026-09-26 |
| 08 | developmental | 48 | 20/1/27 | 44 (333 MB) | FINAL 2026-09-26 |
| 09 | comparison of psychodynamic concepts | 38 (20 NEW + 18 xref) | 11/1/8 | 43 (366 MB) | FINAL 2026-09-27 |
| 10 | psychopathology & psychiatry | 37 (32 NEW + 5 xref) | 5/0/27 | 30 (268 MB) | FINAL 2026-09-27 |
| 11 | individuation process | 22 (7 NEW + 15 xref) | 6/1/0 | 10 (128 MB) | FINAL 2026-09-27 |
| 12 | practical case | 35 (23 NEW + 12 xref) | 14/1/20 | 35 (184 MB) | FINAL 2026-09-28 |
ALL 12 SECTIONS COMPLETE (2026-09-28). Total: 483 positions, 566 files (~2.3 GB). Root view: READING-LIST.md (generated by tools/make_rootlist.py; regenerable, do not hand-edit).
PROJECT COMPLETE (2026-09-28). All 12 sections final. Follow-ups: user-oracle re-probes on ❌ lists (per-section MISSING.md), AGENTS.md variant A (extract sources).
Repository layout
AGENTS.md # this file — plan, conventions, source status
CONTEXT.md # misc findings that fit no other category
data/MASTER-LIST.md # CANONICAL: all 12 sections, cross-section dedup (tools/xref.py --write)
READING-LIST.md # ROOT VIEW: весь список как на сайте ISAP + ссылки на файлы каждой книги
# (генерируется tools/make_rootlist.py из INDEX + MANIFEST + карточек; не редактировать руками)
data/isap-raw/ # RAW-списки ISAP 08–12 (scope pending); 07 в sections/07-complexes/RAW.md
tools/ # search helpers (lg.py, rsl.py, sx.sh, xref.py, audit_anchors.py, …)
authors/ # CANONICAL: one dossier per author (shared across all sections)
sections/0N-<name>/ # one .md per BOOK (chapter-level list items are merged into the book file)
# + AUTHORS-EN.md = THIN INDEX (item → author → link to authors/<slug>.md)
Conventions
- One file per book, named
NN-<slug>.md(NN = order in the reading list). - Status marker at top of each book file:
⬜ not searched yet🔎 searching in progress✅ RU edition(s) found🔶 partial(e.g. only an essay/chapter available in RU)❌ no RU edition found (verified)
- Book card CANON v3 (agreed 2026-09-21; mandatory for new cards, sec05+ from day one):
Rules: (a) Editions-секции симметричны по языкам, порядок: язык списка → RU → прочие; (b) ISBN только в таблицах (из шапки убрано поле EN ISBN); (c) пустые секции не пишутся; (d) ЗАПРЕЩЕНЫ плейсхолдеры:# <listed title> (<listed edition: publisher, year>) **Author(s):** <Last, First> [dates] **Shelf mark:** <ISAP or —> **Section:** <NN name / subgroup> **Original:** <язык оригинала, изд., год — только если ≠ языка списка (ES/DE/…)> **Status:** <✅ | 🔶 | ❌> — <одна фраза: «RU 4 изд.» / «essay only» / «no RU (verified <date>)»> ## Editions — EN ← язык списка; несколько строк = несколько изданий | Title / edition | Publisher | Year | Pages | ISBN | Notes | |-----------------|-----------|------|-------|------|-------| ## Editions — RU ← тот же формат + колонка Translator | RU title | Publisher | Year | Pages | Translator | ISBN | |----------|-----------|------|-------|------------|------| *(нет: одна строка `— нет (verified <date>, gate: MISSING.md §NN)`)* ## Editions — ES ← только если есть изд./оригинал на 3-м языке ## Downloads ← только если есть файлы; симметрично по языкам | lang | file | size | source | |------|------|------|--------| | EN | NN-…-en-<edition>.pdf | 64 MB | lg f/… | ← source: lg f/ | flib b/ | ia <item> | RU | NN-…-ru-<edition>.fb2 | 4.4 MB | flib b/… | ## Catalog records РГБ: rsl.py "<query>": [id] <GOST-строка>; [id2] … РНБ: ✓ 3 (только если sweep трогал; добавления — с данными) МГУ: ✓ 1 / GBS: ✗ 0 / flib: a/<id> (N книг) EN refs: OL (N ed., ISBNs ✓ isbnval) · Google Books 200 · ia: <item> *(❌-карточка — ОДИН compact gate: `РГБ 0 · НРБ 0 · МГУ 0 · flib 0 (verified <date>, gate: MISSING.md §NN)`)* ## Notes - <идентичность, identity-check, cross-refs на досье/секции; датированные append: «2026-09-20: …»>// not searched yet,(empty), дубли## Notes; (e)## Libgenболее нет (был асимметричным); файлы ≠ издания (3 файла одного изд. = 3 строки); (f) EN-каталогами: полные записи не ведутся (WorldCat мёртв) — строкаEN refsпо данным isbnval/OL/ia; (g) ✅-карточка: РГБ обязателен + ≥1 второй каталог если sweep его трогал; (h) старое поле Status-формулировок: единая «маркер + одна фраза». Все карточки мигрированы на v3 (2026-09-21, one-shottools/migrate_cards_v3.py); lintertools/check-md.py= 0 issues. New cards start on v3. - Section files unified 2026-09-21 (CANONs ниже; one-shot скрипты
tools/migrate_*.py,make_authors_table.py,restructure_summary.py— done, не для повторного запуска). Living authors keep open-ended dates[1951-]. - Section file CANONS v1 (agreed 2026-09-21):
SUMMARY.md: H1
# Section NN — <Name>: summary (<date>)→## Final counts(одна строка: статусы + файлы) →## Publication status at a glance(ОБЯЗАТЕЛЬНЫЙ, колонки# | EN book | EN ISBN | RU book | RU ISBN | DL EN | DL RU+ totals-строка + legend ✅/⚠/❌) →## Authors — EN → RU(colitem(s) | EN author | RU name | verified via) →## Notes(датированные append). Запрещено: «Table N»-нумерация, «Table 2 — Books» (дубль карточек), ISBN-validation-таблицы (данные в data/isbn-validation.tsv, строка-ссылка). INDEX.md: одна таблица| # | Group | EN title | Author(s) | Status | File |(Group = A.1/B.1, нет подгрупп →—; Status = только маркер; File = имя карточки). MANIFEST.md (downloads/NN-<name>/): H1# Downloads — Section NN <Name> — FINAL <date>→## Totals(N files, X GB, RU/EN split) →## Files (per item): на каждый item### NN — <title> — <status>+ таблица| file | size | source | verify |(source: lg f/… | flib b/… | ia ). Cross-ref rule (2026-09-26, user): item, чья книга впервые появилась в предыдущей секции, ОБЯЗАТЕЛЬНО получает свой per-item блок в основном списке (заголовок «— files in sec0N» + одна cross-ref-строка| — (cross-ref sec0X #NN: <файлы>) | … |); блок## Cross-references= только резюме-указатель, не замена per-item блокам. →## Not downloadable (verified <date>)таблица| item | reason |→## Cross-section→## Notes. MISSING.md: H1 →## ❌ (N)→ на каждый### NN — <EN title> (<author>)+ 1–3 строки evidence (каталоги 0, verified date) —### NN= ЯКОРЬ для gate-строк карточек (gate: MISSING.md §NN) →## 🔶 (N)→## Re-probes(датированные append). Dossiers (authors/*.md): H1# <Lastname, Firstname> [<dates>]→ bold-поля**Wiki anchor:**/**Dates:** / **Field / identity:** / **RU name (canonical):** / **RU name (candidates):**→## List items (ISAP Zurich)таблица| sec | # | title | status |→## Verified RU worksтаблица| RU title | EN original | publisher, year | ISBN |→## Homonym warnings(ЕДИНОЕ название, legacy-варианты: «Homonym warning (date)», «Homonyms / notes», «Омонимы») →## Notes(датированные append; legacy «Round N» логи сюда). Запрещено:**EN identity:**отдельно,**Dossier status:**. Wiki anchor (2026-09-24, MANDATORY):**Wiki anchor:** QID (URL)— факты identity (даты, родство, поле) БЕЗ якоря = unverified и не пишутся как утверждение. Якорь = Wikidata QID (wbgetentities по QID — точечно, без омонимов) + en/ru-wiki URL. Инструмент:tools/verify_dossiers.py(resumable, TSVdata/dossier-check.tsv, 3s/polite + 429-backoff; статусы: MATCH = QID сверен, CANDIDATE = найден поиском — РЕВЬЮ перед записью якоря, NO-QID = legitimately absent). Предыстория: 2 досье с галлюцинациями (Emma Jung: «1877-1965» + «eldest daughter» вместо 1882-1955 + жена; Verena Kast: «1951» + «daughter of Hans Kast» вместо 1943 + отец Walter Kast) — lesson: identity-факты только с источником. - TODO (later, after 4 sections in order):
tools/liart.py— HALF DONE 2026-09-21: OPAC-Global protocol reverse-engineered (direct.exe/FindView, guest session), SearXNG record-URL fallback works; BLOCKED on search labels + iddb (server InfoDB, admin-only for GUEST) — needs a one-time browser read (see data/liart-params.json). - Author dossiers (
authors/<lastname>-<firstname>.md) are the CANONICAL home for identity, shared across ALL sections (same author can appear in 01, 04, …). One dossier per PERSON, slug<lastname>-<firstname>(lowercase, hyphens). Hold: dates, RU name candidates (all plausible transliterations, German names often have several: Нойманн/Нойман/Нейманн), the list items that cite this author (sec # + status), verified RU works with EN mapping, homonym warnings. sections/0N/AUTHORS-EN.mdis a THIN INDEX only — tableitem(s) → author → link to../../authors/.md``. It must NOT duplicate identity data (no RU-name analysis, no RU works table). If an author appears in two sections, the dossier is written once and both section indexes point to it. Cross-section homonym traps (e.g. the two different "Alice Miller"s) get their own dossier + a warning in both.- Commit per completed book (or when meaningful author info is gained), message like:
sec04: <book> — <finding>. - Update
CONTEXT.mdfor anything useful that fits no category.
Sources and their status (verified 2026)
1. РГБ (Russian State Library) via libgen.vg biblioservice — PRIMARY for existence + editions
- Tool:
tools/rsl.py "Full Name"(JSON, polite, 3s between pages) - Raw URL:
https://libgen.vg/biblioservice.php?value=<query>&type=rsl&format=json - Rich records: title parts, author with life dates
[1905-1960](great identity anchor), translator, publisher, city, year, pages, ISBN, series, UDC tags, even TOC field. - Limitations: no pagination (fixed first page) → always query by full first+last name for people, or 2–3 distinctive words. Loose matching (searches author AND title fields).
- Other
type=values on this mirror are dead: worldcat, googlebooks, isbndb, udc (all return "Nothing found" even for control queries).type=isbndecodes ISBN structure: checksum error + registered publisher group name (e.g. 5-519 → «Клуб Касталия», 5-98712 → «Петроглиф») — cheap, unlimited; used bytools/isbnval.pyas the structure layer.
2. libgen.vg full-text search — PRIMARY for readable editions
- Tool:
tools/lg.py "query"(whole words only, res=100, prints pagination hint) - Raw URL:
https://libgen.vg/index.php?req=<query>&res=100 - Requires a browser User-Agent (otherwise nginx default page / robot block).
- Search semantics (verified): whole-word, case-insensitive. Stems do NOT work reliably
(
происхожден→ 0 hits; fullПроисхождение→ many). Multi-word = AND. → Use full inflected words (nominative, as stored in titles) and 2–3 distinctive words. - Author search works (
req=Нойманн) but pulls homonyms — filter by first name in results. - Pagination:
&page=Nworks. - Search objects (2026-09-25):
objects[]param = что искать (f=files, e=editions, s=series…). Без него = editions-only — файлы, не привязанные к нужному изданию, НЕ находятся (Solomon: editions-поиск = только mislinked 2021-издание, files-поиск = настоящий файл 2018-го). lg.py теперь шлётobjects[]=f,e,s,a,p,w+topics[]=l,c,f,a,m,r,sи печатает[md5=…]в строках →tools/lgdl.py bymd5 <md5>= f_id + метаданные. Это нашло все 6 rescues sec06. - Edition page:
https://libgen.vg/edition.php?id=<id>.
3. cogito-shop.com — current Russian trade (Jungian/psychoanalysis specialists)
- URL:
https://cogito-shop.com/search/?q=<query>(param is q, not query) - Matches author AND title (fuzzy — «исцеляющее сновидение» missed, «Майер» hit). Good for "is it sellable now in RU".
- Product pages work with plain curl (2026-09-19): «Характеристики» block in HTML —
Автор / Издательство / Формат / Вес / Тип обложки / Кол-во стр / Год / ISBN / Код. Search URL
=
/search/?q=<author>, product link =/catalog/<cat>/<slug>/.
4. SearXNG (local) — cross-verification
- Tool:
tools/sx.sh "query"(JSON API at http://localhost:8888; the web_search tool blocks localhost) - Use to verify RU↔EN title pairing in a third party's words, resolve transliterations.
- Keep queries ≤ 3–4 words; «…» quotes only on distinctive RU titles. Unstable relevance — drop junk queries, don't retry blindly.
- Engine status (2026-09-19): container
searxng(docker, config~/.searxng/settings.yml→ mounted /etc/searxng/settings.yml). Currently google only: bing is BLOCKED (responds but ignores query — random promo junk), brave dead (403), duckduckgo CAPTCHA, wikipedia/wikidata dead/timeout — all disabled in config. Diagnosis:docker logs searxng | grep -oE '(ERROR|WARNING):searx.(engines|network).[a-z_]+' | sort | uniq -c- per-engine test
curl --get :8888/search --data-urlencode 'q=...' --data-urlencode format=json --data-urlencode 'engines=google'. Fix = mark enginedisabled: truein settings.yml +docker restart searxng.
- per-engine test
- Always pass explicit
language=ru/language=en(user-confirmed fix; per-request, config stays multilingual).
5. ru.wikipedia / en.wikipedia — author dossiers
- ru-wiki articles often have «…на русском» sections listing actual RU titles.
- en-wiki for the EN bibliography (identity anchor: 3–5 flagship titles).
6. livelib.ru author pages — all RU works of a known author
- e.g.
livelib.ru/author/<id>-<slug>lists every RU edition (server-rendered, scrapable). - Author ID: find via SearXNG
site:livelib.ru <author>. The /search page is JS-only (nope). - Publisher pages
/publisher/<id>-<slug>= the publisher's full book list (server-rendered, book links/book/<id>-<slug>carry title + author slugs) — cheap reverse publisher sweep. Example: /publisher/2076-medkov-s-b (Медков С. Б. — small esoteric press w/ Jung line).
7. cogito-shop person pages — shop stock per author
cogito-shop.com/person/<first>_<last>/(transliterated, first-name first, e.g./person/mariya_luiza_fon_frants/,/person/yaffe_aniela/).- Slug discoverable from the shop's search page (results contain person links) or guessable.
- Lists every edition the shop carries (title variants!). Product pages lack full metadata in HTML (JS-rendered) — use RSL/livelib for publisher/year/ISBN.
8. GNB SPb (State Public Library of St. Petersburg) catalog — SECONDARY bibliographic source
- URL:
https://www.gbs.spb.ru/ru/search/detail/?id=<hash>(rich GOST records: title, responsible parties incl. translators, publisher w/ full legal names, year, pages, ISBN, notes like «Др. кн. авт.») - Complements RSL: RSL (RGB) misses some small-press editions (e.g. Pattis-Zhoya
«Аборты…» Т8/ЦГИ 2017 is in GBS but NOT in RSL). Query via SearXNG
site:gbs.spb.ru <author>.
9. chitai-gorod.ru / ozon / labirint — trade retail pages
- chitai-gorod author pages list RU works (some noise). Product pages: CRW works (2026-09-19) —
«Характеристики» block: Год издания / Кол-во стр / Переводчик / Издательство / ISBN. Plain curl
gives nothing (JS). Find product IDs via SearXNG
site:chitai-gorod.ru "<RU title>". - labirint.ru book pages work with plain curl — «Характеристики» block after
(publisher, year, pages, translator, ISBN). Labirint SEARCH is anti-bot (text= ignored).
- Ozon product pages: blocked from curl (redirect loop) AND via CRW (FAB challenge 2026-09-19); fetch_content also fails. Use SearXNG OZON snippets instead.
- vse-svobodny.com — works with
Cookie: beget=begetok(anti-bot challenge, 2026-09-19); product pages carry publisher/year/pages.
10. flibusta — big e-book library. .su = server-rendered search; .is = OPDS + direct downloads (user fixed access, 2026-07-16)
- .is OPDS (preferred): Tool
tools/flib.py:flib.py authors "query"→/opds/search?searchType=authors&searchTerm=…(id, name, book count)flib.py books "query"→searchType=books(id, title, author, year, format, translator from annotation)flib.py author <id>→/opds/author/<id>/alphabet— book list (20 per page!)flib.py authorall <id>— followsrel="next"pages (/alphabet/1,/2, …) to the end — always use authorall (Jung a/5272 = 128 books; the bare feed shows only the first 20)- Author HTML:
/a/<id>(e.g. a/57639 = фон Франц, a/118921 = Нойманн); book page/b/<id> - Direct downloads:
/b/<id>/fb2|epub|mobi|html|txt(pdf/doc as «/b//download» links on the pages) - Quirks (2026-07-17):
fb2+zipformat actually returns a ZIP (unpack: 1 fb2 + cover inside);/b/<id>/pdfsometimes 302-redirects to the book HTML page (a ~20 KB stub) when the file isn't served directly — check downloaded «pdf»s withfile, fall back to libgen (same editions usually there) or the fb2 twin.
- Converter quirk (2026-07-18):
/b/<id>/epub(or any fmt endpoint) on a book whose NATIVE format is pdf/djvu/doc may return the native file, not an epub — always magic-byte-check downloads (file); some «20 KB HTML stubs» are in fact real native files with the wrong extension. - .su = mirror of .is with reduced functionality (user, 2026-07-17) — use .is OPDS only.
.su
booksearch/?ask=exists but adds nothing; API endpoints blocked (api/search.php → 403). - Value: (1) RU full-text downloads for the download phase; (2) author-book-list = cheap author sweep; (3) negative verification (all same-name authors listed). Verified 2026-07-09 (.su): no «Выявляющий образ», no «Сэндплей» titles, no «Тест Дерево», no «Искусство и творческое бессознательное» (as titled). Flibusta search is fuzzy (whole-phrase not required).
10b. OpenAlex + Crossref (tools/oa.py) — Phase 0 identity/ISBN (free, no key, better than OL for modern academic books)
oa.py "Exact EN Title"(OpenAlex works: full author names, year, DOI, publisher),oa.py --crossref "Title words"(Crossref: ISBN list + publisher + year),oa.py --author "Last, First"(OpenAlex author entity),oa.py both "Title".- Verified 2026-07-17: Evers-Fahey full name "Karen Evers-Fahey" + Routledge 2016; "Coming into Mind" = Margaret Wilkinson (not Michael!) 9781317710578; "The Symbolic Quest" Whitmont 9780691213187 (Princeton).
- Use in Phase 0 INSTEAD of / alongside OL; OL still primary for pre-1990 titles.
10c. DOI pipeline for article items (tools/doi.py) — Crossref → DOI → libgen (2026-09-26, user idea)
- Reading-list items that are JOURNAL ARTICLES:
doi.py "Last" "Title words" --year YYYY= Crossref (query.bibliographic + query.author, year ±2) → top-DOI hits ranked by title-token overlap → for each: libgenreq=<DOI>(libgen indexes papers BY DOI — verified: Krieger 10.1111/1468-5922.12544 → exactly ed 84462696; Meier I 10.1111/1468-5922.12545 → 84462697; Bovensiepen 10.1111/j.0021-8774.2006.00602.x → 15358298; controls 3/3). - Then
lgdl.py dl <f_id> ...+ content-verify: first page must be the PAPER (author + journal + pages matching RAW + abstract), NOT a review of it (user rule 2026-09-26). - Limits: pre-1997 works have no DOIs (1975 Hill paper — fall back to title sweeps); Crossref top hits can be homonym noise (score column disambiguates; pick the [0.8x] Jungian one).
- Recovered sec07 EN via this route: #05, #17 (ed from user), #20. Also found a JAP 2008 Bovensiepen book-REVIEW (10.1111/j.1468-5922.2008.00737_3.x) — the review-check matters.
10d. flibusta.su — the OTHER flibusta (different content, 2026-09-26, user-found)
- flibusta.su has books NOT on flibusta.is (unique subset). Verified: CW3 RU «Психогенез душевных болезней» (АСТ 2025, 978-5-17-139315-1) = b/390112 on .su, ABSENT on .is (a/5272 has 129 books, no CW3). So .is-only author sweeps can miss RU editions (sec07 CW3 gap).
- Book/author pages curl-able (server-rendered titles). Downloads = JS POST
/lang/?b=<id>→"b":"0"= partner-locked (litres exclusive, no file) or a file domain. No static/b/<id>/<fmt>on .su (that's .is). For .su-only books, source the file via libgen/RSL by ISBN, not .su. - Jung .su author page: 31 books (a/877 is the WRONG different-person «Бартольд»; the Jung .su author id is unknown — find via the book page's author link).
11. Open Library — the "English RSL" (author works lists, full names, ISBNs)
- Tool:
tools/ol.py "author:Lastname, First"ortools/ol.py 'title:"Exact EN Title"' openlibrary.org/search.json?q=...&fields=title,author_name,publish_year,language,publisher,isbn- No key, no rate limits observed (1s between calls is polite). Aggregates all editions of a
work into one doc. Best use: resolve full EN author names from initial-only list entries
(
title:"..."→ author_name) and get the EN works list + ISBNs to feed Google Books. - RU coverage is sparse but REAL (romanized titles, e.g. «Chelovek i ego simvoly»); the
search.json?isbn=facet works for RU ISBNs too. Don't use alanguage=rusfilter — it under-reports. RU hunting still mainly RSL/libgen/flibusta/GBS SPb/Google Books.
12. Wikidata — identity anchors (life dates + OFFICIAL RU name)
- Tool:
tools/wd.py "Full Name" wbsearchentities+wbgetentities: QID, P569/P570 (birth/death), EN/DE/RU labels, RU aliases. The RU label is the single most reliable spelling for RSL/libgen queries.tools/rulabel.py(2026-09-24): QID → ru label + ru aliases (canonical Cyrillic spelling). Authority-DB route (VIAF/ISNI/GND) verified DEAD from this box: VIAF 403 + JS-app, ISNI 403, GND API gone (services.dnb.de 404s), BnF SRU 403 — Wikidata ru labels are the working channel (user idea, 2026-09-24). Real catches: Zimmer = «Циммер» not «Зиммер», Corbin = «Корбен» not «Корбин». Niche authors may have NO ru label (MacCulloch) — fall back to RSL author field.- Only useful for Wikipedia-scale names (Jung, Neumann, Kandinsky, Jaffé, …); niche authors are absent — fall back to OL title search or publisher pages.
13. archive.org — EN full texts: search, metadata, downloads ("English libgen"; user-confirmed 2026-07-16)
- Tool:
tools/ia.py search "creator:(Carl Gustav Jung)"/title "Aion"/meta <id> advancedsearch.php?q=...&output=json&fl[]=identifier,title,year,downloads,access-restricted-item,mediatype(Solr syntax; appendAND mediatype:texts— creator: search is loose, includes images/misattributions).archive.org/metadata/<id>→ file list (Text PDF / EPUB / DJVU / FB3 + sizes) — pick the file, then direct download/download/<item>/<file>works (used for sec04 OCR verification; verified: Aioncollectedworksof92cgju= CW9/2 full text, open, 344k downloads).- Uses: (1) EN full-text downloads for ✅ items (Jung CW volumes live here:
collectedworksof*); (2) EN identity/metadata; (3) content verification (OCR_text.pdf→ pdftotext, e.g. the Neumann 4-essay match).access-restricted-item: true= controlled lending (borrow, no direct download). - Item page:
archive.org/search?query=creator%3A%22C.+G.+Jung%22(web UI of the same index). CarlJungCollectedWorksitem (OPEN, found 2026-07-18) = Princeton CW set in one item: vols 1-18 (incl. 9/1, 9/2, 10=Kundalini 1996) + Jung Seminars 1 (Dream Analysis 1984), 520 (Zarathustra), 539 (Analytical Psychology 1925) + Children's Dreams 2012 + Zofingia Lectures + Synchronicity + Psychology of the Unconscious (Hinkle 1916). PDF + EPUB per volume, direct/download/CarlJungCollectedWorks/<file>works. Supersedes per-volumecollectedworksof*items (many of those are borrow-only).
14. Google Books — book page via curl OR fetch_content (API is 429 without key)
books.google.com/books?vid=ISBN<13digits>: plain curl works — 200 +<title>"T - A - Google Книги"if indexed, 404 if not. Good ISBN validator incl. RU (АСТ/Эксмо/БукСМарт hit; small-press 404). googleapis.com/books API stays 429 without a key.- Keep volume LOW: Google may block agents on sustained querying (user warning 2026-07-15). Budget: a few dozen requests per session, 2–5s random pauses, never in tight loops; treat as an auxiliary cross-check, not a bulk source. If 429/403-redirect walls appear — stop and fall back to OL/libgen/RSL.
fetch_contenton the same URL → richer metadata: full author/editor names, edition, series. Resolved Ottmann=Klaus, Goldstein=Ralph, Brutsche=Paul, Elder=George, Moon=Beverly, Killick=Katherine, Pennington/Staples, Rowland=Susan, Bolander=Karen, Acton=Mary via this or OL.- PagePlace preview PDFs (2026-07-17): many Routledge/T&F books have a public preview at
api.pageplace.de/preview/DT0400.<ISBN13>_A<n>/preview-<ISBN13>_A<n>.pdf(~20–30 pp: title page + TOC — perfect for identity/series verification when the full text is closed; find the URL via SearXNG"<EN title>" preview pdf). Used for item 12 (Evers-Fahey, 9781317219583).
15. Shop directory (HSE bookshelf34 pattern) — where RU psychoanalytic books are sold
- HSE «Книжный шкаф» (hse.ru/ma/therapy/bookshelfNN) = a DIRECTORY of shops, not a book list:
NikBook (nikbook.ru — WBS shop,
?q=search is fuzzy/works), Когито-Центр, Скифия (skifiabook.ru — publisher/shop, JS catalog, has a full price list download), marketplaces. - NikBook Jungian finds (2026-07-09): Тоцци «Активное воображение в теории, практике и обучении», Конгер «Юнг и Райх. Тело как Тень», Эдингер «Библия и психе».
- beta2alpha.ru — RU publisher (Tilda site; sells via Ozon seller beta-2-alpha); Jungian stock TBC.
16. Publisher sites — RU full texts and bibliographies
- hi-human.org (Living Human Heritage RU) — publishes book forewords/sections as pages (source of the Neumann EN-works list with RU titles).
- castalia.ru / castaliasilvasacra.ru (Клуб Касталия) — author articles + shop.
Full catalog crawl-able (2026-09-24):
/collection/all?page=N— 48/page, 27 pages = 1285 товаров, plain curl, заголовки в h3/h2/alt; кэшdata/catalogs/castalia-full-catalog-2026-09-24.md(в rutitles DB). (PDF)-твины (/product/<slug>-pdf) = прямые PDF-продажники = download-источник. Product JSON-LD: name/sku(CAS+ISBN)/price; year/ISBN-блок «Характеристики» = JS (не читается). - daimon-verlag.ch / chironpublications.com / spring-publications.com — EN Jungian presses (spring-publications unreachable from this box; use Google Books ISBN instead).
17. alib.ru — RU/KZ book marketplace, 3.2M listings (used + new, incl. small press)
- Tool:
tools/alib.py "query"(stem search! matches all inflections — unlike libgen). - Endpoint:
https://www.alib.ru/find3.php4?tfind=<query>; query and pages are CP1251 (utf8 → mojibake). Phrase mode: double quotes. Also: author-first, year range, ISBN digits, price range.>Купить<links = listing count. - Value: catches small-press / out-of-print editions shops don't carry (found: Furth 2nd ed. 2014, Turner handbook 2015, Neumann КДУ/Маниф/Питер editions). Seller titles can be wrong — verify against RSL before trusting (e.g. «Челокес и миф» Kastaalia 2018 — not in castalia.ru catalog → seller error).
- Also: alib.top (Ukraine), 33ob.ru (vinyl), amarka.ru (stamps) — sister sites.
18. Local CRW renderer (fastcrw.com, Firecrawl-compatible, localhost:3000)
- Tool:
tools/crw.py <url> [--html] [--wait N ms]— POST /v1/scrape {"url","formats":["markdown"]} - Use for JS-heavy sites curl can't read (soznanie.ast-academy.ru festival site rendered fine).
- Limits: Ozon (FAB challenge) and chitai-gorod SEARCH are API/anti-bot protected — CRW returns nothing. chitai-gorod PRODUCT pages DO render via CRW (see source 9).
- Insales shops (castalia.ru) don't need it: product data is server-side JSON-LD (curl OK).
19. NLR / РНБ (National Library of Russia) via Primo (primo.nlr.ru) + CRW
- Tool:
tools/nlr.py "query"— free-text search (all fields, words ANDed), page 1 (20 rows). - Search is client-rendered → the tool renders the
search.do?fn=search&ct=search& vl(freeText0)=QUERY&vid=07NLR_VU1&mode=Basic&initialSearch=trueURL via CRW and parses title/author/year/holding (+ doc IDs when the linked layout renders). - Full record (GOST description, incl. translators): CRW on
display.do?tabs=detailsTab&ct=display&fn=search&doc=<07NLR_LMS#########>&displayMode=full&vid=07NLR_VU1(plain curl gives only the shell — details tab is XHR). - Value: complements RSL — own St. Petersburg holdings + GOST records; author rows carry life dates and NLR auth IDs. Count line parsing is flaky — trust the parsed row list.
- Verified 2026-07-09: all 35 section-04 surnames swept (see sections/04-pictures/SWEEP-R7.md).
20. RSL direct (aleph.rsl.ru, Ex Libris Aleph) — PARTIAL
- find-a flow works: GET /F/-?func=file&file_name=find-a → parse action URL (session token) → POST func=find-b&find_code=WAU&request= → results 10/page with pagination.
- Full record: follow the result row's
full-set-setlink (session-dependent); gives translator, series, ISBN, UDC (verified: Нойманн «Амур и Психея» МИФ 2024, ISBN 978-5-00214-519-5, пер. М. Виноградова). - Quirk: find_code WTI (title) and SYS (ISBN) return EMPTY 0-byte responses; only WAU (author) reliably works so far. libgen proxy (tools/rsl.py) still covers title/ISBN queries.
- search.rsl.ru — same catalog, Yii app, no JSON API found (2026-07-09).
21. ISBN validation pipeline — tools/isbnval.py (2026-07-15)
- Per ISBN: (1) libgen.vg biblioservice
type=isbn— checksum + registered publisher group; (2) Open Librarysearch.json?isbn=— title/author/year; (3) Google Books?vid=ISBNcurl — title tag, 200/404; (4) NLR free-text for RU misses. Random pauses 2–5s, resumable. - Verdicts: VALID (found w/ matching title) / STRUCT-ONLY (real ISBN, prefix registered,
not indexed — small-press RU) / BAD-STRUCT. Output TSV:
data/isbn-validation.tsv. - Found in sec04: 2 bad-check-digit typos (5-89613-003-6, 5-263-00368-2), 2 publisher-prefix mismatches (Ленанд vs URSS; Академический проект vs Азбука-Аттикус), 2 co-publishing prefixes (Петроглиф for ЦГИ 2020; Пальмира for Касталия 2025).
- triumph.ru (РГБ ISBN search, user-provided):
https://www.triumph.ru/html/serv/find-isbn.php?isbn=<13digits>(relative to the poisk-isbn.html page!) → returns links to РГБ/МГУ/РНБ/БЕН РАН/ГПНТБ catalogs pre-filled with the ISBN. Itself no data; МГУ (nbmgu.ru) results are JS-rendered, БЕН РАН (Koha) is scrapable but small collection. NLR absence ≠ ISBN invalid. - isbnsearch.org: nice per-ISBN pages (author/publisher/year) but rate-walls after ~10 requests ("Please Verify to Continue", no recovery in 7 min) — not for batches.
22. MGU Scientific Library (nbmgu.ru) — server-rendered, curl-able (2026-07-15)
- Tool:
tools/mgu.py "query" [FIELD] [pages] [method]or--multi "AUT:X" "TIT:Y". The "JS" advanced search is a plain GET:/search/?adv=1&q1=<q>&f1=<F>&v1=<M>&cat=BOOK[&p=N]. - Fields: ANY/AUT/COA/TIT/KEY/RUB/YEA/PLA/PUB/SER/ISB/ISS/NBM; rows AND between q1..q3. Method v: 0=Слова (stemming, default) · 1=Словосочетание · 2=Начинается с · 3=Дословно.
- "Всего: N" + GOST row (title/authors/notes) + "City : Publisher, Year" + shelf code + uid link. 20 rows/page, p=0-based; out-of-range p silently falls back to page 0 (dedupe by uid!).
- ISB field NOT populated (0 hits, 3 formats × control ISBNs) — no ISBN search; use
AUT/TIT/SER/PUB.
storing.aspx?uid=page = holdings/order only (no full record). - Value: second RU catalog (complements RSL/РНБ), strong SER sweep (e.g. «Библиотека аналитической психологии» → full series list in one query), GOST rows w/ translators.
- Anchor titles in result rows contain a raw
>(title="Хранение
Заказ") — regexes must not use [^>]+ across the anchor attrs.
23. bookvoed.ru — curl-able, ISBN in page title (2026-09-20, user-provided)
https://www.bookvoed.ru/product/<slug>-<id>: HTTP 200 with plain curl;<title>containsНазвание (Автор) (ISBN)— a fast existence + ISBN check. Full specs in HTML.
24. litres.ru book pages — curl-able JSON-LD (2026-09-20, user-provided)
https://www.litres.ru/book/<cat>/<slug>-<id>/chitat-onlayn/: HTTP 200 with curl; JSON-LD has name/author/publisher (translator/ISBN/pages often absent). «Кратко» (Culture-Multur abridgments) are distinct ISBNs from the originals — don't conflate. Good for title-variant discovery.
25. royallib.com / pda.coollib.in — free RU full text, curl-able (2026-09-20, user-provided)
royallib.com/book/<author>/<title-slug>.html— 200 with curl, annotation + text pages;pda.coollib.in/b/<id>-<slug>/readp?p=N— page-by-page text. Use for CONTENT verification (e.g. Eliade «Тайные общества» = Rites and Symbols of Initiation: preface = Haskell Lectures 1956, University of Chicago — fingerprint of the EN original). Both = download supplements.
26. book.ivran.ru + opac.liart.ru — small RU library OPACs, curl-able GOST records (2026-09-20)
book.ivran.ru/book?id=<N>— «Книги сотрудников ИВ РАН» DB: full GOST rows (title, editor, city/year/pages, ISBN). Find via SearXNGsite:book.ivran.ru.- opac.liart.ru (РГБИ) — platform switched to OPAC-Global (DITM-Global) (2026-09-21, probed):
/opacg/guest session: POSTarg0=GUEST&arg1=GUESTE&TypeAccess=PayAccessto/cgiopac/opacg/opac.exe; session9096453/GUEST(hardcoded in JS).- Search = POST
/cgiopac/opacg/direct.exeFLAT fields:_service=opacfindd.FindView(orFindSize),query/body="<label> <value>",iddb=<N>,userId,session,_xsl=_version=2.7.0/1.1.0— full protocol intools/liart.pydocstring. - BLOCKER: valid search labels + iddb come from server InfoDB (admin-only for GUEST;
iddb 1–4 absent, 7+ exist). Record pages
/record/<db>/<hex>ARE curl-able (GOST in